Harbor Realm
Built for real agent reliability, not just launch-day charts. Harbor Realm evaluates frontier models on expert-authored tasks, layered HITL gold labels, and stage-level error analysis.
Tracks
Robo-Edge
In progressEmbodied perception tasks across RGB-D, IMU, and time-linked annotations.
Align
DraftCross-modal alignment for video, audio, and temporal consistency checks.
Methodology
Stage 1: Transcription
Read sensor values and map them to the right fields and timelines.
Stage 2: Intermediate structure
Build contact graphs, object relations, and temporal structure.
Stage 3: Final decisions
Make downstream labels that depend on long chains of earlier decisions.
Blind comparison method
Realm scores pairwise judgments: model A vs model B on the same expert-authored gold task. Reviewers do not see model names. We report win/loss on the specific task, not a star rating.
This page is a method, not a completed league table. No scores are published until a run is finished with reproducible settings.
Cite Harbor Realm
Harbor Realm: a staged benchmark for embodied and multimodal agent reliability on expert-authored gold tasks.
https://harborml.com/benchmarks/harbor-realm
Publishing policy
Harbor publishes benchmark numbers only from completed runs with reproducible settings. Until runs are complete, status is shown as in-progress or draft instead of placeholder scores.