Harbor Realm

Harbor’s flagship benchmark program

Built for real agent reliability, not just launch-day charts. Harbor Realm evaluates frontier models on expert-authored tasks, layered HITL gold labels, and stage-level error analysis.

Tracks

Robo-Edge

In progress

Embodied perception tasks across RGB-D, IMU, and time-linked annotations.

Align

Draft

Cross-modal alignment for video, audio, and temporal consistency checks.

Methodology

  1. 1

    Stage 1: Transcription

    Read sensor values and map them to the right fields and timelines.

  2. 2

    Stage 2: Intermediate structure

    Build contact graphs, object relations, and temporal structure.

  3. 3

    Stage 3: Final decisions

    Make downstream labels that depend on long chains of earlier decisions.

Blind comparison method

Realm scores pairwise judgments: model A vs model B on the same expert-authored gold task. Reviewers do not see model names. We report win/loss on the specific task, not a star rating.

This page is a method, not a completed league table. No scores are published until a run is finished with reproducible settings.

Cite Harbor Realm

Harbor Realm: a staged benchmark for embodied and multimodal agent reliability on expert-authored gold tasks.

https://harborml.com/benchmarks/harbor-realm

Publishing policy

Harbor publishes benchmark numbers only from completed runs with reproducible settings. Until runs are complete, status is shown as in-progress or draft instead of placeholder scores.