Evals
Rubric-based evaluation of model behaviour.
Evaluation work scores model outputs against a published rubric: preference ranking, safety, domain correctness, scenario coverage, missed detections, night and weather edge cases, occlusion. You work in the expert workspace. You do not upload phone video to Collect unless the invitation says that is the sample.
What an invitation contains
- Scope and rubric version
- Estimated effort and timeline
- Confidentiality / NDA level
- Compensation structure (stated on the invite—not guessed from Collect)
Decline without penalty if you cannot meet the rubric. Accepting means you can deliver original human judgment. Bots, bulk templates, and identity sharing are bans. Assistive tools are allowed only when the rubric says so.
Quality
Rework follows the rubric version on the batch. Appeals: seven days, with evidence tied to that version. Fraud and confidentiality breaches are final. See Contributor standards.