by Rapidata
MethodologyAnnotator qualityUpdatesHugging Face281.1k downloadsSuggest a model or benchmarkRequest a private benchmark
© Rapidata, benchmark.ai
MethodologyAnnotator qualityUpdatesHugging Facerapidata.ai

Annotator quality

Can these human judgements be trusted? This page documents the acquisition, validation, and aggregation pipeline. Platform-wide mechanisms apply to all benchmarks; round-specific statistics apply to a single benchmark round.

Platform-wide
1. How annotators encounter tasks
Annotators reach tasks inside third-party apps and advertising contexts, where a short validation task appears in place of a conventional ad. This is how benchmark.ai collects millions of judgements per day from a broad, non-expert population rather than a small panel.
Platform-wide
2. Acquisition to graduation
A new annotator is first shown validation tasks only. Once they answer enough correctly they graduate into the global audience and their judgements begin to count. The funnel below shows the stages; no annotator contributes scored judgements before graduating.
Platform-wide
3. Validation tasks
A validation task is a task with a known correct answer, indistinguishable to the annotator from a scored task. Validation tasks are interleaved continuously, not only at onboarding, so reliability is measured throughout an annotator's activity.
Platform-wide
4. Global audience qualification
The global audience is the pool of annotators who have graduated and remain in good standing. Benchmarks draw from this audience by default, giving a broad, demographically diverse population rather than curator taste.
Platform-wide
5. Continuous reliability measurement
Each annotator carries a user score updated continuously from their validation-task performance and other signals. The user score reflects how much a given annotator's judgements can be trusted.
Platform-wide
6. Interpreting reliability thresholds
Annotators below a reliability threshold are excluded, and their judgements are removed retroactively from affected matches. Thresholds trade coverage against noise; the estimator itself is proprietary and treated here as a black box.
Platform-wide
7. Curated and custom audiences
Beyond the global audience, a curated audience restricts to demographics or expertise, and a custom audience is defined per project. Public benchmarks use the global audience unless a round's protocol states otherwise.
Platform-wide
8. Multiple independent judgements
Every pairwise match is shown to several distinct annotators (by default five). Independent judgements reduce the effect of any single noisy response; the match winner is the majority.
Platform-wide
9. Raw and weighted aggregation
We compute both a raw aggregation (equal weight per judgement) and a reliability-weighted aggregation (weighted by user score). Published leaderboards use the reliability-weighted result; the raw result is available for audit.
Platform-wide
10. Confidence estimation
Uncertainty is estimated by bootstrapping the judgement set (1,000 resamples) and reporting 95% intervals, so every score carries an explicit interval rather than a bare point estimate.
Platform-wide
11. Auditability
Judgements, prompts, and outputs are released publicly so third parties can recompute rankings and inspect the raw votes behind any match. Reproducibility is the core trust mechanism.
Round-specific
12. Annotator demographics
Self-reported demographics vary per round because the audience is sampled fresh.
Acquisition (in-app)
Onboarding validation
Graduated: global audience
Continuously measured
Curated / custom audiences

Figure 1. Acquisition-to-graduation funnel (stages, illustrative widths). Annotators contribute scored judgements only after graduating into the global audience.

The task, as shown to an annotator
Which image do you prefer?
Output A
Output B
A
B

Figure 2. Task interface. Model identities are hidden and left/right order is randomized per judgement.

Validation-task performance+Continuous behavioral signals→Reliability estimator (proprietary)→User score → judgement weight

Figure 3. Reliability estimation as a black box. Inputs and outputs are documented; the estimator's exact form is not disclosed.