Methodology
How the work is actually done
Who evaluates, how they are qualified, what they can and cannot see, what happens to work that fails review, and what our published figures rest on.
Who does the work
Contributors apply, complete a profile, and pass a qualification test for the specific variety they will work in. Evaluating Algerian Darija requires a speaker of Algerian Darija; for some work it requires a speaker from a particular region. Qualifications are per variety and per task type, and they can be withdrawn.
How work is assigned
Units are claimed rather than pushed, subject to eligibility rules: the qualification, the equipment check where one applies, project-specific access, and how much of the project one person has already done. Each claim has a deadline, and an abandoned claim returns to the pool.
Blind evaluation
Evaluators see the response and not the system that produced it. The identity is stored on the unit so results remain traceable, but the field is hidden from the person doing the rating. Knowing the vendor changes the score, and we have no interest in discovering that the hard way.
Several evaluators per item
A run specifies how many independent judgements each item needs, typically three. Every individual judgement is stored, not only the aggregate, so the spread stays visible — disagreement is information about the item, and averaging it away hides that.
Quality assurance
Evaluation work goes through the same review pipeline as any other task on the platform: automatic checks where they apply, then human review against a rubric, with hidden known-answer units mixed into the stream. A rating that fails review is excluded from the aggregate — it was judged wrong, and counting it would defeat the review that caught it.
Consent
Contributors accept terms covering the platform, data processing and the use of their work as AI training and evaluation data. Voice recording carries its own consent. Publishing a recording publicly — on this site, in a talk — requires a further, separate consent, because agreeing that a client may train on your voice is not agreeing that strangers may listen to it.
Anonymisation
Delivered datasets identify speakers by a per-project pseudonym. No name, no contact details, no contributor identifier that survives outside the platform. Demographic fields leave the platform only when the project explicitly allows it, and public sample pages describe speakers by region and age band only.
Benchmark versioning
A benchmark edition is frozen. Changing the items or the method opens a new version rather than editing the old one, so a result published against v0.1 still means what it meant. Every published result names the version and the methodology version it was produced under.
What our figures do not tell you
We publish means with their sample sizes, evaluator counts and exclusions, and a plainly stated agreement figure. We do not claim statistical significance from a few dozen items and three evaluators, we do not present model-scored results as human evaluation, and we do not describe an illustrative run as a measurement. Where no verified run exists, the page says so instead of showing a number.
Peer review
None of this has been through academic peer review, and we do not describe it as research in that sense. It is the operating practice of a data company, published so that a client can judge whether our numbers are worth anything.
Have us evaluate your system this way
Same qualification, same blind design, same quality assurance — against your product and the varieties your users speak.