Evaluation Studio
Human evaluation for AI that needs to understand the real world.
Automated benchmarks tell you how your model does on text somebody else collected. We tell you how it does on the language your users actually speak — measured by people who speak it, with the method written down and the sample size beside every number.
LLM evaluation
Native speakers score your model's answers on the criteria you choose — correctness, helpfulness, naturalness, cultural fit — or compare two models side by side without being told which is which.
Model comparison
Two systems or eight, on the same prompts, fully blinded. Every unordered pair is compared in turn, so what comes out is a win rate a person can check rather than a ranking somebody asserted.
Voice AI
Testers call or use your agent, play a scenario, and report what happened: did it understand the dialect, did they achieve the goal, how long were the pauses, did the voice sound like a person.
Retrieval and RAG
Evaluators read the passages your system retrieved alongside its answer, and judge whether the answer is actually supported by them — not whether it happens to be right.
Multilingual and code-switching
Prompts that move between Darija and French mid-sentence, because that is how most people here actually write. Speech recognition with word and character error rates computed against your own references.
Regional expertise
Evaluators qualified in the specific variety — Algiers, Oran, Constantine, Tamazight — because a system that scores well overall can be quietly weak in one region, and an average will not show you that.
How it works
Six steps, and one of them is a person at LAHJA.
- 01
Add your systems
Register each model or version. Versions are never overwritten, so “did v2 get better” stays a question with an answer.
- 02
Upload your test set
CSV or JSONL — whatever your harness already writes. Your own identifiers come back unchanged, and any extra column becomes a way to slice the results.
- 03
Choose what to measure
Start from one of our rubrics or write your own. Every criterion carries the guidance the evaluator will read.
- 04
Configure the panel
Language, variety, region, qualifications, how many people judge each item, and how blinded they are.
- 05
We confirm the scope
Nothing reaches a contributor until we have agreed the scope and the price with you. That gate is a person, not a setting.
- 06
Read the results
A comparison view, the slices where a system is weak, an agreement figure, and a report you can send to somebody who was not in the room.
Method
What we will not do to make a number look better.
An evaluation is only worth what its method is worth, so ours is written down and published rather than described.
Blinded by default
On a comparison, evaluators see “System A” and “System B” and nothing else. No provider name, no model name. An evaluator who recognises a house style stops judging the answer.
Reviewed, then counted
Every judgement goes through the same quality pipeline as the rest of our work. A judgement a reviewer rejected is not a quieter data point — it is not counted at all. The person who made it is still paid.
Every figure with its sample
No confidence intervals, no p-values, no “significantly better”. The number, what it was computed from, and what was missing. A slice too thin to read says so instead of printing a number.
Evaluators appear in your results only as a pseudonym that exists inside your evaluation and nowhere else. There is no access level, contract term or API parameter that produces a name — not because we withhold it, but because the queries that would return one do not exist.
Test your AI in Algerian Darija
A pilot is usually a couple of hundred prompts and three evaluators each — enough to see a real difference, small enough to run in days. We will scope it with you and tell you what it costs before anything starts.