LLM evaluation
Whether your model is usable by people here
Automatic metrics need a reference, and for Maghrebi varieties agreeing on the reference is most of the work. We use people instead — qualified, blind, and reviewed.
Prompt evaluation
Does the model understand what was asked, in the register it was asked in?
Response ranking
Two or more answers, ordered by native speakers working blind.
Human preference data
Preference pairs with the reasoning recorded, usable for training as well as for measurement.
Cultural adaptation
Whether an answer makes sense for someone here — the right institutions, the right places, the right assumptions.
Dialect naturalness
Whether the output reads as something a person in this country would write, or as translated Modern Standard Arabic.
Safety review
Behaviour on sensitive input, judged by people who understand the local context.
Hallucination review
Confident, fluent and wrong is the characteristic failure in under-resourced languages. It takes a native speaker to catch it.
Multilingual and mixed input
Including the code-switched input that monolingual evaluation sets never contain.
How it runs
From your outputs to a number you can defend
Your model's outputs
Sent as CSV or JSONL against a benchmark item set — ours, or one built for your product.
Native evaluators
Qualified for the variety, working blind, several per item.
Quality assurance
Evaluator work is reviewed like any other task. Ratings that fail review are excluded.
Metrics
A figure per criterion with its sample size, evaluator count and exclusions.
Report and data
The aggregates, the individual judgements behind them, and the evaluation set.
Evaluate your model in Algerian Darija
Against our benchmark, or against an item set built from the situations your users are actually in.