Skip to content

LLM evaluation

Whether your model is usable by people here

Automatic metrics need a reference, and for Maghrebi varieties agreeing on the reference is most of the work. We use people instead — qualified, blind, and reviewed.

Prompt evaluation

Does the model understand what was asked, in the register it was asked in?

Response ranking

Two or more answers, ordered by native speakers working blind.

Human preference data

Preference pairs with the reasoning recorded, usable for training as well as for measurement.

Cultural adaptation

Whether an answer makes sense for someone here — the right institutions, the right places, the right assumptions.

Dialect naturalness

Whether the output reads as something a person in this country would write, or as translated Modern Standard Arabic.

Safety review

Behaviour on sensitive input, judged by people who understand the local context.

Hallucination review

Confident, fluent and wrong is the characteristic failure in under-resourced languages. It takes a native speaker to catch it.

Multilingual and mixed input

Including the code-switched input that monolingual evaluation sets never contain.

How it runs

From your outputs to a number you can defend

  1. Your model's outputs

    Sent as CSV or JSONL against a benchmark item set — ours, or one built for your product.

  2. Native evaluators

    Qualified for the variety, working blind, several per item.

  3. Quality assurance

    Evaluator work is reviewed like any other task. Ratings that fail review are excluded.

  4. Metrics

    A figure per criterion with its sample size, evaluator count and exclusions.

  5. Report and data

    The aggregates, the individual judgements behind them, and the evaluation set.

Evaluate your model in Algerian Darija

Against our benchmark, or against an item set built from the situations your users are actually in.