AI Response Evaluation
Does your model sound right to someone from Oran?
Structured human judgement of model output across naturalness, accuracy, relevance, cultural appropriateness, dialect comprehension, safety and hallucination — with calibrated evaluators and documented rubrics.
What we cover
Coverage and categories
Scope any combination below, or bring your own taxonomy — the platform is configured per project.
Naturalness
Does it read like a human from this region wrote it?
Accuracy
Is the factual content correct for the local context?
Relevance
Does it answer what was actually asked?
Cultural appropriateness
Register, religion, gender, formality and local norms.
Dialect comprehension
Did the model understand the Darija input at all?
Safety
Harmful, unsafe or legally problematic output.
Hallucination detection
Invented institutions, procedures, prices or place names.
Preference ranking
Ordered preference across multiple candidate responses.
Deliverables
What you receive
- Per-item scores with evaluator rationale
- Rubric definitions and calibration set
- Agreement statistics and outlier analysis
- Failure taxonomy with representative examples
Outcomes
Typical specification
- Rubric
- Yours, or co-designed with our linguists
- Scale
- Binary, Likert or ordinal ranking
- Redundancy
- 1–5 evaluators per item
- Evaluator tier
- General, linguist or domain expert