Custom evaluation
Test your AI with LAHJA
Tell us what you have built and what you need to know about it. We will come back with a proposed evaluation design, the number of native evaluators it needs, and what it would cost.
If you would rather read how the evaluation itself works first — blinding, randomisation, agreement, what we refuse to claim — that is on the AI evaluation page.
What happens next
You tell us what you need to know
Not a specification — the decision you are trying to make. Whether your assistant is usable in Algeria is a different question from how it compares to a competitor, and they need different evaluations.
We propose a design
How many items, which varieties, how many evaluators per item, what the criteria are, and what the figures will and will not tell you. With a price and a timeline.
Native speakers do the work
Qualified for the specific variety, working blind, several per item, with quality assurance on the evaluation itself.
You get the numbers and the data
Aggregates with their sample sizes and exclusions, the individual judgements behind them, and the evaluation set itself where the agreement allows.
What we will not do
We will not score your system with another model and call it human evaluation, and we will not publish a figure without the sample size and exclusions behind it. If an evaluation cannot answer your question, we will say so before taking the work.