Skip to content

Voice AI Testing Lab

Your voice agent works in the demo. We find out what happens when an Algerian calls it.

Real speakers, real phones, real rooms with real noise. They play a situation, talk the way they normally talk, and tell us what the system did — and we tell you, with the sessions behind every number.

Live voice agents

A native speaker dials your number, plays a real situation and reports what happened. The closest thing to a customer you can buy — and the only test that catches what a synthetic benchmark cannot.

IVR and contact centres

Menu and routing behaviour, longer multi-turn calls, escalation paths, and the hand-off to a person — tested end to end rather than up to the point where it gets interesting.

Speech recognition and synthesis

You supply the audio and your transcript; native speakers review it, and where you also supply a reference we compute word and character error rates. Synthesis is scored on naturalness, pronunciation and whether the accent fits the variety it claims.

Code-switching

Callers who start in Darija, drop into French for a number, and switch back. It is how people here actually speak on the phone, and it is where most agents come apart.

Latency as the caller feels it

Where your system reports per-turn timings we report the median and the 95th percentile. Where it does not, we say “not measured” rather than timing a person with a stopwatch and calling it your latency.

Regression between versions

The same scenarios against v2, with the same speakers where we can. Each comparison is frozen when it is run, so a verdict you acted on cannot change underneath you.

How it works

Six steps, and one of them is a person at LAHJA.

  1. 01

    Tell us how to reach it

    A number, a web address or an app, plus any test credential. Credentials are stored sealed and shown only to an assigned tester, only while their session is open.

  2. 02

    Write the scenarios

    What a caller is trying to do, and the sentence that decides whether the agent did its job. Start from our Algerian library or import your own.

  3. 03

    Shape the panel

    Ten from Algiers, ten from Oran, ten from Constantine; mobile and speakerphone; quiet rooms and busy streets. A panel is a shape, not a list of names.

  4. 04

    We confirm the scope

    Nothing calls your system until we have agreed the scope, the panel and the price with you. That gate is a person, not a setting — and no API key can open it.

  5. 05

    Speakers run the sessions

    They are told to speak normally and not to help the agent along. A test where the tester did the agent's work for it tells you nothing about your callers.

  6. 06

    Read what happened

    Task success, intent understanding, latency, repeats, the failure breakdown, the weakest scenarios, and a report you can send to somebody who was not in the room.

Method

What we will not do to make a number look better.

A voice test is only worth what its method is worth, so ours is written down and published rather than described.

Your failure is not our failure

A session where the agent answered and did badly is a finding. A session where our tester's connection dropped is not, and it is excluded and re-run. Keeping those two apart is most of what this product is.

Every session accounted for

Counted, pending, excluded by reason, being re-run — the columns add up to the total. You can check the denominator of every rate we show you, which is the only way a rate means anything.

Recording is off unless three things are true

The platform allows it, your project asks for it, and that speaker consented. Silence is not consent on any of the three, and a speaker who declines still runs the session and is still paid.

Speakers appear in your results as a pseudonym that exists inside your test and nowhere else — stable enough to see that one person ran four scenarios, useless for joining two tests together. There is no access level, contract term or API parameter that produces a name, and no endpoint that serves you a recording.

Find out before your customers do

A pilot is usually five scenarios and forty sessions across three cities — enough to see where an agent is weak, small enough to run in a week. We scope it with you and tell you what it costs before anybody dials.