Voice AI
A voice agent that passes in a studio can still fail its first real caller
Recognition degrades on dialect and on telephone audio. The dialogue manager then works from a degraded transcript, asks the caller to repeat themselves, and the caller speaks faster and less clearly. Each part may be acceptable alone and the system still unusable.
Scenarios
What we test against
Telecom support
Outages, plan changes, credit transfers — the calls that dominate real volume.
Banking
Balances, transfers, card problems, branch and payment questions.
Travel
Bookings, schedules, changes and cancellations.
E-commerce
Order status, returns, delivery to regions far from the capital.
Public services
Documents, appointments and administrative procedures.
General customer support
Complaints, escalations and callers who are already annoyed.
Dimensions
What we score
We score the call, not the transcript. A system can produce a clean transcript and still leave the caller without an answer.
Intent understanding
Did the agent work out what the caller wanted?
Dialect comprehension
Did it cope with the variety and the speed the caller actually used?
Speech recognition
Measured on telephone-band audio, not studio recordings.
Response relevance
Was the answer about the caller's problem?
Pronunciation
Is the synthesised voice intelligible and locally plausible?
Naturalness
Does the exchange feel like a conversation or an interrogation?
Latency
Measured as the caller experiences it, including think time.
Task completion
Did the caller end the call with what they rang for?
Safety
What the agent does when it is out of its depth.
Why it has to be people
Synthetic test audio does not reproduce the speed, reduction and code-switching of a real caller, and it never becomes impatient. We test with native speakers running realistic scenarios, on real calls.
Test your voice agent in North Africa
Real callers, real scenarios, scored across every dimension above — with the recordings and the individual scores handed back to you.