Skip to content

Benchmarks

What a system actually understands

Each benchmark is a versioned set of items with a published method and a private answer key. A result appears here only when an evaluation was run and an administrator published it.

Demonstration material. This is an illustrative benchmark, not a statistically validated research benchmark. No AI system has been independently measured against it.

No benchmark has been published yet.

Benchmark your own system

We run the same process against a client's model, speech API or voice agent, with native evaluators and the methodology published alongside the numbers.