Skip to content

Speech

Recordings from people who speak the way your users speak

Collected on real devices in real conditions, transcribed and reviewed by native speakers, and described with the metadata that makes a corpus usable.

Scripted speech

Prompted read-aloud across a controlled text set, for coverage you can specify in advance.

Spontaneous speech

Unscripted talk on a topic — the register people actually use, with the reduction and elision that come with it.

Conversational speech

Two speakers, natural turn-taking, overlap and interruption.

Call-centre style

Telephone-band audio in support scenarios, which is where most voice systems meet their users.

Voice commands

Short directed utterances for wake-word and command recognition.

ASR test sets

Held-out evaluation data with verified transcripts, balanced across regions.

Speaker diversity

Balanced across region, age band and gender, as far as the recruited pool allows — and reported honestly where it does not.

Code-switched speech

Darija with French, annotated at the segment level by speakers who use both.

Samples

Listen to the data

Published recordings only. Each comes from a contributor who agreed separately to public use, and speakers are described by region and age band — never by name.

No public samples yet

Recordings appear here once a contributor has given public-use consent and an administrator has published them. We do not publish project recordings automatically.

Build a North African speech dataset

Tell us the varieties, the regions, the speaking style and the volume, and we will come back with a collection plan and a price.