Speech
Recordings from people who speak the way your users speak
Collected on real devices in real conditions, transcribed and reviewed by native speakers, and described with the metadata that makes a corpus usable.
Scripted speech
Prompted read-aloud across a controlled text set, for coverage you can specify in advance.
Spontaneous speech
Unscripted talk on a topic — the register people actually use, with the reduction and elision that come with it.
Conversational speech
Two speakers, natural turn-taking, overlap and interruption.
Call-centre style
Telephone-band audio in support scenarios, which is where most voice systems meet their users.
Voice commands
Short directed utterances for wake-word and command recognition.
ASR test sets
Held-out evaluation data with verified transcripts, balanced across regions.
Speaker diversity
Balanced across region, age band and gender, as far as the recruited pool allows — and reported honestly where it does not.
Code-switched speech
Darija with French, annotated at the segment level by speakers who use both.
Samples
Listen to the data
Published recordings only. Each comes from a contributor who agreed separately to public use, and speakers are described by region and age band — never by name.
No public samples yet
Recordings appear here once a contributor has given public-use consent and an administrator has published them. We do not publish project recordings automatically.
Build a North African speech dataset
Tell us the varieties, the regions, the speaking style and the volume, and we will come back with a collection plan and a price.