Speech Data Collection
Recorded by real Algerians, in the conditions your product ships into.
Scripted, spontaneous and in-the-wild speech across every major Algerian region, recorded to your technical specification and quality-gated before it reaches you.
What we cover
Coverage and categories
Scope any combination below, or bring your own taxonomy — the platform is configured per project.
Scripted speech
Prompt-driven utterances for ASR training, with controlled phonetic and lexical coverage.
Conversational speech
Natural two-party dialogue, including overlap, hesitation and repair phenomena.
Call-centre speech
Telephony-band recordings that mirror real support, banking and delivery calls.
Noisy-environment speech
Street, café, market, vehicle and household noise captured as it actually occurs.
Wake words & commands
High-repetition trigger phrases and short imperatives across speakers and distances.
Phone recordings
Mobile-captured audio over real Algerian networks and handsets, not studio simulations.
Accent datasets
Targeted collection by regional accent, including Algerian-accented French.
Voice-agent datasets
Turn-level audio designed for training and evaluating conversational voice agents.
Deliverables
What you receive
- WAV audio at your required sample rate and channel layout
- Per-recording metadata (dialect, region, age band, device, environment)
- Speaker-balanced distribution reports
- Anonymised speaker IDs — never personal identifiers
Outcomes
Typical specification
- Audio format
- WAV / PCM, configurable
- Sample rate
- 8 kHz – 48 kHz
- Speakers
- From 50 to 5,000+
- Utterances
- Configurable per speaker
- Regional split
- Quota-enforced by wilaya
- Age & gender
- Balanced to your distribution