Skip to content

Speech Data Production

Speech data built for languages and dialects AI still struggles to understand.

Algerian, Moroccan, Tunisian and Libyan Arabic. Modern Standard Arabic, Kabyle and other Tamazight varieties, French and English. Recorded by people who grew up speaking them, reviewed by people who can hear the difference, and delivered with the hours it actually contains.

Read speech

Native speakers read prompts you supply. The most controllable corpus you can commission, and the fastest to scale — with the per-speaker caps that stop it becoming a recording of ten people.

Spontaneous speech

A topic and nothing else. People talking the way they actually talk — faster, less grammatical, full of the repairs and hesitations a model trained on read speech has never heard.

Conversations

Two speakers, a supplied situation, and no idea who each other are. Each records their own side, so what you receive is speaker-separated audio of a real conversation.

Code-switching

Speech that moves between Darija and French mid-sentence, because that is how most people here actually speak. Tagged to your schema, not ours.

Wake words and commands

One phrase, hundreds of speakers, many environments — plus the near-misses that must not trigger it. Progress counted by distinct speaker, never by recording.

Transcription

Verbatim, clean verbatim or normalised, by people who speak the variety. Darija has no standard orthography; we agree the convention with you and write it into the manifest.

How it works

Six steps, and one of them is a person at LAHJA.

  1. 01

    Tell us what you need

    What kind of speech, in which variety, from how many different people, and what you need to be able to do with it.

  2. 02

    We design the collection

    Prompts, environments, devices, lengths, quality control and transcription. Versioned and signed off, so an argument later has a document to settle against.

  3. 03

    Upload your prompts

    CSV or JSONL. Every row that cannot be used comes back with its line number — and if any row is wrong, the whole file is refused rather than imported with holes in it.

  4. 04

    We confirm the scope

    Nobody is asked to record until we have agreed the design, the speakers and the price with you. That gate is a person, not a setting — and no API key can open it.

  5. 05

    Speakers record

    Recruited, qualified, paid per recording, and told to speak normally rather than carefully. Every recording passes an automatic check and then a person.

  6. 06

    You receive a version

    Audio, transcripts, metadata and a manifest with checksums — plus the hours it contains, the speaker spread, and what the corpus does not evidence.

Method

What we will not do to make a corpus look bigger.

A speech dataset is only worth what its method is worth, so ours is written down and published rather than described.

Hours are measured, not estimated

Every hour we report is the sum of durations of accepted audio. Not the target, not the count times an average. If you ordered forty and we collected thirty-one, we tell you thirty-one before it ships.

Speaker concentration is reported

The share of the audio from the single most prolific speaker — a figure nobody asks for. A corpus where one person contributed a fifth is a recording of that person with other voices in it, whatever the quotas say.

Consent decides what can be delivered

What you may do with the corpus decides what we must collect from every speaker, so it is settled before recording rather than after. A shortfall stops the build — we never quietly ship you a shorter dataset.

We do not claim phonetic balance. We report character and category coverage, which is a real and useful thing to know about a script set, and we label it as what it is. A corpus that genuinely needs phonetic coverage needs a lexicon and a linguist.

Speakers appear in your dataset as a pseudonym that exists inside your corpus and nowhere else — stable enough to reason about speaker balance, useless for joining two datasets together. There is no access level, contract term or API parameter that produces a name.

A licence to train a model is not a licence to reproduce somebody's voice. Building a synthetic voice needs a separate agreement with us and a separate, explicit consent from every speaker — asked on its own, refusable without affecting their other work. No contract term can supply that consent, because it is not ours to give.

Collect Algerian Darija data

A pilot is usually fifty speakers, a couple of hours and three regions — scripted and spontaneous, with audio QA and full metadata. Enough to see whether the data does what you need before you commission a hundred hours of it. We scope it with you and tell you what it costs before anybody records.