Transcription
Orthography decisions made by people who speak the dialect.
Verbatim and clean-read transcription for Arabic, Algerian Darija, French, Arabic/French code-switching and Tamazight — with a documented, consistent transcription convention.
What we cover
Coverage and categories
Scope any combination below, or bring your own taxonomy — the platform is configured per project.
Modern Standard Arabic
Standard orthography with optional diacritisation.
Algerian Darija
Arabic-script convention (with Arabizi normalisation where required), applied consistently across annotators.
French
Native and Algerian-accented French, verbatim or clean-read.
Arabic/French code-switching
Span-level language tagging so training data reflects how people actually speak.
Tamazight
Kabyle, Chaoui, Mozabite and Tuareg in Latin or Tifinagh script, per your spec.
Deliverables
What you receive
- Time-aligned or utterance-level transcripts
- Documented orthography and tagging conventions
- Code-switch span annotation
- Inter-annotator agreement reporting on a sampled subset
Outcomes
Typical specification
- Output
- TXT, JSON, CSV, SRT/VTT
- Granularity
- Utterance or word-level timings
- Tagging
- Speaker, noise, code-switch, disfluency
- QA
- Second-pass review on every batch