Skip to content
← All services

Transcription

Orthography decisions made by people who speak the dialect.

Verbatim and clean-read transcription for Arabic, Algerian Darija, French, Arabic/French code-switching and Tamazight — with a documented, consistent transcription convention.

What we cover

Coverage and categories

Scope any combination below, or bring your own taxonomy — the platform is configured per project.

Modern Standard Arabic

Standard orthography with optional diacritisation.

Algerian Darija

Arabic-script convention (with Arabizi normalisation where required), applied consistently across annotators.

French

Native and Algerian-accented French, verbatim or clean-read.

Arabic/French code-switching

Span-level language tagging so training data reflects how people actually speak.

Tamazight

Kabyle, Chaoui, Mozabite and Tuareg in Latin or Tifinagh script, per your spec.

Deliverables

What you receive

  • Time-aligned or utterance-level transcripts
  • Documented orthography and tagging conventions
  • Code-switch span annotation
  • Inter-annotator agreement reporting on a sampled subset

Outcomes

Consistent labels across thousands of hoursCode-switching your model can actually learn fromTranscripts that match how Algerians write, not textbook Arabic

Typical specification

Output
TXT, JSON, CSV, SRT/VTT
Granularity
Utterance or word-level timings
Tagging
Speaker, noise, code-switch, disfluency
QA
Second-pass review on every batch
Scope this project