Skip to content

Services

Managed human data services for North African AI.

Every service below is delivered end-to-end by LAHJA AI: we scope the specification, recruit and qualify the workforce, run collection, quality-assure the output and hand over a structured dataset.

Speech Data Collection

Recorded by real Algerians, in the conditions your product ships into.

Scripted, spontaneous and in-the-wild speech across every major Algerian region, recorded to your technical specification and quality-gated before it reaches you.

  • WAV audio at your required sample rate and channel layout
  • Per-recording metadata (dialect, region, age band, device, environment)
  • Speaker-balanced distribution reports

Transcription

Orthography decisions made by people who speak the dialect.

Verbatim and clean-read transcription for Arabic, Algerian Darija, French, Arabic/French code-switching and Tamazight — with a documented, consistent transcription convention.

  • Time-aligned or utterance-level transcripts
  • Documented orthography and tagging conventions
  • Code-switch span annotation

AI Response Evaluation

Does your model sound right to someone from Oran?

Structured human judgement of model output across naturalness, accuracy, relevance, cultural appropriateness, dialect comprehension, safety and hallucination — with calibrated evaluators and documented rubrics.

  • Per-item scores with evaluator rationale
  • Rubric definitions and calibration set
  • Agreement statistics and outlier analysis

RLHF & Preference Data

Pairwise preference from people whose language you're aligning to.

A/B preference collection with structured reasons, built for reward modelling and preference optimisation in Darija, Arabic, French and code-switched contexts.

  • Pairwise judgements: A better · B better · Equal · Both bad
  • Structured reason codes plus free-text rationale
  • Annotator calibration and agreement metrics

Voice AI Testing

Real humans testing voice agents under real Maghrebi conditions.

We put your voice agent in front of real Algerian speakers — on real phones, in real noise, with real accents — and report exactly where it breaks.

  • Per-scenario pass/fail with recorded evidence
  • Failure categorisation by cause
  • Severity-ranked defect list

Expert AI Evaluation

Doctors, lawyers and engineers judging domain output in their own language.

Credential-verified professionals evaluate model output in medicine, law, engineering, finance, accounting, education, science and technology — in Arabic, French or Darija.

  • Expert judgements with credential provenance
  • Domain-specific rubrics
  • Risk flags for unsafe or non-compliant guidance