Skip to content
← All services

RLHF & Preference Data

Pairwise preference from people whose language you're aligning to.

A/B preference collection with structured reasons, built for reward modelling and preference optimisation in Darija, Arabic, French and code-switched contexts.

What we cover

Coverage and categories

Scope any combination below, or bring your own taxonomy — the platform is configured per project.

Pairwise comparison

Two candidate responses, one judgement, one reason set.

Reason taxonomy

More natural · more accurate · understands dialect · more relevant · safer · better French/Arabic usage.

Tie handling

Explicit Equal and Both bad options so weak pairs are not forced.

Free-text rationale

Optional written justification for every judgement.

Multi-rater redundancy

Configurable overlap to measure agreement and filter noise.

Deliverables

What you receive

  • Pairwise judgements: A better · B better · Equal · Both bad
  • Structured reason codes plus free-text rationale
  • Annotator calibration and agreement metrics
  • JSONL export ready for reward-model training

Outcomes

Reward models that prefer what Algerians actually preferReduced reward hacking on dialect promptsAuditable preference data with reasons attached

Typical specification

Export
JSONL / CSV / Parquet-ready
Overlap
Configurable per batch
Throughput
Thousands of pairs per week
Quality gate
Gold-standard items seeded per batch
Scope this project