Skip to content

Quality assurance

Quality is the product, not the final inspection.

Four layers sit between a contributor pressing record and a file entering your dataset. Each one is transparent, documented and reportable — you can audit every decision we made on your data.

The four-layer model

Every submission passes through the same four gates.

The layers run continuously while a project is live — not as a batch inspection at the end, when it is too late to re-collect.

Layer 1

Automatic signal checks

Every audio submission is measured the moment it lands, before a human spends a second on it. Anything outside the project specification is flagged or rejected automatically.

  • File integrity and decodability
  • Minimum and maximum duration bounds
  • Silence ratio and leading/trailing dead air
  • Clipping and peak-level detection
  • Sample-rate and channel-layout conformance
  • Estimated noise floor and signal-to-noise ratio
Layer 2

Human review

Trained reviewers judge what a signal check cannot: is this the right dialect, the right register, the right content, recorded by the right person in acceptable conditions?

  • Approve, reject, or request a resubmission
  • Structured rejection reasons, never free-text only
  • Keyboard-driven review queue for throughput
  • Project-specific rubric agreed with the client
  • Full review or defined sampling rate, per contract
Layer 3

Gold standards & agreement

Items with a known correct answer are seeded into evaluation and annotation batches so we can measure reviewers and contributors against ground truth — not just against each other.

  • Gold items blended invisibly into live batches
  • Inter-annotator agreement tracked per batch
  • Agreement tracked per contributor over time
  • Disagreement routed to adjudication
  • Rubric revised when agreement stays low
Layer 4

Contributor scoring

Every review updates the contributor's record. Scores govern who is eligible for sensitive, expert or high-value work — and who needs retraining before their next assignment.

  • Audio quality score
  • Annotation quality score
  • Reliability (accepted assignments completed on time)
  • Completion rate
  • Approval rate
  • Overall score derived from the above

Submission lifecycle

One path, four states, no shortcuts.

A submission cannot enter a deliverable without an explicit approval decision recorded against a named reviewer.

  1. 01

    Pending

    The submission is stored and queued. Nothing is counted toward delivery volume yet.

  2. 02

    Auto-check

    Deterministic signal checks run. Clear technical failures are returned immediately with the exact reason.

  3. 03

    Manual QA

    A reviewer opens the item against the project rubric and records a decision with a structured reason code.

  4. 04

    Approved / Rejected / Resubmit

    Approved items enter the deliverable and the earnings ledger. Rejected items do not. Resubmit returns the task with actionable feedback.

Resubmission returns the task to the contributor with the specific failing check, so the second attempt fixes the actual problem.

Contributor scoring

Six dimensions, updated after every review.

Scores are visible to the contributor. Nobody is penalised by a number they cannot see or a rule they were never told.

Keyboard-driven review queueReviewed continuously, not at hand-off

Score dimensions

Audio quality
Signal metrics and reviewer judgement on recordings
Annotation quality
Accuracy against rubric, gold items and adjudication
Reliability
Accepted work delivered inside the agreed window
Completion rate
Assignments finished versus abandoned
Approval rate
Share of submissions accepted at review
Overall score
Composite used to gate access to sensitive work

QA report

What arrives with your dataset.

Every delivery ships with a quality report. If a number in it looks wrong, we can trace it back to individual submissions and reviewer decisions.

Pass rate

Accepted submissions as a share of total submissions, broken down by batch and by task type.

Rejection reasons

A counted breakdown of why items failed — background noise, clipping, wrong dialect, off-script, duration, annotation error.

Agreement statistics

Inter-annotator agreement and gold-standard accuracy for every evaluation or annotation batch in the delivery.

Per-region coverage

Delivered volume against the agreed regional, dialect and demographic quotas, so you can see exactly where the corpus is thin.

What we deliberately do not do — yet

We do not use machine-learning models to score, rank or auto-reject contributor work at this stage. That is a deliberate choice, not a gap we are hiding. An ML quality classifier is difficult to explain to a client under audit, difficult to justify to a contributor whose work it rejects, and prone to penalising exactly the accents and dialects we exist to represent.

Instead every automatic check we run is a threshold on a measured signal — deterministic, reproducible and inspectable. If a file is rejected, we can tell you the measured value, the configured bound and the timestamp. Model- assisted triage may be introduced later, and if it is, it will be disclosed in the QA report and will never be the sole basis for a rejection.

DeterministicTransparentAuditable

Want the rubric before you commit?

Send us your acceptance criteria and we will map them onto the QA pipeline, including sampling rate, gold-item density and the exact report you receive.

Request a Proposal