Quality assurance
Quality is the product, not the final inspection.
Four layers sit between a contributor pressing record and a file entering your dataset. Each one is transparent, documented and reportable — you can audit every decision we made on your data.
The four-layer model
Every submission passes through the same four gates.
The layers run continuously while a project is live — not as a batch inspection at the end, when it is too late to re-collect.
Automatic signal checks
Every audio submission is measured the moment it lands, before a human spends a second on it. Anything outside the project specification is flagged or rejected automatically.
- File integrity and decodability
- Minimum and maximum duration bounds
- Silence ratio and leading/trailing dead air
- Clipping and peak-level detection
- Sample-rate and channel-layout conformance
- Estimated noise floor and signal-to-noise ratio
Human review
Trained reviewers judge what a signal check cannot: is this the right dialect, the right register, the right content, recorded by the right person in acceptable conditions?
- Approve, reject, or request a resubmission
- Structured rejection reasons, never free-text only
- Keyboard-driven review queue for throughput
- Project-specific rubric agreed with the client
- Full review or defined sampling rate, per contract
Gold standards & agreement
Items with a known correct answer are seeded into evaluation and annotation batches so we can measure reviewers and contributors against ground truth — not just against each other.
- Gold items blended invisibly into live batches
- Inter-annotator agreement tracked per batch
- Agreement tracked per contributor over time
- Disagreement routed to adjudication
- Rubric revised when agreement stays low
Contributor scoring
Every review updates the contributor's record. Scores govern who is eligible for sensitive, expert or high-value work — and who needs retraining before their next assignment.
- Audio quality score
- Annotation quality score
- Reliability (accepted assignments completed on time)
- Completion rate
- Approval rate
- Overall score derived from the above
Submission lifecycle
One path, four states, no shortcuts.
A submission cannot enter a deliverable without an explicit approval decision recorded against a named reviewer.
- 01
Pending
The submission is stored and queued. Nothing is counted toward delivery volume yet.
- 02
Auto-check
Deterministic signal checks run. Clear technical failures are returned immediately with the exact reason.
- 03
Manual QA
A reviewer opens the item against the project rubric and records a decision with a structured reason code.
- 04
Approved / Rejected / Resubmit
Approved items enter the deliverable and the earnings ledger. Rejected items do not. Resubmit returns the task with actionable feedback.
Resubmission returns the task to the contributor with the specific failing check, so the second attempt fixes the actual problem.
Contributor scoring
Six dimensions, updated after every review.
Scores are visible to the contributor. Nobody is penalised by a number they cannot see or a rule they were never told.
Score dimensions
- Audio quality
- Signal metrics and reviewer judgement on recordings
- Annotation quality
- Accuracy against rubric, gold items and adjudication
- Reliability
- Accepted work delivered inside the agreed window
- Completion rate
- Assignments finished versus abandoned
- Approval rate
- Share of submissions accepted at review
- Overall score
- Composite used to gate access to sensitive work
QA report
What arrives with your dataset.
Every delivery ships with a quality report. If a number in it looks wrong, we can trace it back to individual submissions and reviewer decisions.
Pass rate
Accepted submissions as a share of total submissions, broken down by batch and by task type.
Rejection reasons
A counted breakdown of why items failed — background noise, clipping, wrong dialect, off-script, duration, annotation error.
Agreement statistics
Inter-annotator agreement and gold-standard accuracy for every evaluation or annotation batch in the delivery.
Per-region coverage
Delivered volume against the agreed regional, dialect and demographic quotas, so you can see exactly where the corpus is thin.
What we deliberately do not do — yet
We do not use machine-learning models to score, rank or auto-reject contributor work at this stage. That is a deliberate choice, not a gap we are hiding. An ML quality classifier is difficult to explain to a client under audit, difficult to justify to a contributor whose work it rejects, and prone to penalising exactly the accents and dialects we exist to represent.
Instead every automatic check we run is a threshold on a measured signal — deterministic, reproducible and inspectable. If a file is rejected, we can tell you the measured value, the configured bound and the timestamp. Model- assisted triage may be introduced later, and if it is, it will be disclosed in the QA report and will never be the sole basis for a rejection.
Want the rubric before you commit?
Send us your acceptance criteria and we will map them onto the QA pipeline, including sampling rate, gold-item density and the exact report you receive.
Request a Proposal