Training data with the quality measured
Labelling is easy. Labelling consistently is not. We write the guidelines, measure how often our annotators agree, settle the disagreements, and hand back a dataset with its own quality report.
- Target inter-annotator agreement (κ)
- ≥ 0.85
- Ships with a quality report
- Every set
- Label, review, adjudicate
- 3-pass
What the team actually does
Annotation guideline authoring
The guideline is the product. Edge cases, worked examples and a decision tree written before annotation starts, versioned as it evolves.
Text, image and audio labelling
Classification, entity extraction, sentiment and intent, bounding boxes and segmentation, transcription and diarisation.
Model output evaluation
Human grading of LLM and ML outputs against a rubric — helpfulness, accuracy, tone, safety — with scores and written failure notes.
Preference ranking
Pairwise and n-way ranking for preference tuning, run by trained raters with calibration rounds and drift checks.
Adversarial testing
Structured red-teaming for prompt injection, jailbreaks, hallucination and unsafe output, documented as reproducible cases.
Golden sets and regression suites
A held-out evaluation set with expected outputs, so every model or prompt change can be measured rather than eyeballed.
What is included
- Versioned annotation guideline
- Calibration rounds before production labelling
- Three-pass workflow: label, review, adjudicate
- Inter-annotator agreement reporting
- Per-batch quality report
- Golden set and regression suite maintenance
Channels
- Annotation platform
- Secure data room
- Batch delivery
- Evaluation dashboards
Tools we work in
- Label Studio
- Prodigy
- CVAT
- Argilla
- Custom review harnesses
Your instance, your data, named accounts. If your stack is not listed, we learn it — the tool is rarely the hard part.
Service-level targets
| Metric | Target |
|---|---|
| Inter-annotator agreement (Cohen's κ) | ≥ 0.85 |
| Post-adjudication accuracy | ≥ 98% |
| Batch turnaround | Per agreed SLA |
| Guideline version currency | 100% |
Targets, not guarantees. Your current performance is measured in week one and the agreed target is written into your contract.
Questions about this service
If yours is not here, email us. A person answers, usually the same day.
Work happens in a controlled environment: no local downloads, named accounts, access logged, NDAs at individual level, and — where required — a dedicated pod that works on nothing else. Data handling terms are agreed in writing before the first batch.
Yes. We build the rubric with you, calibrate raters against a seed set, then grade at volume and return scores with written failure notes, so you get a regression suite rather than a one-off opinion.

Let's talk about ai data services
Half an hour to walk through the job, then a written scope with volumes, targets and a price. Yours whether or not we work together.