What we deliver
Saolabs delivers data that a model can learn from or be measured against. That covers labelled inputs, such as images, video, audio and sensor logs, and written supervision, such as expert answers, rankings and step-by-step reasoning.
The work is done by people who practise in the field the data describes. A dosing question is reviewed by someone with medical training. A contract clause is read by someone who has drafted contracts. A failing test is diagnosed by an engineer who writes code for a living.
- Labelling: images, video, audio and sensor data, labelled precisely
- Training data: expert-written answers, rankings and reasoning
- Evaluation data: expert-written tests and scoring
- Alignment data: preference judgements and reference answers
Why expertise changes the data
Most labelling errors that matter are not typos. They are plausible answers that a non-specialist would accept and a specialist would not. A model trained on those labels learns to sound right in exactly the places where being wrong is costly.
Take a simple case. A user asks a model: "I missed a dose. Can I take two next time?" The model says yes. A medicine expert flags the answer as unsafe and writes the better one: "Don't double up; take the next dose as usual and check with your doctor." That correction is the data.
Expert labellers also notice when a task is ambiguous, when a guideline does not cover a case, or when the honest answer is "it depends". Those notes are often as useful to a research team as the labels themselves.
Data types we work with
Saolabs works across the modalities that current models are trained and evaluated on. The common thread is that each type needs a human with domain knowledge to judge what correct looks like.
Formats follow your pipeline. We agree the schema up front, whether that is JSONL for SFT pairs, ranked completions for preference training, or frame-level annotations for video.
- Language and code
- Vision, video and voice
- Agent trajectories
- Simulation and robotics sensor logs
- Physical-world sensor data
One team, sourcing through delivery
Saolabs sources pre-vetted experts and produces the data with them. The same team that recruits a cardiologist or a tax lawyer also writes the guidelines they work to and reviews what they produce.
Experts come in through our Hiring OS, outreach agents, recruiter network and jobs marketplace. They are vetted through AI voice interviews, real-world work tests in their field and skills assessments before they are matched to a project.
How an engagement typically works
We start by scoping: what the data is for, what a good label looks like, and where the hard cases are. From there we produce a small sample set so your team can check the output against your own judgement before anything scales.
Once the guidelines are settled, we match experts to the work and move into production, with expert review and quality checks along the way. Delivery comes with notes on known limits, such as categories with thin coverage or cases where experts disagreed, so you can decide how to weight the data.