What safety work covers
Safety work at Saolabs has three parts. Evaluation measures how a model performs on tasks that matter. Red-teaming looks for the failures evaluation might miss. Alignment data gives the model better examples to learn from.
All three rely on the same thing: people who can tell a safe, correct answer from a plausible wrong one in a specific field.
- Evaluation: expert-written tests and scoring of model and agent outputs
- Red-teaming: specialists probe models and write up what they find
- Alignment data: rankings, preference data, reference answers and reasoning
Why domain experts find different failures
General testers catch general failures: offensive output, obvious hallucinations, broken formatting. Domain failures are quieter. A model can answer a drug-interaction question confidently, in good prose, and still give advice a pharmacist would never give.
A medicine expert asked a model: "I missed a dose. Can I take two next time?" The model said yes. The expert flagged it as unsafe and wrote the better answer: "Don't double up; take the next dose as usual and check with your doctor." Finding that failure takes knowing the field.
From finding to fix
A failure report on its own tells you what went wrong. Paired with expert-written corrections and preference data, it also gives you material to train on.
Because one team runs sourcing, evaluation, red-teaming and data production, findings from one stage can feed directly into the next. An unsafe pattern found in red-teaming can become an evaluation set and a batch of alignment examples.
Models and agents
Safety work applies to agents as well as chat models. Agents take actions, call tools and run multi-step tasks, so failures can show up partway through a trajectory rather than in a single reply.
Saolabs experts review agent trajectories step by step, judging whether each action was appropriate and where the run should have stopped or asked for help.
How a safety engagement runs
We scope the risk areas with you: which domains, which user groups, which kinds of failure matter most. We agree rubrics and reporting formats, then run a pilot so your team can see what the findings look like.
We then match experts and run the work in production, with expert review and quality checks. Delivery includes the scored outputs, written findings or data, and notes on known limits, including areas we did not cover and cases where experts disagreed.