Answers written by people who know.For models to learn from.

Saolabs produces expert-written answers, rankings and reasoning for supervised fine-tuning and preference training. Vetted specialists write and judge the data; we run the work end to end.

What this data is

Training data here means written supervision: the prompts, answers and judgements a model learns from after pre-training. Saolabs produces three main kinds.

Each is written or judged by a specialist in the field. A model learning to answer finance questions learns from people who work in finance.

  • SFT data: prompts paired with expert-written reference answers
  • Preference data: expert rankings and pairwise choices between model outputs, for RLHF-style training
  • Reasoning traces: step-by-step working that shows how an expert reaches an answer

Why expert-written answers matter

A fine-tuned model imitates its examples. If the reference answers are fluent but subtly wrong, the model learns to be fluent and subtly wrong.

Experts write answers that are correct, appropriately hedged and safe to act on. They know when the right response is a caveat, a referral or a refusal, and they can explain the choice in a sentence.

Rankings and preference data

Preference training needs judgements about which of two or more responses is better, and why. Those judgements are only as good as the person making them.

Saolabs experts rank model outputs against a rubric agreed with your team: correctness, safety, completeness, tone or whatever your objective needs. Where useful, they add a short rationale, so you can audit the preference rather than just the label.

Reasoning traces

Reasoning data shows the path to an answer, not just the answer. Mathematicians write proofs step by step. Engineers walk through a diagnosis. Lawyers lay out how a rule applies to a set of facts.

These traces help models learn structured problem-solving in specialist domains. They also make errors easier to spot, because a wrong step is visible in a way a wrong final answer is not.

How a training-data project runs

We scope the domain, task mix and output format with you, including what a good answer looks like and which edge cases matter. A sample set comes first, so your team can check the writing and the rubric before production.

We then match vetted experts to the work and produce the data, with expert review and quality checks along the way. Delivery comes in your agreed format with notes on known limits, such as topics with thinner coverage or prompts where experts disagreed on the best answer.

Questions, answered.

What is SFT data?

SFT (supervised fine-tuning) data is a set of prompts paired with high-quality reference answers. A model is trained to produce answers like the references, so their correctness and style shape the model's behaviour directly.

What is preference data for RLHF?

Preference data records which of several model responses a human judge prefers, often with a reason. It is used in reinforcement learning from human feedback (RLHF) and related methods to teach a model which kinds of answers are better.

What are reasoning traces in LLM training data?

Reasoning traces are step-by-step explanations of how to reach an answer, written by someone who can solve the problem. They help models learn structured reasoning and make individual mistakes easier to find.

Who writes Saolabs training data?

Saolabs training data is written and judged by vetted experts in medicine, law, finance, engineering, software, science, mathematics and languages. Experts are vetted through AI voice interviews, real-world work tests and skills assessments.

How is expert-written training data different from synthetic data?

Synthetic data is generated by models and inherits their blind spots. Expert-written data comes from people who can recognise wrong or unsafe answers in their field, which makes it useful for teaching and checking exactly the behaviours synthetic data tends to get wrong.

Every safe modelhas an expert behind it.