What this data is
Training data here means written supervision: the prompts, answers and judgements a model learns from after pre-training. Saolabs produces three main kinds.
Each is written or judged by a specialist in the field. A model learning to answer finance questions learns from people who work in finance.
- SFT data: prompts paired with expert-written reference answers
- Preference data: expert rankings and pairwise choices between model outputs, for RLHF-style training
- Reasoning traces: step-by-step working that shows how an expert reaches an answer
Why expert-written answers matter
A fine-tuned model imitates its examples. If the reference answers are fluent but subtly wrong, the model learns to be fluent and subtly wrong.
Experts write answers that are correct, appropriately hedged and safe to act on. They know when the right response is a caveat, a referral or a refusal, and they can explain the choice in a sentence.
Rankings and preference data
Preference training needs judgements about which of two or more responses is better, and why. Those judgements are only as good as the person making them.
Saolabs experts rank model outputs against a rubric agreed with your team: correctness, safety, completeness, tone or whatever your objective needs. Where useful, they add a short rationale, so you can audit the preference rather than just the label.
Reasoning traces
Reasoning data shows the path to an answer, not just the answer. Mathematicians write proofs step by step. Engineers walk through a diagnosis. Lawyers lay out how a rule applies to a set of facts.
These traces help models learn structured problem-solving in specialist domains. They also make errors easier to spot, because a wrong step is visible in a way a wrong final answer is not.
How a training-data project runs
We scope the domain, task mix and output format with you, including what a good answer looks like and which edge cases matter. A sample set comes first, so your team can check the writing and the rubric before production.
We then match vetted experts to the work and produce the data, with expert review and quality checks along the way. Delivery comes in your agreed format with notes on known limits, such as topics with thinner coverage or prompts where experts disagreed on the best answer.