Safety has to hold in every languageNative speakers who can test it.

Saolabs sources native speakers, translators and linguists to produce expert-labelled training and evaluation data for models used in many languages.

Why languages matter for AI safety

Models are used in many languages, but much of their training and testing has been weighted towards a few. Behaviour checked in one language may not hold in another.

Meaning also lives in register, idiom and culture. A response can be grammatical and still wrong, rude or unsafe for the people reading it. Native speakers hear the difference.

What language experts do

Language experts write and rank responses in their native language, evaluate translations, and annotate where meaning, tone or register goes wrong. They also contribute to voice and video data.

They red-team models in their language to test whether safety behaviour carries over, and produce alignment data on culturally appropriate responses. Linguists help design evaluations that measure what a language actually requires.

  • Native-language answers and rankings
  • Translation evaluation
  • Voice and video data
  • Red-teaming beyond English
  • Alignment data on cultural context

Failure modes language experts catch

Experts catch literal translations that lose meaning, the wrong level of formality, mistakes in dialect, and idioms carried over from another language. They notice when a model sounds translated rather than written.

They also catch safety gaps: a request a model refuses in one language but answers in another, or content that is harmless in one culture and offensive in the next.

How we vet language experts

Language candidates are interviewed by AI voice agents, which also lets us hear how they speak. They then complete real-world work tests, such as translating a passage with the right register or judging a model's response in their language.

Credentials and experience are reviewed as part of vetting. Native speakers, translators and linguists are each assessed on the work their role involves.

Questions, answered.

Where can I find native speakers for AI training data?

Saolabs sources vetted native speakers, translators and linguists to produce training and evaluation data for AI models. They write and rank responses, evaluate translations, contribute voice data and red-team models in their language.

Why red-team AI models in languages other than English?

Safety behaviour checked in one language may not hold in another. Red-teaming in other languages tests whether a model refuses, warns and responds appropriately for every group of people who use it.

What do translators catch that machine translation misses?

Translators catch lost meaning, the wrong level of formality, dialect errors and idioms that do not carry over. They judge whether a response reads as natural to a native speaker, not only whether the words match.

How are language experts vetted?

Language experts are interviewed by an AI voice agent and assessed on real-world work tests, such as translating a passage or judging a model's response in their language. Credentials and experience are reviewed as part of the process.

Every safe modelhas an expert behind it.