Why software matters for AI safety
Models now write code that ships, and agents run commands on real systems. A passing test suite says little about whether the code is secure, maintainable or doing what the user meant.
Judging that takes the same review skills a senior engineer brings to a pull request. Those skills are what we vet for.
What software experts do
Software experts write reference solutions and rank model-generated code on correctness, security and design. They annotate why one solution is better than another, giving the model reasoning to learn from, not only a preference.
They evaluate coding agents on multi-step tasks, red-team models to produce insecure or harmful code, and review architecture decisions where a model's choices have long-term cost.
- Reference solutions and code review
- Ranking model-generated code with reasons
- Evaluation of coding agents on real tasks
- Security-focused red-teaming
- Architecture and design review
Failure modes software experts catch
Experts catch code that works on the happy path and breaks on the edge case. They catch injection risks, unsafe defaults, leaked secrets, deprecated APIs and dependencies that do not exist.
With agents, they catch actions as well as code: a destructive command run without confirmation, a change made outside the task's scope, or a test edited to pass instead of the bug being fixed.
How we vet software experts
Software candidates are interviewed by AI voice agents on how they approach design and debugging. They then complete real-world work tests, such as reviewing code for defects or judging a model's solution to a programming task.
Security specialists are assessed on security work, and architects on design. Credentials and experience are reviewed as part of vetting, and project-specific assessments can cover particular languages or stacks.