About this role
Salary range: $250,000 - $500,000/year + benefits
Description: Transluce is a fast-moving nonprofit research lab building the public tech stack for AI evaluation and oversight. We have contributed foundational research to the study of AI agents and their behaviors, and we put the results where they can change decisions: in the hands of labs, policymakers, and the public.
About the role: We are looking for a frontline evaluator to investigate the honesty and alignment of frontier AI systems. You will develop and run rigorous automated evaluations, conduct novel analyses of massive datasets, and surface behaviors of interest, writing up your results for technical, policy, and lab audiences. Example behaviors of interest include misreporting results, falsely claiming success, evaluation awareness, and memetic effects within AI swarms. Your work will uncover risks that might otherwise go unnoticed, turning observations into evidence for the public to decide how AI is built, deployed, and governed.
As an early member of a highly collaborative team, you will learn and grow quickly, and work with our governance and infrastructure teams to scale your impact and technical reach. Your work will be with frontier labs, with governments, and on key topics of public interest — for example, rapid-response investigations of major incidents, public reports on frontier model behavior, and serving as an independent evaluator for governments such as the EU. It may include embedded evals within frontier AI labs as those opportunities arise. As we further develop this approach to evaluating frontier AI systems, we expect the role to evolve.
Core responsibilities:
- Develop and conduct evaluations to investigate misalignment and unexpected behaviors in AI systems
- Write code to quickly build and run evaluation and analysis workflows, including data-science tools, LLM-as-a-judge pipelines, and evaluation environments
- Analyze agent transcripts, datasets, and evaluation results, finding notable patterns or unexpected behaviors to investigate
- Perform investigations under time and access constraints, iterating quickly while validating conclusions
- Translate findings into clear, rigorous written reports
- Collaborate with teammates and external technical stakeholders to conduct evaluations, communicate progress, and share relevant findings
What we're looking for:
- Strong empirical judgment to extract meaningful findings from data, design follow-up experiments, and identify anomalies.
- Proficiency in Python to implement data analysis, experiments, and evaluation tooling.
- Ability to turn an ambiguous concern into a testable question.
- Operational resourcefulness, adaptability, and sound prioritization in the face of incomplete information.
- Ability to iterate quickly and balance between scrappiness and thoroughness based on the impact needs of a project.
- Strong communication skills, including the ability to clearly explain technical findings to various readers.
- Collaborative orientation, low ego, and openness to both giving and receiving feedback.
Preferred qualifications (nice to have):
- Experience with evaluation (e.g. model-based judges or agent evaluations), red-teaming, or rapid-response and incident-analysis work on frontier systems
- A track record of insightful empirical investigations shared through reports, blog posts, research, or independent projects.
- Practical understanding of how model training and deployment can affect behavior and evaluation results.
- Experience in customer-facing, consulting, or forward-deployed roles translating ambiguous partner needs into concrete deliverables.
- Experience delivering technical work under tight deadlines or in unfamiliar or constrained environments.
We are hiring at all levels of experience and would encourage those enthusiastic about the role who do not meet all of the qualifications to apply. We are located in San Francisco and excited to work together in-person. We are open to sponsoring international visas.