About this role
About Handshake
Handshake's mission is to organize expert human knowledge to advance the AI economy. Handshake AI works directly with frontier labs on their most consequential data, evaluation, and post-training challenges, building the systems that turn expert human knowledge into the data and evaluations that make frontier models better.
This is an unusual moment to join: Handshake AI works directly with frontier labs on their most consequential data, evaluation, and post-training challenges. You will work alongside engineers, researchers, operators, and builders from organizations including Scale AI, Meta, xAI, Notion, Coinbase, and Palantir—and help build the systems that make expert human knowledge useful for advancing AI.
###
About Handshake Labs
Handshake Labs is building external AI products, research platforms, and customer-facing AI systems. We are evolving work that is often custom-built for an individual partner into reusable products and platforms that improve with every deployment.
Our work spans the full post-training loop: designing evaluations and training environments, building high-quality data and feedback systems, running experiments, and turning what works into durable infrastructure. For example, we are developing agents that can analyze long, complex coding-agent sessions in days rather than weeks—with expert review and calibration built into the system.
###
The Role
We are hiring a Member of Technical Staff, Coding Evals to help define how frontier coding agents are measured, understood, and improved. This is a high-ownership role for someone who is deeply passionate about building software with agents—and about creating the evaluations that reveal where those agents truly succeed and fail.
You will design and publish coding benchmarks that meaningfully challenge state-of-the-art agents. You will work with AI researchers, software engineers, and domain experts to turn difficult, real-world software tasks into high-signal evaluation environments, datasets, verifiers, and feedback systems. The work will help shape both how the frontier evaluates coding agents today and where the field goes next.
Early members of the team will have unusual influence over our technical direction, operating culture, and the open-source benchmarks, software, and research products we build. We care more about demonstrated technical depth, judgment, and a builder's mindset than a specific title, degree, or career path.
Location: San Francisco & Mountain View preferred; we are open to exceptional candidates in other locations.
###
What you'll do
- Design, build, and publish coding benchmarks that measure meaningful progress in frontier coding agents.
- Create realistic, difficult software-engineering tasks, repositories, environments, and test harnesses that expose agent capabilities and failure modes.
- Develop reliable verifiers, graders, reward signals, and evaluation methodology for agentic software development.
- Research how to make coding-agent evaluations representative, difficult, robust, and resistant to shortcutting or benchmark contamination.
- Analyze coding-agent behavior and trajectories to understand where agents fail, what feedback is useful, and which capabilities matter next.
- Partner directly with AI researchers, software engineers, and expert contributors to develop high-signal tasks, data, and evaluation methods.
- Run fast, rigorous iteration loops: prototype, evaluate, interpret results, diagnose failure modes, and turn learnings into the next benchmark or system.
- Identify repeatable patterns across engagements and productize them into reusable software, benchmarks, datasets, and platforms.
- Raise the technical bar through strong design judgment, clear communication, code quality, and mentorship.
- Contribute to the field through open benchmarks, open-source tools, research, and technical writing where it creates leverage.
###
What we're looking for
- Deep enthusiasm for agentic software development, with clear evidence that you actively build, experiment with, or think seriously about coding agents.
- Strong software engineering skills and the ability to write clean, reliable, maintainable code.
- Strong Python skills and comfort working with modern ML tooling, evaluation infrastructure, and data workflows.
- Sound experimental judgment: you can form hypotheses, choose meaningful metrics, diagnose failures, and distinguish genuine capability improvement from evaluation artifacts.
- Experience designing systems—not only implementing specifications—including the ability to make tradeoffs around validity, quality, scale, reliability, and reuse.
- Comfort operating in an ambiguous, fast-moving environment with substantial ownership.
- Collaborative, low-ego communication and the ability to work effectively with researchers, engineers, domain experts, and customers.
You should also have at least one of the following:
- Experience working on a widely used coding-AI benchmark or evaluation suite.
- Published research on AI for coding or code generation at a leading venue, such as NeurIPS, ICML, ICLR, or COLM.
- Software engineering experience at a top technology company, or a strong public GitHub profile, paired with deep knowledge of agentic AI and a clear passion for building software with agents.
###
Especially compelling experience
- Building or maintaining coding benchmarks, coding-agent environments, repository-level evaluation suites, or open-source developer tools.
- Developing automated graders, test harnesses, programmatic verifiers, reward models, or reinforcement-learning environments for software tasks.
- Researching code generation, program synthesis, AI agents, reinforcement learning, post-training, or software-engineering productivity.
- Experience with post-training methods such as reinforcement learning, RLHF, preference optimization, supervised fine-tuning, or reward modeling.
- Strong public contributions through GitHub, papers, benchmarks, technical writing, or developer communities.
- Experience turning research prototypes or repeated customer work into robust, reusable products or platforms.
###
Why join
- Work on the benchmarks and environments that shape how the world's best coding agents are built and evaluated.
- Help define the next generation of agentic software-development tasks alongside frontier labs and expert practitioners.
- Publish open benchmarks, software, research, and technical writing that can shape the broader AI ecosystem.
- Help build an early technical organization where your work shapes the roadmap, standards, and culture.
- Join a company building durable infrastructure for careers in the AI economy.
# Perks
Handshake delivers benefits that help you feel supported—and thrive at work and in life.
The below benefits are for full-time US employees.
🎯 Ownership: Equity in a fast-growing company
💰 Financial Wellness: 401(k) match, competitive compensation, financial coaching
🍼 Family Support: Paid parental leave, fertility benefits, parental coaching
💝 Wellbeing: Medical, dental, and vision, mental health support, $500 wellness stipend
📚 Growth: $2,000 learning stipend, ongoing development
💻 Office: Commuting support, free lunch, and gym in our SF office
🏝 Time Off: Flexible PTO, 15 holidays + 2 flex days
🤝 Connection: Team outings & referral bonuses
Compensation
$200K – $350K • Offers Equity