About this role
What you'll do
- Help develop the technical roadmap and architecture for the evaluation platform, from offline benchmarking to online/production monitoring of agentic and LLM-based systems.
- Design eval methodologies appropriate to different stages of the pipeline: golden/regression test sets, human-in-the-loop review workflows, LLM-as-judge approaches, and automated metrics for task success, safety, and hallucinations.
- Build the data infrastructure evaluation depends on: annotation and labeling pipelines, dataset versioning, data quality checks, and tooling that lets researchers and product teams run and interpret experiments without needing platform team help.
- Partner closely with Research, Product, and Platform teams to productize experiments into robust AI solutions
- Represent the eval platform to stakeholders outside the immediate team- set expectations on what "good" looks like for a model/agent release, and report on platform health and coverage.
- Stay current with advancements in ML, NLP, voice, and LLM systems, and contribute actively to technical discussions across teams.
- Mentor and support other engineers through design reviews, feedback, and knowledge sharing.
What you'll need
- Deep, hands-on experience building and operating evaluation systems for modern ML/LLM/agentic systems- not just consuming existing eval tools.
- Demonstrated experience leading the technical direction of a project or small team: setting architecture, driving design reviews, and being accountable for a system's long-term health (not just shipping features).
- Strong architectural skills, with proven experience designing complex, data-intensive software systems and production experience with Python, AWS, Kubernetes, and/or Docker.
- Experience designing data pipelines for ML evaluation- labeling/annotation workflows, dataset versioning and quality control, and reproducible benchmarking.
- A Bachelor’s Degree in CS or other related fields
- Demonstrated technical mentorship of junior and mid-level engineers, driving adoption of best practices and architectural alignment for scalability and extensibility.
- Desire to learn, teach, and collaborate closely with cross-functional peers.
What we'd like to see
- Experience building and evaluating agentic systems at scale.
- Experience with voice/audio quality evaluations.
- Production experience with LLM-centric services (e.g., inference, orchestration, evaluation, monitoring)
- Familiarity with large-scale ML experimentation, benchmarking, or simulation frameworks.
- Experience with conversational/customer-support AI domains (e.g., containment rate, conversation quality, goal completion).
- Knowledge of techniques for optimizing model architectures for faster inference.
- Experience with AWS, CI/CD, Kafka, Athena
Compensation package also includes a performance bonus on top of the listed salary range
Separately, we also offer a compelling equity grant comprised of stock options