About this role
You will take on the following responsibilities:
- Partner with ML researchers to explore model architectures, training methodologies, LLM and agentic engineering, and experimental objectives, translating modeling needs into engineering priorities.
- Translate prototypes into reusable, maintainable, production-ready code, ensuring research insights scale seamlessly into operation.
- Design and optimize data loading and pipeline infrastructure for training and inference at scale, including GPU utilization, batching/bin-packing strategies, and other large-scale compute optimizations.
- Extend core research libraries to support agentic workflows, integrating modern agent frameworks while maintaining a high quality bar for AI-assisted or AI-generated code and output.
- Lead structural improvements to keep a growing, multi-contributor codebase coherent, testable, and maintainable as usage and team size scale, evolving internal libraries toward open-source-caliber engineering practices.
You should possess the following qualifications:
- BS in Computer Science, Mathematics, Physics, or related technical subject area
- Minimum 1 year of experience required; 7-15 years of experience preferred
- Experience in quantitative software engineering
- Strong software engineering skills using Python, AI coding tools, version control, testing frameworks, and CI/CD practices
- Hands-on experience building high-throughput GPU ML systems, including low-level performance optimization, batching, scheduling, orchestration, and training infrastructure
- Experience collaborating across research and engineering functions, applying the right level of rigor from prototype-stage code to production systems.
Preferred Skills
- Experience with CUDA
- Experience building on and extending leading ML frameworks such as PyTorch, TensorFlow, or Hugging Face
- Experience using LLMs and agent frameworks to automate research and engineering workflows in ways that are reliable and useful in practice
- Experience in agent evaluation frameworks, RLHF, RLVR, policy optimization, or synthetic data generation.