About this role
About the Role
Build core reinforcement learning systems from the ground up as a founding engineer on an early-stage AI team. You will work directly with the founders and own the process from environment design and model training through evaluation, helping shape the company's technical direction.
What You'll Do
- Design and build reinforcement learning environments, reward functions, and training pipelines.
- Train and fine-tune models using methods such as PPO, GRPO, DPO, RLHF, and RLAIF.
- Develop evaluation frameworks to measure model and agent performance.
- Run experiments, interpret results, and decide which approaches to explore next.
- Scale training on GPU clusters and maintain reliable pipelines.
- Turn research ideas into production systems.
- Help establish engineering culture and hire future engineers.
What We're Looking For
- At least 2 years of hands-on experience in reinforcement learning or machine learning engineering, with relevant experience potentially ranging from 2 to 10 or more years.
- Strong Python skills and deep experience with PyTorch or JAX.
- Practical experience training models with reinforcement learning, including policy gradient methods, reward modeling, or RLHF.
- Comfort with distributed training and GPU infrastructure, including tools such as Ray, CUDA, and Kubernetes.
- Experience with LLM post-training or agent training is a plus, as is familiarity with Gymnasium, Ray RLlib, Isaac, MuJoCo, TRL, verl, or OpenRLHF.
- A degree in computer science, mathematics, physics, or a related field is sought; a master's or PhD is a plus. Publications or open-source work in RL are also valued.
- A hands-on builder who enjoys moving quickly in a small, early-stage team.
Compensation & Benefits
Compensation is $125,000 to $200,000 USD annually.
Location
On-site in San Francisco, California, United States.