About this role
Founding Large-Scale Mid-Training/RL Infrastructure Engineer
Location: Onsite in Palo Alto
Compensation: Competitive Salary + Equity
## About Peano AI
Peano AI is building the infrastructure and application stack for the next generation of agentic AI systems.
We believe token usage will grow exponentially over the coming years, but routing all inference and training through closed model providers will remain too expensive for many users and enterprises. Our thesis is that agentic applications require a vertically integrated stack: high-throughput, cost-efficient serving and training infrastructure paired with an application layer designed for long-running, agentic workloads.
Peano AI is building the Agent Cloud, a serving and training infrastructure platform purpose-built for agentic workloads, long-context inference, large-scale open-source model deployment, and pretraining of our own foundation model. By combining infrastructure and application design, we aim to make open-source models and custom foundation models significantly more performant, practical, and competitive.
## About This Role
We are looking for a Large-Scale Mid-Training/RL Infrastructure Engineer to help build, optimize, and scale out our in-house foundation model training stack , with an emphasis on both core pretraining and RL-enhanced methods. This role is deeply technical and directly impacts Peano AI's core model product.
You will architect, scale, and tune the distributed and accelerator-aware training infrastructure for massive language models, including data pipelines, model sharding, checkpointing, rollout and reward-based RL, and overall system throughput. Experience with large-scale model pretraining and mid-training interventions is required.
Preferred experience with frameworks such as Megatron, Transformer-Engine, verl, slime, and related large-scale training and RL toolkits. Familiarity with training pipeline scaling, mixed-precision, memory optimization, and deep understanding of both supervised and RL-based training cycles is essential.
## What You'll Do
- Build, optimize, and scale the distributed training infrastructure for Peano AI's own foundation models
- Own and improve throughput, efficiency, scalability, cost, and reliability of end-to-end pretraining and RL pipelines
- Architect distributed data loading, sharding, checkpointing, batching, and accelerator utilization for multi-node LLM training
- Integrate and optimize RL components: rollout, reward modeling, environment orchestration, and mid-training signal injections
- Work with and extend frameworks such as Megatron, Transformer-Engine, verl, slime , and other high-performance training libraries
- Tune memory usage, mixed-precision, and runtime performance for massive models and long-context workloads
- Debug and profile training performance bottlenecks—across model code, distributed compute, networking, and infrastructure
- Collaborate with application, infrastructure, and research teams to ensure our foundation models deliver on performance and functionality
- Translate research prototypes and experimental features into production-ready, scalable training systems
## Qualifications
- Significant experience with large-scale deep learning model training and distributed system design
- Proven track record of pretraining and/or RL-based fine-tuning of large models (LLMs or comparable scale)
- Deep familiarity and hands-on experience with frameworks such as Megatron, Transformer-Engine, verl, slime , or similar large-scale/foundation-model toolkits
- Experience with data pipeline design, model/data/optimizer sharding, checkpointing, rollout/reward design, and training operations at scale
- Strong accelerator (GPU/TPU) and memory optimization skills
- Excellent debugging, profiling, and performance-tuning abilities in distributed environments
- Comfort working in deeply technical, high-ownership, early-stage startup settings
## Cultural Fit
- Hands-on technical excellence and strong engineering judgment
- End-to-end ownership—from design to implementation to production
- Bias for action: ship quickly, learn from failures, iterate
- High intensity during critical milestones, focused on customer and product outcomes
- Ability to work deeply and sustain high execution
- Clear communicator, low ego, strong collaborator
- Thrives in ambiguity, rapid change, and taking on multiple roles
- Lifelong learner with a belief that capability compounds with time and effort
If you are excited to build the core training and RL infrastructure for Peano AI's vertically integrated foundation model, tackle the hardest scaling and optimization challenges, and help push the boundaries of agentic AI, we'd love to talk.