Post-Training Research Scientist

Two SigmaNew York City, New YorkOn-sitePart-timeListed 4 hours ago

Apply now

About this role

You will take on the following responsibilities:
- Lead post-training efforts for LLMs applied to financial time series and quantitative reasoning
- Design and execute RLHF, DPO, and related alignment methods at scale, including deployment of substantial compute budgets (O($100mm))
- Build infrastructure for preference data collection, reward modeling, and policy optimization on financial datasets
- Drive research agenda connecting post-training methods to quantitative finance applications
- Collaborate with quant researchers to define task distributions and evaluation frameworks
- Unblock production systems dependent on post-training capabilities

You should possess the following qualifications:
- BS or equivalent work experience in Science, Technology, Engineering or Math (an MS is a plus).
- Minimum 1 year of experience required; 1-10 years of experience preferred (ideally 1-5 years) at a frontier AI lab (OpenAI, Anthropic, DeepMind, Meta FAIR, or equivalent)
- Shipped post-training systems in production: RLHF, DPO, RLAIF, or related methods
- Deep understanding of distributed training infrastructure: multi-node GPU clusters, training stability, checkpointing
- Track record managing large-scale compute: budgeting, experiment design, ablations
- Publications or demonstrated expertise in alignment, preference learning, or reward modeling
- Hands-on implementation skills: PyTorch/JAX, distributed frameworks (DeepSpeed, FSDP, etc.)