About this role
This role owns the question of what our models learn from: how we generate, filter, weight, and scale training data across pretraining and midtraining. You will design and run synthetic data pipelines at scale, build rigorous methods to measure whether a data intervention actually improves the model, and run the scaling and ablation experiments that decide what goes into the next training run. The work is end to end: from a hypothesis about data, to a generation or curation pipeline, to a controlled training experiment, to a verdict that changes the recipe.