About this role
This role owns how non-text modalities enter the pretraining run: the data we train on, the encoders and fusion architecture that carry it into the language model, and the capabilities we get out. You will build and scale multimodal data pipelines (images, video, 3D, physical / scientific data), run the architecture research that decides how modalities are tokenized, encoded, and interleaved with text, and define the evaluations and ablations for validating your data and architecture. The work spans the full stack, from a data or architecture hypothesis to a controlled training experiment to a verdict that lands in the flagship recipe.