Deep Learning Research Engineer Intern

DataRobot FZ-LLCSan Francisco, Boston, California, Washington, MassachusettsRemoteInternshipListed 4 hours ago

Apply now

About this role

Job Description:

DataRobot delivers AI that maximizes impact and minimizes business risk. Our platform and applications integrate into core business processes so teams can develop, deliver, and govern AI at scale. DataRobot empowers practitioners to deliver predictive and generative AI, and enables leaders to secure their AI assets. Organizations worldwide rely on DataRobot for AI that makes sense for their business — today and in the future.

This effort is about building a multi-target joint probabilistic foundation model that can be used across industries to tackle some of the hardest real-world problems. The ambition goes beyond what tabular and time-series foundation models usually do: one model should support temporal forecasting, unordered tabular regression and classification, missing-data completion, and mixed-modality inputs and outputs while learning coherent joint structure across connected variables, rows, horizons, and scenarios. The primary objective is to create real value in serious use cases rather than optimize benchmark scores in isolation, although strong results on standard benchmarks are expected to follow closely.

This role is designed as a research-engineering internship for someone operating at post-doctoral level, or very close to it, who wants to shape the architecture of a probabilistic foundation model. The center of gravity is the model itself: the encoders that read temporal, tabular, and mixed-modality inputs, the attention and sequence mechanisms that carry structure across variables, rows, and horizons, and the distributional output heads and decoding schemes that turn hidden states into coherent joint predictions. We are looking for someone who can reason precisely about what an architecture can and cannot represent, and who then implements, trains, and evaluates the resulting PyTorch models in a production-quality codebase. A background in stochastic modeling is a strong plus rather than a prerequisite: the model is trained on synthetic data drawn from stochastic dynamics, so a candidate who can also reason about stochastic differential equations and design synthetic data generators can contribute on both sides of the training loop.

Required Skills:

- Strong PyTorch skills and hands-on experience building and training deep models.
- Solid understanding of Transformers, attention variants, and long-context or state-space sequence architectures, including the tradeoffs between quality, memory, and latency.
- Experience with probabilistic modeling in neural networks: distributional output heads, likelihood-based or proper-scoring-rule losses, mixture or flow models, or related uncertainty-aware learning setups.
- Strong foundation in probability and statistics, enough to reason about densities and masses, dependence between variables, and calibration.
- Ability to connect architectural ideas to working GPU-native implementations, controlled experiments, and diagnostics that show why a change helped.
- Strong engineering habits: readable code, tests, reproducible experiments, and disciplined evaluation of model changes.
- Ability to debug training instability, reason about why results changed, and iterate quickly from hypothesis to evidence.

Strong Plus:

- Working knowledge of stochastic processes and stochastic differential equations: drift and diffusion, jumps, regime switching, mean reversion, heavy tails, and how such dynamics are simulated numerically.
- Experience designing synthetic data generators or simulation-based training curricula, and an understanding of how the training distribution shapes what a model learns.
- Depth in a domain with rich stochastic structure such as finance, energy, commodities, or a similarly quantitative field.
- Familiarity with tabular or mixed-modality deep learning: categorical, ordinal, count, and bounded targets alongside continuous ones.

What You Can Expect To Work On:

- Designing and improving the core model architecture: input encoders for temporal, unordered tabular, and mixed-modality data; attention factorizations across variables, rows, and horizons; the distributional output heads and decoding strategies that produce coherent joint samples.
- Running controlled architecture studies, from ablations and scaling behavior to memory and throughput profiles, and turning the results into design decisions.
- Building scalable PyTorch implementations that support larger input and output spaces, better throughput, and tighter memory budgets.
- Studying how architectural choices, data-generation choices, and inference constraints interact to change benchmark quality and real-world usefulness.
- For candidates with the stochastic-modeling background: extending the synthetic-data engine with richer stochastic dynamics, constraints, dependence structures, heavy tails, and regime behavior.
- Turning research ideas into robust implementations and credible empirical results.

What You Should Expect:

- Work that sits close to the core of the project rather than at the edges.
- A fast research loop where good ideas can move quickly from hypothesis to implementation to benchmark.
- A real chance to contribute to research that aims for top-tier, state-of-the-art outcomes rather than only incremental internal work.
- A team that values rigorous thinking, clean code, and practical usefulness at the same time.
- Exposure to problems that combine deep learning, stochastic modeling, and real deployment constraints instead of isolating only one of those dimensions.
- A role that fits someone with broad interests who wants to work across both advanced modeling and serious implementation work rather than staying narrowly specialized.

The talent and dedication of our employees are at the core of DataRobot’s journey to be an iconic company. We strive to attract and retain the best talent by providing competitive pay and benefits with our employees’ well-being at the core. Here’s what your benefits package may include depending on your location and local legal requirements: Medical, Dental & Vision Insurance, Flexible Time Off Program, Paid Holidays, Paid Parental Leave, Global Employee Assistance Program (EAP) and more!

DataRobot Operating Principles:

- Wow Our Customers
- Set High Standards
- Be Better Than Yesterday
- Be Rigorous
- Assume Positive Intent
- Have the Tough Conversations
- Be Better Together
- Debate, Decide, Commit
- Deliver Results
- Overcommunicate

Research shows that many women only apply to jobs when they meet 100% of the qualifications while many men apply to jobs when they meet 60%. At DataRobot we encourage ALL candidates, especially women, people of color, LGBTQ+ identifying people, differently abled, and other people from marginalized groups to apply to our jobs, even if you do not check every box. We’d love to have a conversation with you and see if you might be a great fit.

DataRobot is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. DataRobot is committed to working with and providing reasonable accommodations to applicants with physical and mental disabilities. Please see the United States Department of Labor’s EEO poster and EEO poster supplement for additional information.

Use of Artificial Intelligence in Our Hiring Process

DataRobot uses approved AI-powered tools to support the hiring process in selected regions. These tools may assist in writing job descriptions, reviewing applications, assessing qualifications, and evaluating candidate materials. All decisions regarding applications are made by members of the DataRobot team.

All applicant data submitted is handled in accordance with our Applicant Privacy Policy .