Applied Scientist II, AWS Neuron Science - Core Algorithm

AmazonCupertino, CaliforniaOn-siteFull-timeJunior, 1–2 yearsListed 53 minutes ago

Apply now

About this role

Description

The AWS Neuron Science Core Algorithm team is looking for talented Applied Scientists to push the frontier of hardware-aware machine learning for Trainium and Inferentia, the AWS Machine Learning accelerators. In this rare role at the intersection of LLM modeling, large-scale training systems, and hardware/datatype co-design, you own model and algorithm decisions jointly with AWS custom silicon. You will own solutions end-to-end from research through production, publish at top venues, and work alongside distinguished engineers and scientists in a strategic growth area for AWS.

We actively work on these areas:

- Low-precision training and inference: MXFP8, MXFP4, and sub-4-bit training and inference recipes, stochastic rounding, and Trn4 datatype exploration.

- Trn-friendly architectures: model architectures that exploit hardware strengths without sacrificing quality.

- System-aware optimizers & efficient distributed systems: efficient optimizers and distributed system that gives best accuracy, co-designed with the hardware.

- Foundation-model pre-training accuracy: end-to-end validation across model scales, catching training divergence early, and equivalence-checking tooling.

- GenAI for systems: RL post-training for NKI kernel generation, mitigating reward-hacking and accelerating under low precision on Trn.

Key job responsibilities

- Own scientific problems end-to-end - from research and experimentation through production impact - applying rigorous evaluation to complex, ill-defined problems at large scale.
- Develop production-quality code in PyTorch or JAX and integrate scientific components into large-scale training and inference systems with operational excellence and efficient resource usage.
- Partner with foundation-model, engineering, and hardware-architecture teams so your findings directly inform what gets built into Trainium and shipped in the product stack.
- Mentor fellow scientists and interns, give constructive peer reviews, and help shape team goals, priorities, and the technical roadmap.
- Author and publish research at top peer-reviewed venues (ICLR, NeurIPS, ICML, MLSys) and engage the broader scientific community.

A day in the life
You might start your morning reviewing large-scale training runs — checking accuracy at a new low-precision datatype or debugging a divergence before it costs a run — then join a design discussion with engineering partners on how to land your recipe in the production stack. After lunch you could be whiteboarding a Trn-friendly architecture variant or an RL post-training approach for kernel generation with a teammate, then writing code to prototype it on Trainium. You will regularly present findings to the team and to leadership, review peers' and interns' work, and stay connected with the academic community.

About the team
AWS Neuron is the software of Trainium and Inferentia, the AWS Machine Learning chips. Inferentia delivers best-in-class ML inference performance at the lowest cost in the cloud to our AWS customers. Trainium is designed to deliver the best-in-class ML training performance at the lowest training cost in the cloud, and it's all being enabled by AWS Neuron. Neuron is a software that includes an ML compiler and native integration into popular ML frameworks. Our products are being used at scale with external customers like Anthropic and Databricks as well as internal customers like Amazon FMR, Amazon AGI, Amazon Bedrock, Amazon Robotics, Amazon Ads, and many more.

Basic Qualifications

- PhD, or Master's degree and 4+ years of CS, CE, ML or related field experience
- Experience in patents or publications at top-tier peer-reviewed conferences or journals
- 3+ years of building models for business application experience
- Experience programming in Java, C++, Python or related language
- Experience in state-of-the-art deep learning models architecture design and deep learning training and optimization and model pruning

Preferred Qualifications

- Experience in professional software development
- Experience in any of the following areas: algorithms and data structures, parsing, numerical optimization, data mining, parallel and distributed computing, high-performance computing
- Knowledge of standard speech and machine learning techniques

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, CA, Cupertino - 171,600.00 - 222,200.00 USD annually