About this role
Description
The Artificial General Intelligence (AGI) team is looking for a passionate, talented, and inventive Software Development Engineer to help advance the quality and capabilities of industry-leading large language models and Generative AI systems.
In this role, you will build scalable solutions for discovering, processing, curating, transforming, and delivering high-value data used across model evaluation, training, and post-training workflows. You will work with large-scale and heterogeneous datasets, including data produced by partner teams, long-context data, and other data sources that can help improve model development.
You will collaborate closely with scientists, ML engineers, evaluation teams, and other engineering teams to identify valuable data sources, understand downstream data requirements, and develop mechanisms that make data easier to discover, standardize, validate, transform, and consume at scale.
Your work will help expand the breadth, quality, and usability of data available for Generative AI development and establish scalable data infrastructure, quality standards, and workflows that enable downstream teams to evaluate and improve models more effectively.
This is an opportunity to solve challenging problems at the intersection of data infrastructure, evaluation data, and Generative AI, and to directly influence the data foundation supporting Amazon’s next generation of AI models and experiences.
Key job responsibilities
Quickly learn and apply emerging technologies, methodologies, and best practices in Generative AI, data processing, and large-scale model development.
Design, build, and maintain scalable platforms and infrastructure for discovering, scanning, processing, curating, and delivering data used to support LLM evaluation, training, and model improvement.
Develop systems and mechanisms to identify useful evaluation data produced by partner teams, understand its characteristics and usability, and make relevant data easier to discover, standardize, validate, and consume.
Expand the breadth and quality of data available for model development, including evaluation datasets, long-context data, training data, and other high-value data sources.
Build scalable capabilities to transform heterogeneous data from different sources and formats into standardized, high-quality datasets that can be reliably consumed by downstream evaluation and training workflows.
Partner closely with Applied Scientists, ML Engineers, evaluation teams, and other engineering teams to understand data needs, identify gaps in available data, and translate those needs into scalable data solutions.
Investigate design approaches, prototype new technologies, evaluate technical feasibility, and drive data infrastructure solutions from concept to production.
Build reliable, high-quality data discovery and processing pipelines that operate at massive scale while improving data quality, coverage, observability, efficiency, and developer productivity.
Influence technical architecture and establish engineering best practices for data discovery, curation, transformation, and delivery across Generative AI model development workflows.
A day in the life
As an SDE with the AGI team, you will build scalable systems that discover, process, curate, transform, and deliver high-value data for model evaluation, training, and post-training workflows.
You will work with large-scale, heterogeneous data sources and partner closely with scientists, ML engineers, and evaluation teams to understand downstream data needs and make useful data easier to identify, standardize, validate, and consume.
You will influence technical architecture, establish engineering best practices, and build reliable infrastructure that improves data quality, usability, and developer productivity across Generative AI model development.
About the team
Join our AGI team and work at the forefront of Generative AI. We build scalable data and model development infrastructure that helps teams discover, process, curate, transform, and use high-value data to support the development of large language models and other foundation models.
You will collaborate with talented engineers and scientists working across model development, evaluation, data, and large-scale distributed systems. Together, we solve challenging technical problems at massive scale and build mechanisms that make valuable data easier to identify, standardize, validate, and consume across evaluation, training, and post-training workflows.
This is an opportunity to influence how next-generation AI systems are developed, expand the breadth and quality of data available to model development teams, and help shape the future of artificial intelligence at Amazon.
Mission of the team:
We build scalable data discovery, curation, transformation, and infusion capabilities that accelerate the development of high-quality foundation models. By identifying high-value data across 1P and other relevant sources—including data produced through evaluation workflows—we make that data easier to discover, standardize, validate, and use across evaluation, training, and post-training.
Our mission is to expand the breadth, quality, and usability of data available for model development, enabling downstream teams to more effectively evaluate models, address data gaps, and improve model performance.
Basic Qualifications
- 3+ years of non-internship professional software development experience
- 3+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- 3+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or distributed software applications, tools, systems, and services using: C#, C++, Java, or Perl experience
- 3+ years of Object Oriented Design experience
- Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field
Preferred Qualifications
- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .
USA, WA, BELLEVUE - 143,700.00 - 194,400.00 USD annually