Senior Software Engineer, AI/ML, Ads Training

GoogleMountain View, CaliforniaOn-siteFull-timeSenior, 5–8 yearsListed 39 minutes ago

Apply now

About this role

Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.

The Ads Training Runtime team enables the core machine learning stack for the Ads Training Infrastructure. Our mission is to empower Ads machine learning teams with a flexible, high-performance infrastructure and intuitive tooling, enabling rapid innovation and seamless adoption of cutting-edge hardware and software technologies to maximize performance and deliver excellent business outcomes.

In this role, you will drive TPU efficiency at scale for our core training stack. You will lead vital efforts in model stability and hardware enablement, managing the technical challenges of migrating large-scale Ads models to JAX and supporting next-generation TPUs. By leading cross-stack optimizations, robust debugging infrastructure, and resource-efficient ML initiatives, you will directly deliver major efficiency wins, SWE cost savings, and business impact.
Google Ads is at the forefront of AI innovation, applying machine learning and Generative AI models like Gemini to power a multi-billion dollar global business.

Our work directly impacts billions of users by protecting users from harm, improving ad quality, and optimizing campaigns for advertiser return-on-investment. We foster a culture of deep collaboration, partnering closely with teams like Google Research and DeepMind to solve complex challenges. Join us to work on state-of-the-art AI, take on problems at an unparalleled scale, and build the next generation of advertising technology.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits

Learn more about benefits at Google (https://www.google.com/about/careers/applications/benefits/).

Minimum qualifications:

- Bachelor’s degree or equivalent practical experience.
- 5 years of experience programming in Python or C++.
- 3 years of experience with Machine Learning infrastructure, ML execution frameworks (e.g., TensorFlow, JAX, PyTorch), or hardware accelerators (e.g., TPUs, GPUs).
- 3 years of experience testing, maintaining, or launching software products, and 1 year of experience with software design and architecture.
- Experience with large - scale distributed systems and performance debugging.

Preferred qualifications:

- Master's degree or PhD in Computer Science or related technical field.

- 5 years of experience with data structures and algorithms.

- 1 year of experience in a technical leadership role.

- Experience developing accessible technologies.

- Write and test product or system development code.

- Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies (e.g., style guidelines, checking code in, accuracy, testability, and efficiency).

- Contribute to existing documentation or educational content and adapt content based on product/program updates and user feedback.

- Triage product or system issues and debug/track/resolve by analyzing the sources of issues and the impact on hardware, network, or service operations and quality.

- Design and implement solutions in one or more specialized ML areas, leverage ML infrastructure, and demonstrate expertise in a chosen field.