Staff Software Engineer, AI/ML GenAI, Search Intelligence

GoogleMountain View, CaliforniaOn-siteFull-timeStaff, 8–12 yearsListed 2 hours ago

Apply now

About this role

Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.

We are seeking a Staff Software Engineer with a deep background in machine learning and data science to serve as a technical lead. You will evaluate AI Overviews for factual accuracy and its impact on user experience, playing a critical role in delivering the most factual AI globally.

You will set the architectural direction and scientific excellence for evaluating Search’s AI products. Bridging advanced statistical methodology with production-scale software engineering, you will architect trusted automated evaluation systems and deploy company-wide standards at Search-wide scale.
In Google Search, we're reimagining what it means to search for information – any way and anywhere. To do that, we need to solve complex engineering challenges and expand our infrastructure, while maintaining a universally accessible and useful experience that people around the world rely on. In joining the Search team, you'll have an opportunity to make an impact on billions of people globally.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits

Learn more about benefits at Google (https://www.google.com/about/careers/applications/benefits/).

Minimum qualifications:

- Bachelor's degree or equivalent practical experience.

- 8 years of experience in software development.

- 5 years of experience testing, and launching software products, and 3 years of experience with software design and architecture.

- 5 years of experience leading ML design and optimizing ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).

- 2 years of experience with state of the art GenAI techniques (e.g., LLMs, Multi-Modal, Large Vision Models) or with GenAI-related concepts (language modeling, computer vision).

Preferred qualifications:

- Master’s degree or PhD in Computer Science, Applied Mathematics, Statistics, Data Science, Machine Learning, or a quantitative field.

- 8 years of experience architecting and operating large-scale distributed systems or ML infrastructure in Python, C++, or Java.

- Experience with statistical modeling, hypothesis testing, sampling theory, reliability metrics, and data science tools (e.g., Pandas, PyTorch).

- Experience designing measurements that capture GenAI user behavior, conversational experiences, and multidimensional product quality.

- Experience leading major engineering initiatives to production deployment and driving technical alignment across teams.

- Define the long-term technical roadmap and architectural standards for automated evaluation infrastructure.

- Establish company-wide metric definitions and taxonomies that faithfully capture user behavior and the multidimensional quality of GenAI experiences.

- Lead the design and implementation of highly scalable, production-grade autorater systems.

- Drive advanced statistical methodologies to validate autorater accuracy against production user signals and labeled ground truth.

- Partner with cross-functional teams to integrate statistically significant evaluations into the systems, while mentoring engineers in measurement best practices.