Principal Engineer, Nexus Capacity Engineer

GoogleSunnyvale, Seattle, California, WashingtonOn-siteFull-timePrincipal, 12–15+ yearsListed 1 hour ago

Apply now

About this role

Google runs one of the largest computational fleets in the world, with compute, storage, networking and dedicated accelerators spread across all continents. As a Principal Software Engineer for Nexus Capacity Engineering, you will pioneer the technological vision for a critical mandate. You will discover the defining technologies and forge the architectural path for Google’s planet-scale capacity delivery, ultimately empowering our largest, most demanding internal and external AI/ML customers. This position demands an industry-recognized expert capable of making high-stakes technical architecture decisions and deep-diving into complex distributed systems.

You will guide a high-impact full-stack team through technical challenges towards our mission of scalable and efficient machine and infrastructure planning for Google's data centers. You will learn about and interact with many infrastructure software and operational teams within Cloud and TI such as: Colossus, Data Science, Supply Chain, Operations Research. You will have an impact on everything related to machine capacity delivery. We do everything from cycle time reduction (using automation) to continuous optimization of deployed capacity in the fleet. This is a truly exciting time to be part of the organization.

Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $307000 - $427000 (USD) + 30% bonus target + equity + benefits

Learn more about benefits at Google (https://www.google.com/about/careers/applications/benefits/).

Minimum qualifications:

- Bachelor's degree in Computer Science or similar technical field, or equivalent practical experience.

- 15 years of experience as a software engineer.

- Experience delivering large-scale capacity planning, IaaS/PaaS solutions, or fleet management systems.

Preferred qualifications:

- Master's degree or PhD in Computer Science or a field related (e.g., Networking or Security Systems).

- Experience architecting, leading, and delivering large-scale transformations from concept to deployment.

- Deep understanding of modern AI/ML infrastructure demands (TPU/GPU topologies, accelerators) with the ability to integrate AI-driven solutions.

- Deep understanding of the latest AI and ML capabilities with a proven ability to upskill teams and integrate AI-driven solutions to improve large-scale systems.

- Ability to influence and lead without direct authority, building strong cross-organizational relationships across disparate teams (e.g., Hardware, Software, Supply Chain, and Product Management).

- Lead the design and evolution of next generation global fleet planning, defining a multi-year engineering roadmap to deliver 10x more data center capacity with ML, compute & storage resources.

- Partner across AI & Infrastructure organizations to define a unified capacity management product suite meeting AI training and inference for internal and external customers.

- Integrate cutting-edge AI/ML, operations research, and advanced mathematical optimization into capacity planning workflows to compress capacity cycles, eliminate stranded capacity, and maximize data center utilization.

- Pioneer the technological vision for a critical mandate. Discover the defining technologies and forge the architectural path for Google’s planet-scale capacity delivery, ultimately empowering our largest, most demanding internal and external AI/ML customers.