About this role
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to users' needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance.
Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale which are unique to Google, while using your expertise in coding, algorithms, complexity analysis and large-scale system design.
SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.
To learn more: check out our books on Site Reliability Engineering or read a career profile about why a Software Engineer chose to join SRE.
The Platforms and Devices team encompasses Google's various computing software platforms across environments (desktop, mobile, applications), as well as our first-party devices and services that combine the best of Google AI, software, and hardware. Teams across this area research, design, and develop new technologies to make our user's interaction with computing faster and more seamless, building innovative experiences for our users around the world.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google (https://www.google.com/about/careers/applications/benefits/).
Minimum qualifications:
- Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
- 8 years of experience analyzing global distributed systems.
- 3 years of experience managing people or teams.
Preferred qualifications:
- Experience with mobile development, application deployment.
- Proficiency in algorithms, data structures, complexity analysis and software design or expertise in Unix/Linux systems, IP networking, performance and application issues.
- Ability to inspire and motivate the engineering team to work together as a cohesive and highly productive unit. Experience in recruiting and managing a team of experienced engineers on projects.
- Ability to set and drive the strategy while providing technical guidance to the team, enabling them to execute effectively and deliver products on time and within budget.
- Capable of technical deep dives and strategic discussions with Googles executive leadership team.
- Lead a team of software and systems engineers, including iteration and task planning, manage availability and performance of mission-critical services and build automation to prevent problem recurrence.
- Lead by example and establish credibility with the quality of the teams' technical execution.
- Build relationships and influence internal customers and partner teams (distributed globally), and manage on-call rotations across continents, using a follow-the-sun model.
- Work with other engineering teams to reuse and understand existing frameworks, drive technical projects, provide leadership and take responsibility for the overall planning, execution and success of technical projects.
- Develop and grow engineering talent through mentoring, coaching, succession planning and strategies, work with other engineering teams to reuse and understand existing frameworks, and partner closely with product management and developer teams to create, drive and deliver service level objectives that enable product reliability.