Emulation Operations Engineer

GoogleSunnyvale, CaliforniaOn-siteFull-timeStaff, 8–12 yearsListed 1 hour ago

Apply now

About this role

Systems Development Engineering (SDE) at Google is a role where you manage services and systems at scale. SDEs creatively put their engineering discipline to use automating the mundane and reducing toil. We don’t just write code to fix bugs, but emphasize the development of tools and solutions that fix classes of problems. We know it’s hard to control what you can’t measure – so we focus on observability: instrumenting first, then turning data into knowledge, and finally knowledge into action. We know that the operational efficiency of Google systems, services, virtual compute environments and the operating systems that power them impact the environment, not just the bottom line. We know that working together we can do more, and that community matters.

Google brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.

Together we engineer and build the infrastructure, tools, access and telemetry for systems that enable orchestration of Google-scale services. Come build things that matter.

As a Systems Development Engineer, you will help maintain and improve existing systems, while helping develop next-generation solutions for future projects.

You will play a key role in ensuring the smooth operation and optimal performance of our emulation lab. This involves managing and maintaining the lab's IT infrastructure, including servers, storage, networking, and security systems.

The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $138000 - $197000 (USD) + 15% bonus target + equity + benefits

Learn more about benefits at Google (https://www.google.com/about/careers/applications/benefits/).

Minimum qualifications:

- Bachelor's degree in Computer Science or IT-related field, or equivalent practical experience.

- 3 years of experience with IT operations, Linux administration, and scripting.

Preferred qualifications:

- 3 years of experience with Linux operating systems internals and administration.

- 3 years of experience in IT infrastructure management, in an emulation lab or high-performance computing environment.

- 3 years of experience with Cloud systems design.

- 3 years of experience with one or more programming/scripting languages (Python, Bash).

- Design, implement, and maintain the emulation lab's IT infrastructure (servers, storage, networking, and workstations) and manage vendor relationships for hardware, software, and support agreements.

- Explore, gather team feedback on, and implement new emulation workflows and methodologies.

- Automate systems management with code to streamline operations, propose process improvements, and reduce support toil.

- Monitor production distributed infrastructure, participate in on-call rotations, and troubleshoot advanced technical issues within a large-scale Linux environment.