Customer Solutions Engineer, Compute, Google Cloud

GoogleSingaporeOn-siteFull-timeMid level, 2–5 yearsListed 1 hour ago

Apply now

About this role

Our Customer Solutions Engineers for AI Infrastructure own complex customer issues and provide specialized support to other teams. In this role, you will be a part of a global team that provides 24x7 support to ensure customers can seamlessly deploy their AI and ML workloads on AI Infrastructure products. When customers encounter deep technical issues, you will ensure we have the expertise, tools, and processes to resolve the issue. You will troubleshoot technical problems with a mix of hardware and software debugging, networking, Linux system administration, coding/scripting, and updating documentation. You will help our customer’s success in the AI/ML space by making improvements to the product, internal tools, processes, and documentation. You'll help drive business growth by recognizing and advocating for our customers challenges related to AI deployments.

Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.

Minimum qualifications:

- Bachelor's degree in Science, Technology, Engineering, Mathematics, or equivalent practical experience.

- Experience reading/debugging code written in a general purpose coding language (e.g., Java, C, C++, Python, Shell, Go or JavaScript, etc.) and in virtualization and orchestration frameworks.

- Experience in system administrator with Linux/Unix systems and debugging issues across the hardware/software on enterprise-grade server infrastructure.

- Experience troubleshooting and advocating for customer needs, and triaging technical issues across the stack (e.g., hardware faults, low-level software, networking, virtualization, kernel drivers, firmware, performance).

Preferred qualifications:

- Experience working directly with AI/ML computing hardware, including GPUs or other accelerators.

- Experience with ML frameworks (e.g., TensorFlow, PyTorch), and understanding of the AI/ML training and inference lifecycle.

- Experience working with large-scale distributed systems, and familiarity with common solutions, design patterns, or best practices.

- Familiarity with containerization and orchestration technologies like Kubernetes or Slurm in an on-prem or cloud environment.

- Manage customer’s problems through effective diagnosis, resolution, or implementation of new investigation tools to increase productivity for customer issues on AI/ML infrastructure.

- Develop an in-depth understanding of AI/ML workloads and underlying hardware architectures by troubleshooting, reproducing, and determining the root cause for customer reported issues, and build tools for faster diagnosis.

- Act as a consultant and subject matter expert for internal stakeholders in engineering, sales, and customer organizations to resolve complex deployment and operational obstacles in AI infrastructure environments.

- Work closely with multiple Product and Engineering teams to find ways to improve the product, and interact with our Site Reliability Engineering (SRE) teams to drive high-quality production.

- Participate in rotating on-call schedules including during nights, weekends and holidays, to ensure prompt and proper resolution of customer-impacting technical challenges.