Senior Site Reliability Engineer, Production Engineering

JobgetherIndiaOn-siteFull-timeSenior, 5–8 yearsListed 49 minutes ago

Apply now

About this role

Accountabilities:

- Support production Kubernetes services as part of a global 24/7 production engineering operation, including flexibility to work split-weekend shifts.

- Administer and maintain large-scale Kubernetes clusters, systems, and infrastructure while protecting service availability, integrity, reliability, and SLAs.

- Automate operational processes and continuously identify opportunities to reduce manual tasks and improve engineering efficiency.

- Use monitoring, observability, alerts, and alarms to proactively detect, prevent, investigate, and respond to production incidents.

- Analyze logs, metrics, system behavior, and infrastructure signals to troubleshoot complex issues and determine root causes.

- Lead incident management calls, coordinating timely detection, escalation, investigation, and resolution of critical production issues.

- Engage subject matter experts, service owners, and cross-functional engineering teams to resolve complex incidents efficiently.

- Develop and improve monitoring, alerting, and reliability mechanisms in collaboration with development teams.

- Perform systems administration and security monitoring across large-scale infrastructure environments.

- Apply deep knowledge of Linux, networking, Kubernetes, and cluster infrastructure to maintain reliable production services.

- Contribute to the architecture, deployment, and ongoing improvement of Kubernetes environments operating at significant scale.

- Continuously evaluate emerging infrastructure and high-performance computing technologies and identify opportunities for innovation.

Requirements:

- 7+ years of demonstrated experience administering large-scale production Kubernetes environments within high-availability Internet, cloud, or data-center environments, with strong on-premises experience preferred.

- Bachelor's degree in Computer Science, Engineering, Mathematics, or a related discipline, or equivalent professional experience.

- Advanced hands-on expertise with Kubernetes, SLURM, and large-scale cluster management.

- Familiarity with GPU/DPU hardware and high-performance computing cluster environments.

- Strong Linux systems administration experience, including DNS, DHCP, IP tables, routing, firewalls, and core Linux networking.

- Proven ability to troubleshoot and maintain services across large-scale bare-metal infrastructure.

- Experience with CI/CD technologies and tools such as Jenkins and ArgoCD.

- Scripting or programming experience in Python, Golang, or Rust is preferred but not mandatory.

- Strong understanding of observability, incident management, reliability engineering, and production operations.

- Excellent analytical and troubleshooting skills, with the ability to work effectively under pressure during complex incidents.

- Strong communication and interpersonal skills, including the ability to clearly present technical information and influence cross-functional stakeholders.

- Ability to learn new technologies quickly and adapt to evolving infrastructure environments.

- Experience architecting, building, and deploying Kubernetes environments at large scale is highly valuable.

- Passion for innovation and advanced high-performance cluster technologies is an advantage.

Benefits:

- Full-time opportunity with a remote working option in India.

- Opportunity to work on large-scale production Kubernetes and infrastructure environments.

- Exposure to advanced cloud, bare-metal, GPU/DPU, and high-performance computing technologies.

- Opportunity to work alongside SRE, DevOps, security, development, and other specialized engineering teams.

- Significant technical ownership across reliability, automation, observability, incident response, and infrastructure operations.

- Opportunity to solve complex engineering challenges at global scale.

- Continuous exposure to emerging technologies and opportunities to develop advanced infrastructure expertise.

- 24/7 production engineering environment offering substantial experience in incident management and high-availability operations.

- Compensation, healthcare, leave, and other employment benefits are provided according to the applicable employment package and location-specific terms.

How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
 Why Apply Through Jobgether? 
 
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
 
 
#LI-CL1