About this role
Job Responsibilities:
Cloud Infrastructure & Operations:
- Build, operate, and troubleshoot infrastructure across AWS and Azure in support of production workloads.
- Operate and maintain Kubernetes clusters, including deploying and maintaining Helm charts for the services you support.
- Participate in on-call rotation, respond to incidents, and drive them to resolution within your area of ownership.
- Contribute to capacity planning, cost optimization, and resilience improvements for the systems you support.
Automation & Continuous Delivery:
- Build and maintain GitOps-based deployment pipelines using Argo CD/Argo Workflows, including rollout and promotion configuration across environments.
- Write and maintain Infrastructure-as-Code (Terraform, OpenTofu) for the infrastructure you own, following team module standards.
- Build and maintain CI/CD pipelines in Jenkins, improving build/deploy automation and reliability.
- Support progressive delivery practices (blue-green/canary, automated rollback) for the services you support.
Reliability & Observability:
- Build and maintain Datadog dashboards, monitors, and alerts for the services you support, tuning alert thresholds to reduce noise.
- Contribute to defining SLIs/SLOs for your services and help track them over time.
- Participate in postmortems for incidents you're involved in, and follow through on assigned remediation items.
Collaboration & Mentorship:
- Partner with engineers across the SRE team and with product engineering teams to troubleshoot issues and improve system design.
- Share knowledge with and mentor less-experienced engineers on the team (CloudOps Engineer I III) on cloud infrastructure, Kubernetes, and CI/CD practices.
- Contribute to documentation, runbooks, and onboarding materials for the systems you support.
Required Qualifications:
- 10+ years of experience in Cloud Operations, Site Reliability Engineering, DevOps, or Infrastructure Engineering roles.
- Hands-on experience with AWS — you can build, troubleshoot, and operate cloud infrastructure directly.
- Hands-on experience with Kubernetes and Helm — deploying, operating, and troubleshooting workloads in production clusters.
- Hands-on experience with Argo CD/Argo Workflows for GitOps-based continuous delivery.
- Hands-on experience with Infrastructure as Code (Terraform, OpenTofu).
- Hands-on experience with Jenkins and Rancher for CI/CD pipeline development and maintenance.
- Hands-on experience with Datadog (or equivalent observability platform), including building dashboards, monitors, and alerts.
- Experience participating in an on-call rotation and responding to production incidents.
- Strong communication skills and the ability to work effectively across teams.
Preferred Qualifications:
- Experience supporting payments, fuel/retail, or loyalty platforms, or other systems with PCI DSS or similar compliance obligations.
- Relevant certifications such as CKA/CKAD, AWS Certified Solutions Architect – Associate, Microsoft Certified: Azure Administrator, or HashiCorp Terraform Associate.
- Experience with messaging systems (Kafka/SQS/SNS) and multi-region/multi-AZ resilience patterns.
- Prior experience mentoring junior engineers or leading small technical initiatives.
Behavioral Competencies:
- Cultivates Innovation
- Decision Quality
- Manages Complexity
- Drives Results
- Business Insight