About this role
Job Responsibilities:
CI/CD & GitOps :
- Design, build, maintain, and enhance complex CI/CD pipelines using Jenkins, GitHub Actions, Azure DevOps, or similar tools.
- Implement reusable pipeline libraries, templates, quality gates, security scanning, automated testing, packaging, and deployment workflows.
- Implement and support GitOps-based deployment workflows using Argo CD, Flux, or equivalent technologies.
- Maintain declarative application and platform configuration, automated reconciliation, environment promotion, drift detection, rollback, RBAC, and secrets integration.
- Implement deployment strategies such as blue/green, canary, progressive delivery, and automated rollback where appropriate.
- Troubleshoot complex CI/CD, GitOps, and deployment issues and implement permanent corrective solutions.
AWS & Cloud Infrastructure:
- Implement and improve scalable, secure, resilient, and cost-efficient infrastructure primarily in AWS.
- Work with AWS services including EC2, ECS, EKS, IAM, VPC, ALB/NLB, S3, RDS, Route 53, CloudWatch, SSM, Lambda, and related platform services.
- Troubleshoot complex infrastructure, networking, IAM, DNS, security, performance, and application connectivity issues.
- Implement established cloud architecture, security, tagging, governance, reliability, and operational standards.
- Contribute implementation expertise, technical options, and proof-of-concept validation during solution and design reviews.
Infrastructure as Code:
- Design, develop, test, and maintain reusable infrastructure using Terraform/OpenTofu.
- Create and enhance modular IaC frameworks used across multiple products, environments, and AWS accounts.
- Implement Terraform state management, testing, validation, versioning, and lifecycle practices.
- Integrate infrastructure provisioning with CI/CD and GitOps workflows.
- Troubleshoot complex Terraform plan/apply issues, state issues, drift, and dependency problems.
- Participate in and provide technical feedback during peer reviews of Infrastructure as Code changes.
Automation & Platform Engineering :
- Develop automation and platform tooling using Python, Go, PowerShell, Bash, TypeScript, Ansible, or equivalent technologies.
- Build reusable automation, APIs, utilities, templates, and workflows that reduce manual engineering and operational activities.
- Implement Internal Developer Platform capabilities and self-service workflows based on established platform architecture and standards.
- Build and improve paved roads/golden paths that simplify infrastructure provisioning and application delivery.
- Maintain and improve artifact repository, configuration-management, and engineering automation capabilities.
Containers & Kubernetes:
- Build, configure, operate, and improve Kubernetes platforms, particularly Amazon EKS and Azure AKS where applicable.
- Implement standards for cluster configuration, networking, workload isolation, identity, security, scaling, storage, observability, and lifecycle management.
- Build and maintain reusable Helm charts and Kubernetes deployment patterns.
- Troubleshoot complex Kubernetes, container, networking, storage, identity, resource, and application deployment issues.
- Automate Kubernetes provisioning, upgrades, configuration, and operational activities.
Monitoring, Reliability & Troubleshooting :
- Implement monitoring, logging, metrics, tracing, dashboards, and alerting using Datadog, Grafana, Prometheus, CloudWatch, or similar platforms.
- Lead troubleshooting and root-cause analysis for complex infrastructure, platform, CI/CD, and deployment issues.
- Implement high-availability, resiliency, disaster-recovery, scalability, and performance patterns defined for platform workloads.
- Identify recurring operational issues and implement automation or permanent corrective actions.
- Create and improve operational runbooks, health checks, recovery procedures, and platform documentation.
- Participate in an on-call rotation where applicable.
AI & Engineering Automation:
- Apply AI and agentic capabilities to practical DevOps and Platform Engineering use cases, including automation, CI/CD, troubleshooting, incident analysis, documentation, and developer self-service.
- Build or integrate AI-enabled workflows using approved enterprise AI services such as Amazon Bedrock, Azure OpenAI/OpenAI, or equivalent platforms.
- Implement secure integrations between AI capabilities and approved cloud services, APIs, repositories, and engineering tools.
- Apply appropriate identity, authorization, observability, guardrails, cost controls, and human-in-the-loop mechanisms to AI-enabled automation.
Technical Leadership & Developer Enablement:
- Provide technical guidance within Platform Engineering initiatives and lead defined technical workstreams when required.
- Mentor engineers and share expertise across cloud, DevOps, GitOps, Kubernetes, automation, and platform engineering practices.
- Participate in design, code, infrastructure, and operational-readiness reviews and provide actionable technical feedback.
- Collaborate with architects and senior engineers to evaluate implementation approaches and architectural trade-offs.
- Partner with Engineering, Platform, CloudOps, SRE, Security, and other teams to deliver reliable platform capabilities.
Required Skills & Experience:
- 7+ years of experience in DevOps, Cloud Engineering, Platform Engineering, SRE, Infrastructure Engineering, Software Engineering, or a related role.
- Strong hands-on experience with AWS and enterprise cloud infrastructure.
- Advanced experience with Terraform/OpenTofu and Infrastructure as Code, including reusable modules and multi-environment implementations.
- Strong experience designing, building, and troubleshooting CI/CD pipelines using Jenkins, GitHub Actions, Azure DevOps, or equivalent platforms.
- Strong hands-on experience with GitOps deployment tools such as Argo CD, Flux, or equivalent technologies.
- Strong hands-on experience with Docker, Kubernetes, EKS and/or AKS, and Helm.
- Experience with configuration management and automation using Ansible or equivalent technologies.
- Strong scripting/programming skills using one or more of Python, Go, PowerShell, Bash, TypeScript, or equivalent languages.
- Experience with Git and pull-request-based development workflows.
- Strong understanding of IAM/RBAC, cloud security, networking, DNS, load balancing, VPCs/subnets, routing, and security groups/firewalls.
- Experience with observability platforms such as Datadog, Grafana, Prometheus, CloudWatch, or equivalent tools.
- Ability to independently troubleshoot complex infrastructure, Kubernetes, CI/CD, GitOps, and deployment issues.
- Demonstrated experience mentoring engineers and providing technical guidance.
Preferred Skills :
- Experience building or contributing to Internal Developer Platforms and developer self service capabilities.
- Experience with multi-account AWS environments and enterprise cloud governance.
- Experience with Azure and AKS or other multi-cloud infrastructure.
- Experience with AI/LLM platforms and agentic automation, including Amazon Bedrock, Azure OpenAI/OpenAI, or equivalent technologies.
- Experience with policy-as-code, security scanning, SonarQube, Trivy, or similar engineering controls.
- Experience with cloud cost optimization and FinOps practices.
- AWS, Terraform, Kubernetes, or other relevant certifications are a plus.
- Experience working in an Agile/Scrum environment.
Key Competencies:
- Strong technical troubleshooting and problem-solving skills.
- Automation-first and platform-engineering mindset.
- Ability to independently own complex technical requirements from design through implementation and operational readiness.
- Ability to translate architecture and engineering standards into reliable production implementations.
- Ability to mentor engineers and lead defined technical workstreams without formal people-management responsibility.
- Strong collaboration skills across Engineering, Platform, CloudOps, SRE, Security, and Architecture teams.
- Good written and verbal communication skills and ability to explain technical trade-offs.
- Continuous-learning mindset and ability to evaluate emerging technologies through hands-on implementation.
Level III Expectations A DevOps Engineer III - Platform Engineering should be able to:
- Independently own and deliver complex DevOps and Platform Engineering work.
- Design and implement advanced CI/CD and GitOps workflows using established enterprise standards.
- Build and enhance reusable Terraform/OpenTofu modules, automation, and platform tooling.
- Implement and troubleshoot complex AWS and Kubernetes/EKS platform solutions.
- Build self-service capabilities and reusable paved-road patterns for Engineering teams.
- Lead root-cause analysis and drive permanent technical fixes for complex platform issues.
- Contribute implementation expertise and technical recommendations to architecture and design decisions.
- Apply AI-assisted and agentic automation to appropriate engineering use cases.
- Mentor other engineers and lead technical workstreams when required.
- Improve existing frameworks, standards, automation, reliability, and developer experience.
Behavioral Competencies:
- Cultivates Innovation
- Decision Quality
- Manages Complexity
- Drives Results
- Business Insight