About this role
Job Responsibilities:
Platform & Cloud Architecture:
- Define and evolve Platform Engineering architecture and strategy across AWS and Azure where applicable, ensuring scalability, availability, reliability, security, operability, and cost efficiency.
- Define standardized reference architectures, engineering standards, and reusable platform patterns across networking, compute, storage, IAM, Kubernetes, serverless, observability, and application delivery.
- Architect standardized cloud platforms and paved roads that enable Engineering teams to consume infrastructure and platform capabilities through secure self-service mechanisms.
- Lead architecture and design reviews and ensure solutions align with enterprise engineering, security, governance, and operational standards.
- Evaluate emerging cloud-native technologies and establish architectural direction based on technical and business requirements.
Infrastructure as Code & Automation:
- Define architecture and standards for reusable Infrastructure as Code using Terraform/OpenTofu, including modular design, versioning, testing, governance, and lifecycle management.
- Establish IaC patterns that can be consistently adopted across multiple products, environments, and cloud accounts.
- Define automation strategies that reduce manual infrastructure, deployment, and operational activities.
- Establish architecture patterns for integrating infrastructure provisioning with CI/CD and GitOps workflows.
- Define approaches for infrastructure lifecycle management, configuration validation, drift detection, policy enforcement, and automated remediation.
- Guide development of platform automation, APIs, utilities, and integrations using appropriate programming and scripting technologies.
GitOps & Continuous Delivery:
- Define and drive GitOps architecture, standards, and operating models using Argo CD, Flux, or equivalent technologies.
- Establish standards for declarative configuration, repository architecture, environment promotion, automated reconciliation, secrets management, drift detection, rollback, RBAC, and policy enforcement.
- Integrate GitOps practices with Kubernetes, Infrastructure as Code, CI/CD pipelines, security controls, and enterprise governance.
- Define reusable deployment patterns including blue/green, canary, progressive delivery, automated rollback, and policy-driven deployments.
- Guide teams in resolving complex GitOps, CI/CD, and deployment architecture challenges.
Kubernetes & Cloud-Native Platforms:
- Define architecture and engineering standards for enterprise container platforms, including Amazon EKS and Azure AKS where applicable.
- Establish Kubernetes architecture patterns covering cluster design, networking, workload isolation, identity, security, scalability, observability, storage, and lifecycle management.
- Define reusable patterns for deploying and operating cloud-native applications and shared platform services.
- Evaluate Kubernetes ecosystem technologies and establish standards based on enterprise requirements.
- Provide architectural guidance for complex Kubernetes platform, networking, security, scaling, and reliability challenges.
Platform & Developer Enablement:
- Architect Internal Developer Platform capabilities that simplify infrastructure provisioning and application delivery for product Engineering teams.
- Define reusable platform services, templates, APIs, workflows, and self-service capabilities.
- Establish golden paths/paved roads that abstract infrastructure complexity while maintaining security, governance, reliability, and operational controls.
- Partner with Engineering teams to understand developer needs and evolve platform capabilities and standards.
- Define platform success metrics covering adoption, developer productivity, automation coverage, reliability, deployment efficiency, and reduction of manual effort.
AI & Agentic Engineering:
- Define architecture patterns for incorporating AI and agentic capabilities into DevOps and Platform Engineering workflows.
- Identify high-value opportunities for AI across infrastructure automation, CI/CD, troubleshooting, incident analysis, operational support, documentation, and developer self-service.
- Evaluate enterprise AI technologies such as Amazon Bedrock, Azure OpenAI/OpenAI, or equivalent platforms and define secure integration patterns.
- Establish patterns for AI agents to interact securely with cloud services, APIs, engineering platforms, repositories, and operational tooling.
- Define identity, authorization, observability, governance, cost controls, guardrails, and human-in-the-loop mechanisms for AI-enabled automation.
- Guide reference implementations and proofs of concept to validate AI-enabled platform capabilities before broader adoption.
Security, Governance, FinOps & Reliability:
- Define cloud governance frameworks covering IAM, RBAC, security, compliance, policy as-code, tagging, and cost controls.
- Establish architectural guardrails that enable teams to operate independently while maintaining enterprise security and compliance requirements.
- Define architecture patterns for high availability, disaster recovery, scalability, resiliency, performance, and operational excellence.
- Ensure observability, including monitoring, logging, metrics, tracing, and alerting, is incorporated into platform architecture by design.
- Drive cloud cost optimization architecture and FinOps practices including visibility, allocation, budgeting, forecasting, and optimization.
- Define repeatable architecture and automation patterns supporting cloud migration and modernization initiatives.
Technical Leadership:
- Provide technical leadership and architectural direction for complex Platform Engineering and DevOps initiatives.
- Lead cross-team technical initiatives from architecture and design through implementation and operational readiness.
- Mentor engineers and help develop cloud, DevOps, GitOps, automation, Kubernetes, and architecture capabilities.
- Facilitate architecture and design discussions and guide teams through complex technical trade-offs.
- Influence engineering standards and technical decisions across teams and product areas.
- Coordinate technical efforts across engineering teams when required while continuing to operate as a senior individual contributor.
- Collaborate with Engineering leadership, Product Enablement, SRE, Security, Architecture, and product Engineering teams to align platform strategy with organizational objectives.
Required Skills & Experience:
- 10+ years of experience in DevOps, Cloud Engineering, Platform Engineering, Infrastructure Engineering, Software Engineering, or a related role.
- Extensive experience designing enterprise-scale cloud and platform architectures.
- Deep expertise in AWS architecture, services, and best practices.
- Expert-level experience with Terraform/OpenTofu, including modular architecture, reusable frameworks, governance, testing, and enterprise-scale implementations.
- Strong experience with Kubernetes and enterprise container platforms, particularly EKS and/or AKS.
- Deep understanding of Kubernetes architecture, networking, security, identity, scaling, observability, storage, and lifecycle management.
- Strong understanding of GitOps architecture and operating models with experience designing solutions using Argo CD, Flux, or equivalent technologies.
- Strong understanding of CI/CD architecture and enterprise software-delivery practices.
- Deep understanding of IAM, RBAC, cloud security, least-privilege architecture, policy-as code, compliance, and audit controls.
- Strong understanding of distributed systems, scalability, high availability, disaster recovery, resiliency, and performance architecture.
- Strong automation and scripting/programming skills using Python, Go, PowerShell, Bash, TypeScript, or equivalent technologies.
- Experience with observability, SRE, reliability, operational excellence, and cloud cost/FinOps practices.
- Demonstrated technical and architectural leadership across multiple engineering teams and complex initiatives.
- Ability to translate business and engineering requirements into scalable platform architecture, reference patterns, and technical standards.
Preferred Skills :
- Experience designing Internal Developer Platforms and enterprise developer self-service capabilities.
- Experience establishing enterprise paved roads/golden paths across multiple product teams.
- Experience implementing GitOps at scale across multiple Kubernetes clusters, environments, and teams.
- Experience with large-scale cloud migration and modernization programs.
- Experience implementing enterprise cloud governance across multi-account AWS environments.
- Experience with Azure/AKS or other multi-cloud environments.
- Experience designing AI-assisted or agentic DevOps and Platform Engineering workflows using Amazon Bedrock, Azure OpenAI/OpenAI, or equivalent technologies.
- Experience with platform engineering metrics, developer experience, and engineering productivity measurement.
- AWS, Terraform, Kubernetes, architecture, or other relevant certifications are a plus.
Key Competencies:
- Enterprise architecture and systems-thinking mindset.
- Strong ability to evaluate architectural trade-offs and establish scalable technical direction.
- Ability to balance standardization, developer experience, security, reliability, and cost.
- Demonstrated technical leadership, mentoring, and influence across engineering teams without requiring formal people-management authority.
- Ability to lead complex cross-team technical initiatives and drive alignment among multiple stakeholders.
- Strong problem-solving and decision-making skills for ambiguous and complex platform challenges.
- Strong written and verbal communication skills with the ability to communicate architecture to engineers, architects, and leadership.
- Continuous-learning mindset and ability to evaluate emerging cloud, platform, and AI technologies.
Level IV Expectations A DevOps Engineer IV - Platform Engineering should be able to:
- Define and evolve enterprise Platform Engineering architecture, reference patterns, and technical standards.
- Architect scalable, secure, resilient, and cost-efficient cloud platform capabilities across AWS and related technologies.
- Define enterprise approaches for Terraform/OpenTofu, GitOps, Kubernetes, CI/CD, automation, governance, observability, and developer self-service.
- Lead architecture and design reviews and guide teams through complex technical trade offs.
- Drive reusable paved roads and Internal Developer Platform capabilities that can be adopted across multiple product teams.
- Define architecture patterns for AI-assisted and agentic Platform Engineering capabilities with appropriate enterprise guardrails.
- Provide technical leadership for complex cross-team initiatives from design through operational readiness.
- Mentor engineers and raise engineering and architecture standards across the organization.
- Validate architecture through reference implementations, proofs of concept, and sufficient hands-on technical depth.
- Influence technical direction at an organizational level while operating as a senior individual contributor rather than a formal people manager.
Behavioral Competencies:
- Cultivates Innovation
- Decision Quality
- Manages Complexity
- Drives Results
- Business Insight