About this role
Job Summary
Performs advanced systems and infrastructure engineering functions, including planning, designing, and implementing cloud and on-premise platform solutions. This includes container orchestration (Kubernetes/OpenShift), servers, storage, networking, and backup solutions.
Leads the design and implementation of highly available, scalable, and resilient infrastructure aligned with enterprise architecture standards, including multi-region deployments and disaster recovery strategies. Develops and maintains Infrastructure as Code ( IaC ) and Configuration as Code to automate the provisioning and management of infrastructure and application platforms.
This role blends systems engineering, software development, DevOps practices, and emerging AI-enabled automation capabilities to deliver reliable, scalable, and intelligent solutions across the enterprise.
Primary Activities and Responsibilities
Systems Engineering & Platform Ownership
- Design, implement, and manage enterprise Kubernetes/OpenShift platforms with a focus on resiliency, scalability, and operational efficiency
- Define and maintain platform standards, including lifecycle management, patching, upgrades, and policy enforcement
- Engineer infrastructure and operational workflows as code, driving automation, self-healing, and scalability (using tools like Terraform, Kubernetes Operators, etc .)
- Develop, maintain, and evolve internal platforms and tools to simplify deployment, monitoring, and operational health, improving the experience for application development teams.
- Collaborate across teams ( MidTier , Dev, Security) to deliver secure, performant, and highly-available systems.
Cloud, DevOps & Automation
- Design and implement CI/CD pipelines using tools such as Jenkins
- Support GitOps deployment models leveraging tools such as ArgoCD
- Develop Infrastructure as Code (Terraform) and Configuration as Code (Ansible) to automate provisioning and management
- Automate installation, configuration, and lifecycle management of Kubernetes and supporting infrastructure
- Implement monitoring, logging, and reporting to ensure system health and performance
Software Engineering & Developer Enablement
- Develop and maintain automation tools, services, and APIs using Java and/or Python
- Build scripts and reusable automation to improve platform reliability and developer productivity
- Collaborate with development teams to enable efficient application deployment and integration
- Apply software engineering best practices, including modular design, testing, and version control
AI-Enabled Engineering & Intelligent Automation
- Leverage AI-assisted development tools to accelerate engineering workflows and improve productivity
- Apply prompt engineering techniques to create effective, repeatable interactions with AI tools for automation and problem solving
- Contribute to development and reuse of standardized AI “skills” or prompt patterns for engineering tasks
- Demonstrate familiarity with agentic AI concepts, including multi-step workflows driven by AI agents and tool integration
- Understand how external tools and services can be integrated into AI-driven workflows (e.g., MCP-style tool integration and context sharing)
- Evaluate opportunities to incorporate AI-driven automation into infrastructure and DevOps processes
Architecture, Planning & Delivery
- Lead or contribute to infrastructure design aligned with enterprise architecture, including high availability and disaster recovery
- Establish project milestones and drive execution to meet time, cost, and quality objectives
- Analyze, design, and implement solutions in complex and evolving environments
- Provide technical leadership and guidance across teams
Operational Excellence & Support
- Investigate and resolve complex service issues across systems, platforms, and teams
- Perform root cause analysis and implement preventative solutions
- Participate in on-call rotation and provide production support
- Maintain operational documentation, standards, and runbooks
Leadership & Continuous Improvement
- Mentor team members and share knowledge across the organization
- Influence stakeholders, partners, and peers on infrastructure and platform best practices
- Stay current with industry trends, including DevOps, cloud-native platforms, and AI engineering practices
- Drive continuous improvement and innovation across systems and processes
Minimum Qualifications
- Bachelor’s degree in Computer Science , Engineering, Information Systems, or related field (or equivalent experience)
- 5+ years of experience in systems engineering, infrastructure engineering, or DevOps roles
- Strong experience with Kubernetes and/or OpenShift
- Strong experience with Linux system administration
Equivalent Minimum Qualifications
- High School Diploma/GED with 10+ years of relevant experience
Preferred Qualifications
- 5+ years of Kubernetes/OpenShift administration in enterprise environments
- Experience with cloud platforms (Azure preferred)
- Microsoft Azure Administrator (AZ-104) certification
- Red Hat Certified Engineer (RHCE) or equivalent Linux certification
- Experience with CI/CD and GitOps tooling (Jenkins, ArgoCD )
- Experience developing automation using Java or Python
- Exposure to AI-assisted development, prompt engineering, or intelligent automation frameworks
Knowledge and Skills
Systems & Infrastructure
- Experience deploying and maintaining enterprise container platforms (OpenShift, Rancher, EKS, GKE)
- Strong Linux administration skills (filesystem, networking, permissions, logging, scheduling)
- Understanding of networking fundamentals including switching, routing, and load balancing
- Familiarity with infrastructure performance, capacity planning, and forecasting
DevOps & Automation
- Experience with Infrastructure as Code (Terraform) and configuration management tools (Ansible, Ansible Automation Platform)
- Experience with CI/CD pipelines and automation frameworks (Jenkins, Git-based workflows)
- Experience implementing GitOps patterns using tools such as ArgoCD
- Experience leveraging APIs and scripting to automate operational processes
Software Engineering
- Proficiency in Java and/or Python for building automation, services, and integration tooling
- Experience with scripting and automation (Python, Ansible, PowerShell)
- Understanding of application architecture and development best practices
AI, Prompting & Emerging Capabilities
- Understanding of prompt engineering concepts and how to structure inputs to improve AI output quality
- Familiarity with reusable AI “skills,” prompt patterns, or workflow-based automation
- Awareness of agentic AI concepts, including orchestration of multi-step tasks using AI agents
- Familiarity with tool integration models (e.g., MCP-style approaches) for connecting AI systems to external services
- Interest in applying AI to improve developer productivity, automation, and operational efficiency
Job Requirements
- Participation in a rotating on-call schedule, including support outside standard business hours
- Work schedule may vary based on operational needs
- This is an on-site position located at CSX corporate headquarters