About this role
As the Senior Site Reliability Engineer, you will serve as a trusted technical resource responsible for deploying, validating, and operationalizing AI, HPC, Kubernetes, and enterprise infrastructure environments. This role transforms newly installed hardware into production-ready platforms through standardized provisioning, automation, testing, and infrastructure validation activities. Working as part of a holistic team strategy, you will support large, complex customer deployments and ensure infrastructure environments are ready for operational handoff and long-term success.
Responsibilities:
- Provide technical expertise and engagement to support infrastructure readiness, platform engineering, and deployment activities across customer environments.
- Deploy, configure, and validate AI, GPU, and High Performance Computing (HPC) infrastructure solutions.
- Prepare and administer Kubernetes platforms, container runtimes, storage integrations, networking components, and cluster infrastructure.
- Install, configure, and validate NVIDIA technologies including GPU drivers, CUDA, GPU Operators, AI Enterprise prerequisites, and telemetry solutions.
- Validate accelerated networking technologies including InfiniBand, RoCE, RDMA, and GPU-to-GPU communications.
- Perform infrastructure readiness assessments, burn-in testing, operational acceptance testing, and performance validation activities.
- Configure and support server infrastructure including iDRAC, iLO, BMC, firmware, storage, and networking components.
- Deploy and administer Windows, Linux, VMware ESXi, Hyper-V, and KVM-based environments.
- Apply security hardening standards, compliance requirements, and operational best practices throughout deployment and validation activities.
- Develop and maintain automation workflows utilizing PowerShell, Python, Bash, and Infrastructure-as-Code methodologies.
- Create customer-facing deployment documentation, technical reports, readiness assessments, and operational validation deliverables.
- Troubleshoot complex hardware, operating system, virtualization, containerization, networking, and AI platform issues.
- Participate in advanced technical training and continued education to maintain expertise in cloud, infrastructure, AI, and platform technologies.
- Support technical engagements across customer environments and collaborate with internal engineering, architecture, and service delivery teams.
Qualifications:
- Associate degree (U.S.)/College Diploma (Canada) or equivalent combination of education and technical experience required.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related technical discipline preferred.
- 5+ years of experience in Infrastructure Engineering, Platform Engineering, Site Reliability Engineering (SRE), Systems Administration, or related technical roles.
- Experience deploying, supporting, or validating AI, GPU, HPC, or large-scale enterprise infrastructure environments.
- Experience with Kubernetes, container platforms, and enterprise Linux administration.
- Strong knowledge of server provisioning, virtualization, storage, networking, and infrastructure operations.
- Experience with VMware ESXi, Hyper-V, KVM, or related virtualization technologies.
- Experience developing automation and scripting solutions using PowerShell, Python, Bash, or similar tools.
- Knowledge of Infrastructure-as-Code and automated deployment methodologies.
- Experience with NVIDIA GPU technologies, CUDA, AI Enterprise, or related AI infrastructure platforms preferred.
- Knowledge of InfiniBand, RDMA, RoCE, or high-performance networking technologies preferred.
- Demonstrated troubleshooting, root-cause analysis, and problem-solving skills.
- Possess a customer-centric mindset and strong written and verbal communication skills.
- Possess intermediate computer skills, including proficiency with Microsoft Office applications.
- Ability to travel up to 25%.
Preferred Certifications
- Certified Kubernetes Administrator (CKA)
- Red Hat Certified System Administrator (RHCSA) or equivalent Linux certification
- NVIDIA certifications related to AI, GPU, or DGX platforms
- VMware Certified Professional (VCP) or equivalent
#LI-VR1 #Hybrid