Engineer, Storage and Data Protection

AHEADGurugram, HaryanaOn-siteFull-timeJunior, 1–2 yearsListed 4 weeks ago

Apply now

About this role

Key Responsibilities

- Provide enterprise-level operational support to Managed Services customers for incident, problem, and change management activities

- Administer parallel and distributed filesystems such as Lustre, GPFS, BeeGFS, Ceph, Weka, or Vast

- Optimize storage performance, throughput, metadata operations, and data locality for AI training and inference

- Build and maintain automation for storage provisioning, monitoring, alerting, quota management, and lifecycle operations

- Plan and perform maintenance activities

- Assess customer environments for performance and design issues and propose resolutions

- Work across technical teams to troubleshoot complex infrastructure issues

- Create and maintain detailed documentation

- Serve as a subject matter expert and escalation point for storage technologies

- Work with vendors to resolve storage issues

- Communicate with customers and internal team with transparency

- Support data movement workflows including ingest, replication, caching, tiering, and archiving

- Troubleshoot storage, Linux, network, and I/O bottlenecks across storage clusters and fabrics

- Partner with infrastructure, platform, and research teams to support production AI/HPC workloads

- Evaluate new storage architectures and technologies for scalability, resilience, and cost efficiency

- Communicate with customers and internal team with transparency

- Participate in on-call rotation

Required Qualifications

- 5+ years of experience with HPC, AI infrastructure, or large-scale storage engineering

- Bachelor’s degree or equivalent Information Systems or related field. Unique education, specialized experience, skills, knowledge, training, or certification may be substituted for education

- Strong experience with Linux systems administration

- Hands-on experience configuring, managing, and tuning distributed or parallel filesystems

- Experience tuning storage for performance-sensitive workloads

- Knowledge of HPC schedulers such as Slurm and/or container platforms such as Kubernetes

- Familiarity with high-speed interconnects such as InfiniBand or RDMA

- Ability to troubleshoot complex issues across storage, compute, and networking layers

- Understanding of data protection mechanisms, including data replication, backup strategies, and disaster recovery in HPC environments

- Experience with machine learning or data science workflows in HPC environments

- Managed Services or consulting experience

- Strong background with customer service

- High level problem-solving and communication skills

- Strong oral and written communications skills

- Managed Services or consulting experience

Preferred Qualifications

- Experience supporting storage solutions for GPU clusters and AI/ML workflows

- Familiarity with object storage such as S3, MinIO, or Ceph Object Gateway

- Experience with Terraform, Ansible, Helm, or GitOps workflows

- Knowledge of observability platforms such as Prometheus and Grafana

- Experience with multi-petabyte environments, caching architectures, and storage isolation in multi-tenant systems

- Experience with machine learning or data science workflows in HPC environments

- Scripting or programming experience with Python and Bash

- Related Storage certifications are a bonus