Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)

UvationRomaniaOn-siteFull-timeJunior, 1–2 yearsListed 3 weeks ago

Apply now

About this role

Job Overview

We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms . This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.

This is not a DevOps-focused role . We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms .

The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.

Key Responsibilities & Required Skills

Linux & Bare Metal Infrastructure

- Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
- Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
- Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
- Strong understanding of server hardware, including: BIOS/UEFI
- RAID controllers
- Firmware management
- iLO/iDRAC/IPMI
- NICs and SmartNICs
- HBA cards
- Hardware diagnostics and troubleshooting

- Experience designing, implementing, and supporting enterprise Linux infrastructure at scale

AI Factory & GPU Infrastructure

- Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
- Understanding of NVIDIA GPU technologies including: A100, H100, H200, B200, or equivalent GPU platforms
- NVIDIA DGX and OEM GPU servers
- GPU provisioning and lifecycle management
- GPU monitoring and performance optimization

- Knowledge of AI Factory architecture and infrastructure requirements
- Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
- Understanding of: GPU resource allocation and scheduling
- Multi-GPU systems
- GPU networking requirements
- High-bandwidth, low-latency infrastructure design

- Familiarity with NVIDIA ecosystem technologies such as: CUDA
- NCCL
- GPUDirect Storage
- NVIDIA Fabric Manager
- NVIDIA Base Command (preferred)

Enterprise Storage & Data Platforms

- Advanced Linux storage administration: LVM
- XFS, EXT4
- NFS
- iSCSI
- Fibre Channel SAN
- Multipath I/O

- Strong hands-on experience with Ceph , including: Cluster architecture
- MON, OSD, MDS
- RBD, CephFS, RGW
- Capacity planning
- Performance tuning
- Failure recovery

- Experience with high-performance AI storage platforms such as: WEKA
- VAST Data
- Dell PowerScale
- Pure Storage FlashBlade
- NetApp

- Understanding of: NVMe-over-Fabrics (NVMe-oF)
- RDMA
- GPUDirect Storage
- Parallel file systems
- AI data pipelines

Networking & Infrastructure

- Strong networking knowledge: Bonding
- VLANs
- Routing
- MTU optimization
- DNS
- DHCP

- Experience with high-performance data center networking: 100G/200G/400G Ethernet
- RoCE
- RDMA
- Spine-Leaf architectures

- Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
- Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting

Operations & Reliability

- Experience with high availability, clustering, and disaster recovery
- Strong troubleshooting skills across: Linux operating systems
- Hardware platforms
- GPU infrastructure
- Networking
- Enterprise storage

- Experience supporting mission-critical production environments
- Bash and Python scripting for automation and operational efficiency
- Experience creating operational documentation, runbooks, and infrastructure standards

Nice to Have

- Kubernetes infrastructure (especially AI/ML and GPU integration)
- KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
- Ansible automation
- NVIDIA Base Command Manager
- Slurm or HPC workload schedulers
- Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
- Data Center Infrastructure Management (DCIM) tools
- IPAM solutions
- AWS, Azure, or hybrid cloud exposure

We Are Not Looking For

- Candidates whose experience is primarily CI/CD pipeline engineering
- Engineers focused mainly on Terraform, GitOps, or application delivery pipelines
- Cloud-only administrators with limited bare metal, storage, or hardware experience
- Professionals whose primary expertise is software development rather than infrastructure engineering

Ideal Candidate

Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.