About this role
About the Team
We are a systems software team building the foundational software for large-scale compute platforms. We work at the hardware/software boundary across the Linux kernel, accelerators, storage, firmware, and platform validation. We value rigorous engineering, clear interfaces, measurable performance and reliability, and upstream collaboration where appropriate. The team partners closely with hardware, architecture, product, validation, and production engineering groups to move new capabilities from design through dependable deployment.
About the Role
You will analyze the performance bottleneck and optimize the Linux storage path for high-performance, reliable compute platforms. Your work will span the block layer, file systems, NVMe and device drivers, memory and I/O interactions, observability, and failure recovery. You will own substantial storage features and difficult production issues from measurement through implementation and validation, with clear accountability for latency, throughput, availability, and data integrity.
Responsibilities
- Analyze the performance bottleneck and optimize Linux storage components across block I/O, file systems, NVMe, device drivers, caching, and relevant user-space services.
- Diagnose failures or exceptions across devices, PCIe, firmware, kernel, memory, networking, runtimes, and applications; drive root-cause analysis through verified resolution.
- Characterize workloads and use tracing, profiling, telemetry, logs, and hardware counters to identify latency, throughput, utilization, and scalability bottlenecks.
- Improve reliability through fault detection, isolation, retry, timeout, reset, failover, recovery, regression detection, and graceful degradation mechanisms.
- Own storage features from workload requirements and architecture through coding, integration, qualification, rollout, and sustained support.
- Create automated correctness, performance, endurance, stress, fault-injection, and compatibility tests for new hardware and software releases.
- Partner with device, firmware, kernel, server, network, validation, production, and workload teams to deliver storage capabilities into production.
- Review designs and code, document performance and recovery behavior, and contribute upstream where maintained community interfaces are the right solution.
Minimum Qualifications
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
- 3+ years of professional experience in C or C++ systems programming on Linux.
- Deep understanding of the Linux I/O path and hands-on experience in at least two relevant areas: NVMe, HDD, Sata SSD, AHCI, HBA, RAID, block layer, file systems, device drivers, memory management, caching, or I/O scheduling.
- Experience analyzing storage performance with reproducible benchmarks, tracing tools, telemetry, and statistically sound comparisons.
- Demonstrated ability to debug production failures across kernel, firmware, device, and service boundaries.
- Experience delivering well-tested software in collaboration with hardware, validation, and production teams.
Preferred Qualifications
- Experience with high-performance or distributed storage for data-center, cloud, high-performance computing, database, or AI workloads.
- Knowledge of PCIe, NVMe over Fabrics, direct I/O, asynchronous I/O, DMA, NUMA, persistent memory, or GPU-direct storage paths.
- Experience contributing to upstream Linux storage or file-system communities.
- Evidence of delivering measurable improvements in tail latency, throughput, recovery time, device utilization, or fleet reliability.