Software Architect, Network System Validation

NVIDIARa'anana, Central DistrictOn-siteFull-timePrincipal, 12–15+ yearsListed 3 weeks ago

Apply now

About this role

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

Within NVIDIA, the Networking Business Unit (NBU) builds the high-speed interconnect — Ethernet, InfiniBand,   NVLink , and   BlueField   DPUs — that   switches   thousands of GPUs into a single AI supercomputer, moving data at the scale and speed the most demanding workloads requ ire.

NVIDIA is   looking for a  Software   Architect  to join our   Network   System   Validation   group .   You will work on   developing systems   for   valida ting   advanced networking solutions across  NVIDIA complex AI   cluster   environments.   The group is a high-performance engineering force that treats   validation   as a first-class software problem. We build   systems , frameworks, and   benchmarks   that prove our network's correctness   and performance   at scale.   This is a senior, deeply hands-on role for a technology leader who can own the   software architecture,   development   roadmap, mentor a team of   high-performance   engineers, and push NVIDIA's   network   to its   speed-of-light   limits.   This role combines   the   design   of validation methodologies   with   hands-on   development   of   testing frameworks,   automation   infrastructure,   and   advanced   debugging, analysis, and investigation   tools for   network   performance and   functionality at scale.

What you’ll be doing:

- Design and implement network validation methodologies for system-level testing
- Build the systems that orchestrate millions of concurrent IO operations, inject chaos at the infrastructure layer (latency, congestion, losses, and hardware failures), and expose the hardest-to-find system faults
- Building the systems that capture, store and process TBs of telemetry data to produce autonomies root-cause-analysis and system level insights
- Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines
- Establish engineering practices — design docs, production-grade code reviews, testing philosophy, and cross-team technical alignment
- Produce clear, data-driven reports and insights

What we need to see:

- B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience
- 12+ years of experience in developing large scale software projects
- Strong distributed system experience and system-level debugging skills
- Strong software skills: C/C++/Rust, Python (must), Bash
- Experience building automation frameworks and validation tools
- Experience building large-scale infrastructure platforms, internal developer platforms, or reliability engineering systems
- Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system behavior under stress
- Proven track record leading complex technical initiatives from architecture through delivery
- Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed

Ways to stand out from the crowd:

- Background with RDMA / RoCE / AI networking (NCCL)
- Experience with L2/L3/L4 networking protocols
- Familiarity with NVIDIA networking solutions (ConnectX , SpecX , BlueField )
- Background in performance analysis, Kubernetes, or cloud environments
- Background in chaos engineering, fault injection, or simulation systems

We have some of the most forward-thinking and hardworking people working for us. If you're creative and autonomous, we want to hear from you!

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, disability status or any other characteristic protected by law.