About this role
As part of Hardware Design and Release to Production (HDRTP) within Meta's Hardware Engineering organization, our team works across all NPI server projects (including cutting-edge AI training and inference hardware platforms) from early bring-up into large-scale production. We build the test automation, tooling, and diagnostics that validate GPU and accelerator platforms at scale, collaborating closely with internal teams and external vendors.
This is a 12 week internship based in Menlo Park, CA.
Responsibilities
Develop and extend test automation tooling across hardware servers and distributed systems to improve test coverage, reliability, and platform insight.
Collaborate with cross-functional teams and external vendors to integrate feedback and validate system behavior against ground-truth datasets.
Analyze large-scale internal and external datasets to identify gaps, enhance tool accuracy, and quantify performance improvements.
Present technical findings to engineering stakeholders, document implementations, and hand off project deliverables at the conclusion of the internship.
Qualifications
Currently enrolled in a full-time degree program in Computer Science, Computer Engineering, Electrical Engineering or a related field, with an expected graduation date after the internship
Programming experience in Python
Coursework or project experience in distributed systems, computer networks or computer architecture
Experience working in a Linux environment
Must be available to work onsite in Menlo Park for a 12-week internship in summer 2027 Experience with test automation frameworks, build systems or CI tooling
Familiarity with datacenter or cluster networking concepts such as topology, fabrics and interconnect
Experience querying, cleaning and analyzing structured datasets and data analysis tooling
Exposure to hardware validation, manufacturing test data or systems bring-up
Prior internship or research experience involving large-scale systems, telemetry or infrastructure tooling
