About this role
We're part of Hardware Design and Release to Production (HDRTP) in Meta's Infrastructure organization. Meta designs its own servers, and our team supports them for the 5+ years they run in Meta's data centers. That includes the compute, database and storage servers that run Facebook, Instagram, WhatsApp and Messenger, along with the GPU and AI accelerator systems that train and serve Meta's Muse models. When these servers fail, we find out why and land durable fixes. Engineers on the team specialize in one of several areas: in-depth hardware root cause analysis, automation that makes debugging and triage faster, hardware diagnostics, or datasets and pipelines that detect anomalies earlier. Your internship project will focus on one of them.
This is a 12-week internship based in Menlo Park, CA.
Responsibilities
Learn how Meta's servers are designed, including what each component does and how they connect to each other.
Get familiar with the internal systems Meta uses to provision servers, run health checks and remediate faults.
Investigate failures on production servers and trace them to a root cause.
Build software that helps the team find and fix hardware failures faster, such as triage automation, hardware diagnostics or anomaly-detection pipelines.
Own your project from design to rollout: work with the engineers who will use it, document it, and present your results to the team.
Qualifications
Currently enrolled in a full-time degree program in Computer Science, Computer Engineering, Electrical Engineering or a related field, with an expected graduation date after the internship
Programming experience in Python
Coursework or project experience in distributed systems, computer networks or computer architecture
Experience working in a Linux environment
Must be available to work onsite in Menlo Park for a 12-week internship in summer 2027 Experience debugging hardware or low-level software, such as firmware, drivers or the Linux kernel
Familiarity with server architecture: CPUs, GPUs, memory, storage and the interconnects between them, like PCIe or NVLink
Experience building tools that automate debugging, monitoring or operations work, including personal or class projects
Experience analyzing large datasets or system telemetry with SQL or Python data tools
