Hardware Systems Engineer

MetaAustin, TexasOn-siteFull-timeMid level, 2–5 yearsListed 2 hours ago

Apply now

About this role

Meta is seeking a Hardware Engineer to support the validation and production readiness of AI and high-performance computing infrastructure deployed across Meta's global data centers. In this role, you will work at the intersection of AI silicon, server systems, and data center operations, partnering with hardware design, firmware, software, and capacity engineering teams to bring AI accelerators and GPU-based server platforms from new product introduction through full-scale deployment. Your work will directly influence the reliability, scalability, and performance of the AI infrastructure that powers Meta's machine learning and AI initiatives.

Responsibilities

Lead hardware system validation and qualification activities for AI server platforms, GPU clusters, and AI accelerator systems entering Meta's production data center environment
Drive new product introduction efforts for AI/HPC hardware by defining acceptance criteria, coordinating bring-up activities, and managing cross-functional readiness reviews with silicon, firmware, and software teams
Develop and execute test plans covering thermal, power, and electrical characterization of AI server systems at component, board, and rack levels
Investigate and root-cause complex hardware failures across the full stack, including AI silicon, firmware, system software, and thermal subsystems
Collaborate with AI silicon vendors and original design manufacturers to resolve hardware issues, drive engineering change orders, and ensure production quality standards are met
Define and improve hardware qualification methodologies and failure analysis processes for AI/HPC platforms to scale across large fleet deployments
Partner with AI platform and capacity engineering teams to support hardware deployment planning, rack integration, and production ramp execution for GPU and AI accelerator systems
Analyze fleet telemetry, failure rate data, and field return trends to identify systemic reliability risks in AI infrastructure and drive corrective actions
Contribute to AI hardware design reviews and provide feedback on system architecture, component selection, and manufacturability to improve production outcomes
Communicate hardware qualification status, risk assessments, and deployment readiness for AI/HPC systems to cross-functional stakeholders and engineering leadership

Qualifications

Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
6+ years of experience in hardware engineering with a focus on AI servers, GPUs, AI accelerators, or high-performance computing systems in a data center environment
6+ years of experience in hardware validation, system bring-up, or new product introduction for AI/HPC platforms
Experience debugging hardware failures across multiple domains including power delivery, thermal management, signal integrity, or high-speed interconnects such as PCIe, NVLink, or HBM in AI systems
Experience developing test specifications, qualification procedures, and failure analysis documentation for AI/HPC hardware systems
Experience collaborating with silicon vendors, original design manufacturers, or contract manufacturers to resolve hardware quality and reliability issues for AI infrastructure Experience with scripting or automation for hardware test workflows, data collection, and failure triage in a Linux-based environment
Familiarity with hardware telemetry frameworks, out-of-band management protocols such as IPMI or Redfish, and fleet-level diagnostics tooling for AI infrastructure
Experience with AI/HPC hardware platforms including GPU clusters, AI accelerators, or custom AI server designs at rack or pod scale
Knowledge of high-speed interconnect standards such as PCIe Gen 4/5, NVLink, CXL, or high-bandwidth memory in the context of AI/HPC system validation