About this role
We are seeking an experienced software engineer to join the MTIA (Meta Training and Inference Accelerator) software team, to work on simulators and other early software development platforms. The simulators you build will be critical for pre-silicon development, allowing the whole team to develop software before hardware is available. This is an opportunity to work with a dedicated engineering team, collaborating with a large set of cross-functional partners. The program carries top-level executive visibility as a key AI infrastructure initiative, and the team values deep technical contributions with extensive knowledge-sharing.
Responsibilities
Contribute to our developer infrastructure, including simulation and HW emulation platforms, to enable performance measurement and optimization for Meta’s in-house accelerator programs
Support networking and compute hardware acceleration techniques to improve ML inference and training model performance
Implement simulation models for Meta’s Accelerator ASICs, develop and analyze various scenarios to evaluate data center performance and identify potential improvements
Perform an architectural analysis to ensure that system designs meet performance, scalability, and reliability requirements
Collaborate with architects and engineers to integrate simulation results into system design processes
Use instruction set simulators to define performant firmware for Meta's training/inference accelerators
Collaborate with hardware teams to ensure accurate modeling and simulation of accelerator functionalities
Qualifications
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
Masters or PhD in Computer Science, Computer Engineering, or any other relevant technical field
5+ years of experience in developing C++ codebase
5+ years experience in developing Python codebase
Understanding of performance and benchmarking measurement and optimization on collective communications and distributed at-scale model training and inference Experience with SystemC, Qemu, Virtual Platform
Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
Experience in one or more of the following machine learning/deep learning domains: hardware accelerators, AI Infrastructure, and/or high performance computing (HPC), particularly pertaining to interconnect and collective
Full-stack experience and understanding of AI/HPC systems, from HW/infrastructure through the application layer, performance optimizations, including familiarity with relevant tools, libraries, and frameworks (e.g., NCCL, PyTorch, CUDA)
