About this role
Meta is seeking a Software Engineer to join the MTIA (Meta Training & Inference Accelerator) Software Tooling team, which develops and maintains the tooling ecosystem for Meta's in-house AI accelerator ASICs. The Tooling team provides debugging, profiling, memory analysis, and monitoring capabilities for the whole MTIA Ecosystem, redefining ML accelerator tooling by leveraging Meta's full-stack ownership from silicon specs to fleet observability.
In this role, you will design, build, and maintain developer tools that help engineers debug, profile, measure, and monitor AI workloads running on MTIA hardware at scale. You will work at the intersection of compilers, runtime, hardware, and ML frameworks, collaborating with cross-functional partners to deliver a high-quality developer experience for Meta's custom AI accelerators.
Responsibilities
Design and develop software tools across the MTIA tooling ecosystem, including debugging, profiling, performance analysis, memory analysis, and monitoring infrastructure
Own features and components end-to-end, from design through implementation, testing, and deployment
Build and improve tooling infrastructure that enables engineers to efficiently develop, test, and optimize AI workloads on MTIA hardware
Collaborate with the MTIA compiler, runtime, kernel, and hardware teams to integrate tooling hooks and support new platform capabilities
Proactively identify gaps in the tooling ecosystem and propose solutions to improve developer productivity
Contribute to technical design discussions, write design documents for medium-scope projects, and participate in code reviews
Partner with internal product teams across advertising, recommendations, and generative AI to understand developer pain points and improve tool usability
Qualifications
Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
Bachelor's degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience
2+ years of experience in software engineering, with exposure to systems software, developer tooling, or infrastructure
Proficiency in C++ and Python, including systems-level programming concepts
Experience working across multiple layers of a system stack (e.g., application, runtime, OS/driver, or hardware interfaces)
Track record of independently delivering software projects from design through production deployment
Experience debugging and resolving issues in complex software systems (e.g., using log analysis, stack traces, or system-level diagnostic tools) Exposure to accelerator ecosystems (GPU/CUDA, TPU, custom ASICs) or heterogeneous computing environments
Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
Experience building developer tools such as debuggers, profilers, build systems, CLI tools, monitoring dashboards, or diagnostic utilities
Contributions to open-source projects demonstrating tooling or system software interest
Experience in using data-driven methods to evaluate tooling effectiveness and to prioritize improvements
Familiarity with Linux debugging and profiling tools (gdb, perf, strace, eBPF, valgrind) or similar diagnostic infrastructure
Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
Familiarity with ML frameworks (PyTorch, TensorFlow) or compiler infrastructure (LLVM, MLIR, TVM)
Experience with distributed systems debugging, profiling, or monitoring at scale
Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
