About this role
Meta designs and deploys its own AI systems. MTIA, the Meta Training and Inference Accelerator, is Meta's family of in-house AI accelerator ASICs, deployed in production across Meta's data centers and expanding workload coverage from ranking and recommendation to generative AI. The MTIA Software team is part of the AI & Compute Foundation organization. Because the hardware is ours, the software is ours too: we build the entire stack a chip vendor would normally supply: compiler and toolchain, runtime, kernel libraries, developer tooling, and deep PyTorch integration, co-designed with the silicon teams generation over generation.
The GenAI Frameworks org owns the layer where a model meets the machine. We take frontier models, including large language, multimodal and diffusion models, and make them run efficiently on MTIA. Our work determines how fast those models respond, how much context they can afford to hold, how many of them a given amount of silicon can serve, and how quickly a newly released model can be running and trusted in production.
We are hiring a Software Engineer to own a key area of GenAI inference on MTIA, spanning serving frameworks, distributed inference, graph-mode execution and compilation, and core PyTorch integration. This is a domain-expert role: you will set technical direction for your area, drive ambiguous multi-quarter programs across team boundaries, and own the quality bar for accuracy, stability, performance, and test coverage on custom silicon.
Responsibilities
Serve as technical owner and domain expert for a key GenAI inference framework area: serving runtime, distributed inference, graph-mode execution and compilation, or core PyTorch integration
Lead ambiguous, multi-quarter technical programs end to end: technical design, execution, test strategy and CI, rollout, and production hardening across teams and org boundaries
Design, implement, and ship features in generative AI inference frameworks and the MTIA integration layers beneath them, from prototype through production deployment
Own performance for production inference workloads end to end: profile across frameworks, runtime, compiler, and kernel boundaries, identify where time actually goes, and drive the fix to the correct layer rather than the convenient one
Build and extend the distributed inference substrate: hierarchical KV caching, cross-host transfer paths, disaggregated prefetch and decode, and parallelism strategies for long-context and mixture-of-experts serving
Optimize the runtime and execution path: graph capture and replay, host-side latency, memory allocation and placement, and throughput under real traffic conditions
Enable frontier models on MTIA and validate accuracy against GPU baselines, closing correctness gaps and distinguishing real regressions from stale references
Own the quality bar for your area, ensuring accuracy, stability, latency, and throughput, and test coverage with regression detection that keeps a win from quietly eroding
Partner with kernel, compiler, and silicon teams on hardware/software co-design, producing framework-level evidence that shapes the future silicon while the design can still change
Provide technical leadership: mentor other engineers, raise the design quality through reviews, and communicate decisions clearly through design documents and cross-team reviews
Qualifications
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
Bachelor's degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience
8+ years of professional experience in systems software, ML infrastructure, performance engineering, or compiler/framework development
5+ years of hands-on experience with generative AI inference or training frameworks, or production model-serving systems
Proficiency in Python, C++, or Rust, including low-level systems code and performance-critical paths
Experience with GenAI inference optimization: prefill and decode optimization, KV-cache management and compression, batching and scheduling, and reasoning about latency and throughput tradeoffs
Hands-on experience with deep learning framework internals beyond model authoring — operator registration and dispatch, eager execution, autograd, or graph capture and compilation paths
Experience with runtime-level optimization: graph-mode execution, host-side latency reduction, memory allocation and placement
Demonstrated experience leading technical design and end-to-end delivery of frameworks or infrastructure projects, including test strategies and CI, across team boundaries
Experience working across multiple layers of a system stack (framework, compiler, runtime, hardware) and root-causing issues across component boundaries
Experience debugging numerical accuracy issues and defending accuracy and stability bars, not just performance Experience with long-context inference techniques: context and sequence parallelism, ring or streaming attention, hierarchical or offloaded KV cache
Experience with AI accelerator platforms (GPU, TPU, or custom ASICs) and bringing up a new hardware backend in a major machine learning framework
Experience with low-precision numerics and quantization (FP8, block-scaled formats, INT8/INT4), including calibration and accuracy error analysis, debugging tools for distributed systems
Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
Experience with distributed inference or training at scale: tensor, pipeline, and expert parallelism, collective communication primitives, RDMA transports, and multi-host serving, scale-up and scale-out design
Familiarity with the PyTorch compilation stack (TorchDynamo, FX IR, Torch Inductor, dynamic shapes, Triton, AOT I) and custom operator integration
Experience with performance optimization in production environments, with a track record of using profiling and tracing to understand and resolve performance bottlenecks
Experience with LLM inference serving stacks, including continuous batching, paged or chunked attention, speculative decoding, prefix caching, disaggregated prefill and decode, or mixture-of-experts routing and load balancing
Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
