About this role
Fleet Management builds the automation layer for Meta's data center fleet. We own the systems that plan, orchestrate, and execute the operational workflows that keep hundreds of thousands of machines healthy, available, and efficient across every region, and increasingly across public cloud providers as well.
Our platforms cover the full lifecycle of work performed on the fleet:
- Planned maintenance and rollouts: firmware, kernel, OS, and hardware maintenance executed safely and predictably at fleet scale.
- Unplanned work and repair: detecting failures, deciding what to do about them, and driving the repair workflow through until capacity returns to production.
- Fleet lifecycle operations: turn-ups, decommissions, server and service moves, rack logistics, and capacity rebalancing.
- Safety and capacity control: deciding how much of the fleet may be unavailable at once, resolving conflicts between competing operations, rate limiting, graceful cancellation, and emergency stop.
- Autonomy: replacing human-driven operational decisions with automation and agentic workflows, backed by the observability and quality signals needed to trust them.
Practically, this means we build large-scale workflow orchestration engines: systems that model operational intent, schedule it against fleet constraints, execute it across dozens of downstream systems, and decide what to do when steps fail. The engineering problems are distributed systems problems - state machines, scheduling, conflict resolution, idempotency, partial failure, and strong safety guarantees over irreversible physical actions.
The fleet is growing fast, and the operational volume with it. We are not going to staff our way through that, so our mandate is to make fleet operations scale by an order of magnitude without a corresponding increase in human effort.
Responsibilities
Own technical direction for a significant problem area of fleet automation – define the multi-year architecture, decide what to build, consolidate or deprecate, and drive it through to production
Design and evolve orchestration engines: workflow modelling, scheduling and conflict resolution, execution state machines, and the contracts between our platform and the many systems that actually perform the work
Make large-scale automation safe. Design the guardrails – unavailability budgets, rate limiting grounded in real fleet health, graceful cancellation, handbrakes, blast-radius control – for operations that physically change production capacity
Lead through others. Decompose a broad ambiguous area into workstreams other engineers can own, review their designs, and be the technical owner the technical owner that the team's most consequential decisions route through route through
Drive alignment across a wide cross-functional surface – capacity management, provisioning, remediation, hardware and data center operations, and external infrastructure providers. Much of the work is negotiating and formalising contracts and SLAs between systems and organisations
Change the operating trajectory. Replace escalation paths, one-off scripts and per-case integrations with platforms and abstractions; make structural reductions in operational load rather than incremental patches
Raise the engineering bar: mentor other engineers, set standards, and bring credible technical framing into roadmap and investment discussions
Qualifications
BS/MS in Computer Science or equivalent practical experience
8+ years of professional software engineering experience, including significant time as the technical owner of a large production system
Demonstrated staff-level scope: you have led the design and delivery of systems spanning multiple teams, across multiple planning cycles, with measurable impact
Deep distributed systems expertise – workflow or orchestration engines, control planes, scheduling, state machines, or resource management systems
Proficiency in one or more of Python, C++, Java, Rust or Go, demonstrated through professional experience building and maintaining production systems in one or more of Python, C++, Java, Rust or Go
Track record of operating critical systems: oncall ownership, incident leadership, observability and SLO design, and structural reliability improvement
Ability to lead ambiguous, cross-organisational work to completion, and to communicate clearly in writing to both engineers and senior leadership
Experience growing other engineers through mentoring, design review and technical direction-setting Experience with constraint-based scheduling, planning or solver-backed systems
Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
Experience applying AI/agentic automation to operational workflows in production
Background in infrastructure automation, fleet or capacity management, hardware lifecycle, data center operations, or cloud infrastructure
Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
Comfort working in a large mature codebase with heavy cross-team dependencies
Experience building workflow/orchestration or job execution engines used by many internal customers
Experience consolidating or deprecating legacy systems while keeping them running
Experience integrating with heterogeneous external infrastructure APIs, including public cloud provider lifecycle and maintenance models
