About this role
About the Role
This role sits at the intersection of macOS systems engineering, fleet management, and physical data center infrastructure. You will own everything from machine provisioning and monitoring to networking and hardware lifecycle management, forming a foundational pillar of how AI infrastructure operates at scale.
What You'll Do
- Design and expand a Mac fleet with thoughtful decisions around hardware, capacity, reliability, and operational efficiency.
- Build management agents, internal tools, and automation for provisioning, configuration, deployments, health checks, and remote recovery.
- Develop reliable workflows for device enrollment, MDM, configuration profiles, security policies, and macOS updates.
- Plan and operate data center infrastructure including rack layouts, cabling, switching, routing, network segmentation, power distribution, UPS systems, and cooling.
- Diagnose and resolve problems across macOS system services, permissions, networking, application behavior, and hardware.
- Improve monitoring, incident response, recovery procedures, and documentation as infrastructure scales.
What We're Looking For
- 5 or more years in infrastructure, systems engineering, or platform engineering building and operating production systems at scale.
- Deep understanding of macOS architecture, system administration, and troubleshooting, including launchd, system services, permissions, and diagnostic tools.
- Hands-on experience building and debugging software in C, Swift, and Objective-C.
- Practical experience with MDM, Apple Business Manager, automated device enrollment, and configuration profiles.
- Experience designing or operating data center infrastructure: networking, power distribution, cooling, and hardware lifecycle management.
- Familiarity with fleet-scale challenges such as configuration drift, staged updates, failure isolation, remote recovery, and capacity planning.
- Strong automation instincts and the judgment to turn recurring operational work into reliable software.
- Comfort moving fluidly between writing code, diagnosing network issues, and working directly with hardware.
- Experience designing infrastructure with limited off-the-shelf support is a plus.
Location
On-site in San Francisco, CA.