AIML - Quality Engineering Manager, Responsible AI and Safety

AppleCupertino, CaliforniaOn-siteFull-timeStaff, 8–12 yearsListed 1 hour ago

Apply now

About this role

Apple's Responsible AI and Safety team focuses on innovative technologies, methodologies, and research to enable fantastic user experiences and to push the frontier of machine learning. Our team is looking to hire a leader with a strong track record in Applied Research, who is passionate about ML and foundation models with a focus on responsibility, fairness, and safety. In this role, you will lead the research and application of ML methods for technologies that power breakthrough user experiences while upholding Apple's values, privacy, and quality standards.

This role leads Apple's Quality Platform Engineering function within Responsible AI — the team responsible for giving engineering teams across Apple Intelligence fast, trustworthy safety signal before code merges, and for extending safety testing coverage to every hardware platform and form factor Apple ships. Most of this role's impact will come from two things: building a strong team and a healthy team culture from the ground up, and establishing the flywheel that connects Quality Platform Engineering to Product Evaluations & Research and Post-Ship Insights — so fast signal, coverage gaps, and pipeline improvements consistently translate into real engineering decisions and clear leadership visibility, rather than one-off dashboards nobody acts on.You should be technically fluent — comfortable with evaluation pipelines, production ML and agentic systems, and cross-platform testing — enough to earn credibility with the team, ask sharp questions, and evaluate hard tradeoffs with good judgment. This role moves at a fast pace, and you should be comfortable making product recommendations in ambiguous situations, often with limited or imperfect data. Just as important are excellent communication, strong product sense to prioritize a team's limited capacity against Apple's highest-risk platforms and use cases, and a track record of hiring and developing technical teams.Prior experience managing or building technical teams is required. Prior exposure to ML evaluation, test infrastructure, or safety-adjacent engineering work is strongly preferred.

Minimum Qualifications

5+ years of technical team management or leadership experience
Experience with ML evaluation, production ML systems, or test/CI infrastructure at scale
Strong engineering skills and experience writing production-quality code (Python or similar)
Experience working across multiple platforms or hardware form factors, or a demonstrated ability to ramp quickly across unfamiliar platforms
Experience working with human-labeled or crowd-sourced evaluation data, including reasoning about label noise and inter-rater agreement

Preferred Qualifications

Experience working on Responsible AI, AI safety, or trust & safety-adjacent engineering
Experience with generative model evaluation and common failure modesStrong organizational and operational skills working with large, multi-functional, diverse teams
MS or PhD in Computer Science, Machine Learning, Statistics, or related field, or equivalent experience Familiarity with hardware/platform-specific testing considerations (e.g., on-device constraints, new form factors)