About this role
iPhone is the most popular camera in the world, and the Aesthetics Science team is focused on infusing photographic knowledge and personal preferences into the features and algorithms that power it. We bring together aesthetics research, image processing, and machine learning to help advance photography while upholding what makes it a beloved medium for self-expression. You'll have real ownership over how these creativity-enhancing features come together into a coherent, best-in-class experience.
We're seeking a self-driven, innovative research engineer to work on image-related machine learning solutions fundamental to visual experiences on iPhone and across Apple's creative ecosystem. You'll research, implement, and evaluate novel computational imaging features and multimodal solutions that shape future Apple products.
You bring solid experience designing, training, and analyzing multimodal models and evaluation frameworks end-to-end — combined with a passion for photographic aesthetics and high-quality software for visual creativity and image generation. Day-to-day responsibilities include prototyping novel algorithms, training and evaluating ML/generative AI (genAI) models, and designing custom data collection efforts.
This role involves close collaboration with our research and user study teams to design thoughtful experiments and translate the results of complex subjective studies into feature development, working alongside subject matter experts across visual media to build responsible, intelligent systems. Your work will be highly collaborative, both within the team and cross-functionally, requiring you to communicate technically and creatively toward a smarter, more inclusive camera.
What sets you apart from a traditional ML software engineer is the ability to integrate aesthetic values beyond your own into your work, and to operate confidently in spaces without clear-cut metrics. This role blurs the line between art and science and requires deep respect for both.
Minimum Qualifications
3+ years of demonstrated experience identifying novel problems and delivering viable solutions in multimodal vision-language models, computational photography, image processing, or computer vision, in industry or academia
Ability to collaborate across multi-functional teams, communicate effectively, and reason through ambiguous problems deeply and thoroughly
Excellent coding skills in Python and deep learning frameworks
Preferred Qualifications
PhD or Master's degree in Computer Science, Machine Learning, Graphics, or a related field, or equivalent industry experience
Experience developing and training models for agentic workflows, tool calling, or personalized interactions
Past involvement building or running evaluation platforms for AI use cases
Personal or professional experience with content creation, media curation, or photo/video editing
Appreciation for concepts and trends in photography, videography, and visual storytelling
Proficiency in Swift, Objective-C, or Metal
Fluency in additional languages (e.g., Mandarin) to support user research across global markets.
Ability to travel internationally for work