About this role
<b>Overview</b><br><div><p><a href="https://www.microsoft.com/en-us/research/lab/spatial-ai-zurich/" target="_blank" rel="noreferrer noopener"><span>The Spatial AI Lab</span></a><span> is part of the </span><a href="https://www.microsoft.com/applied-sciences/?msockid=26afc5ddb79c66f43bb3d283b6636727" target="_blank" rel="noreferrer noopener"><span>Applied Sciences Group</span></a><span><span>,</span><span> a Microsoft research and development organization dedicated to creating next-generation human–computer interaction technologies, leveraging the latest AI developments and exploring new hardware capabilities and device form factors. Our scientists and engineers bring deep expertise in computer vision and multimodal AI, with a particular </span><span>focus on spatial and embodied AI.</span></span><span> </span></p></div><div><p><span>About the internship</span><span> </span></p></div><div><p><span>We are seeking Research Interns in Computer Vision, Machine Learning and Robotics, broadly defined, for our Spatial AI Lab in Zurich. As an intern, you will collaborate with one or more mentors and use the cutting-edge hardware and software we are creating. During the 12-week internship, you will work on a stimulating, open-ended research project that seeks to advance the state of the art. An internship at Microsoft will not only enhance your career but also let you take part in exciting research breakthroughs. Interns work, learn, and network with some of the world’s top researchers and engineers.</span><span> </span></p></div><div><p><span><span>This call covers four research directions, described below. The exact project will be shaped jointly by you and your mentors, based on research opportunities and your background.</span><span> The start date of the internship </span><span>could be the 1</span></span><span>st</span><span> of each month, no later than 1</span><span>st</span><span> of April 2027.</span><span> </span></p></div><div><p><span>Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.</span><span> <br></span></p><div><div><p><span><strong>Direction A</strong> – Multimodal Embeddings</span><span> </span></p></div><div><p><span>Text–image embedding models such as CLIP and SigLIP are foundational building blocks for search, retrieval, and multimodal AI. In this direction, you will study how these models organize information in their embedding spaces and develop new methods that make embeddings more expressive, flexible, efficient, and interpretable, with applications such as semantic search.</span><span> </span></p></div><div><p><span>Example research topics</span><span> </span></p></div><div><ul style="list-style-type: disc;"><li><p><span>Geometry and expressivity of multimodal embedding spaces (e.g., cone effect, modality gap)</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Updating embedding models while preserving compatibility with existing embeddings</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Multi-vector, hierarchical, and composable embeddings</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Explainable and steerable embeddings</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Generative search, efficient retrieval, and quantization for inference</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Benchmarking retrieval beyond web-scale data, such as private collections</span><span> </span></p></li></ul></div><div><p><span>What we look for</span><span> </span></p></div><div><p><span>Research experience in multimodal representation learning, text–vision embedding models, contrastive learning, neural information retrieval, or embedding adaptation and compression (e.g., knowledge distillation, quantization).</span><span> </span></p></div></div><div><div><p><span><strong>Direction B</strong> – Image Understanding and Editing</span><span> </span></p></div><div><p><span><span>Advances in visual intelligence are creating new opportunities for models that interpret rich image content and reason about what makes an image meaningful. In this direction, you will investigate models that build a context-aware understanding of images (scene composition, geometry, appearance, salient regions, objects, and their relationships) and explore how this understanding can power </span><span>downstream applications</span><span>. </span></span><span> </span></p></div><div><p><span>Example research topics</span><span> </span></p></div><div><ul style="list-style-type: disc;"><li><p><span>Richer visual representations and combining complementary visual signals</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Context-aware scene and image understanding, and its use for image editing</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Generalization, robustness, and reliable evaluation of visual understanding</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Efficient architectures, model compression, or on-device deployment (optional extension)</span><span> </span></p></li></ul></div><div><p><span>What we look for</span><span> </span></p></div><div><p><span>Research experience in image understanding, visual representation learning, computational photography, multimodal learning, or generative image editing, with a strong background in modern discriminative, generative, or vision–language models. Experience with efficient neural networks is a plus.</span><span> </span></p></div><div><p><span><span><strong>Direction C</strong> – Robot Learning for Manipulation </span><span>and </span><span>Real-to-Sim-to-Real</span></span><span> </span></p></div><div><p><span>Real-world observations and demonstrations are a rich source of experience for robot learning. In this direction, you will work on manipulation with stationary and mobile bimanual robots, exploring how physical experience, such as videos and demonstrations, can inform simulation, and how simulated experience can improve real-world policies.</span><span> </span></p></div><div><p><span>Example research topics</span><span> </span></p></div><div><ul style="list-style-type: disc;"><li><p><span>Real-to-sim-to-real learning and policy robustness</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Video–action models, world–action models, and vision–language–action (VLA) policies</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Learning from human demonstrations, including 3D reconstruction and tracking of hand–object interactions</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Physics-based and differentiable simulation for manipulation</span><span> </span></p></li></ul></div><div><p><span>What we look for</span><span> </span></p></div><div><p><span>Research experience in robot learning (imitation or reinforcement learning, sim-to-real transfer), bimanual or mobile manipulation, 3D reconstruction and pose estimation, or physics-based simulation; familiarity with tools like PyTorch, ROS, and MuJoCo; and experience deploying policies on real robots.</span><span> </span></p></div><div><p><span>Note: Robot-learning projects in this direction involve hands-on work with physical robots and require presence at our Zurich lab for the full duration of the internship.</span><span> </span></p></div><div><p><span><strong>Direction D</strong> – Image Generation and Efficient Visual Generative Models</span><span> </span></p></div></div><div><div><p><span><span>In this direction, you will explore compact and efficient visual generative models across image </span><span>and </span><span>video</span> <span>generation, and action-conditioned visual prediction. You will choose a focus based on your interests and background.</span></span><span> </span></p></div><div><p><span>Example research topics</span><span> </span></p></div><div><ul style="list-style-type: disc;"><li><p><span>Video generation: extending image generation to coherent sequences conditioned on text or images</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>World modeling: lightweight models that predict future visual observations from context and actions</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Architecture distillation: transferring the capabilities of large image generators into different, smaller architectures, including latent-to-pixel-space distillation</span><span> </span></p></li></ul></div><div><p><span>What we look for</span><span> </span></p></div><div><p><span>Research experience in generative modeling (e.g., diffusion or flow-based models), image or video generation, model architecture distillation, temporal learning, or world models.</span><span> </span></p></div></div><p><span> </span></p></div><br><br><b>Responsibilities</b><br><div><ul><li><span>Review relevant work and establish strong, reproducible baselines within the agreed project scope.</span><span> </span></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Prototype, train, and evaluate new models and methods for your research direction.</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Design controlled experiments covering quality, model behavior, robustness, and efficiency.</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Own the research workflow end to end, from problem formulation and implementation through analysis, iteration, and clear documentation.</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Work closely with mentors and other researchers, seek feedback, contribute to a shared codebase, and present your findings at the end of the internship.</span><span> </span></p></li></ul></div><br><br><b>Qualifications</b><br><p><strong><span>Required qualifications</span><span> </span></strong></p><div><div><ul style="list-style-type: disc;"><li><p><span>Currently pursuing a PhD in Computer Vision, Machine Learning, Robotics, AI, or a related field.</span><span> </span></p></li></ul></div><div><p><strong><span>Preferred qualifications (one or more of the following)</span><span> </span></strong></p></div><div><ul style="list-style-type: disc;"><li><p><span>At least one year of research experience in an area related to one of the research directions above.</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>A strong background in deep learning and modern neural network architectures.</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Skills in algorithmic problem-solving and software development (e.g., Python, C++).</span><span> </span></p></li></ul></div></div><div><div><ul style="list-style-type: disc;"><li><p><span>Experience with tools like PyTorch, TensorFlow, and OpenCV (for robotics: ROS, MuJoCo, or Unity).</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Publication(s) in top-tier conferences or journals in related fields (e.g., CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, IJCV, TPAMI, CoRL, RSS, ICRA, IROS, IJRR, T-RO).</span><span> </span></p></li></ul></div><div><ul style="list-style-type: disc;"><li><p><span>Excellent communication, collaboration, and writing skills.</span><span> </span></p></li></ul></div></div> <br><p>This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.</p><br><hr><br><p>Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about <a href="https://careers.microsoft.com/v2/global/en/accessibility.html"><b><u>requesting accommodations.</u></b></a></p>