About this role
The FAIR Security, Privacy, and Reliability team breaks and fixes agents and the foundational models that power them. We conduct fundamental research to discover novel attacks, measure privacy, and find non-intuitive breaks in robustness. We frequently contribute those benchmarks to Meta's model Evaluation Reports and have published many of them in top research venues, earning top-1% recognition at ICLR and ICML. The team also lands novel post-training mitigations for these risks in both the open-source line of models and has multiple opportunities for direct research-to-production.
Responsibilities
Conduct fundamental research to discover novel safety and security failures, robust approaches to measuring memorization risk or reliability failures in AI agents and foundational models,
Develop and contribute benchmarks to Meta's foundational model evaluation suite and/or the Evaluation Reports
Design and implement novel post-training mitigations for safety, security, privacy, and reliability risks in foundation models
Collaborate with cross-functional teams to translate research into production systems
Qualifications
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
PhD in Computer Science, Machine Learning, or related field, or equivalent practical experience
Research experience in at least one of the following areas: safety, security, privacy, or robustness of AI models; adversarial machine learning; indirect prompt injections or jailbreaks; contextual integrity or memorization; reward hacking or other agent reliability failures; testing or mitigating foundation models for catastrophic risk (CBRNE, cyber, loss of control); or developing mitigations in any of these areas Track record of publications in top-tier research venues
