About this role
About the Role
This Research Engineer role sits at the intersection of privacy engineering and AI infrastructure, owning the systems that make sensitive, real-world data safe for AI training. You will design and build end-to-end anonymization pipelines that protect privacy without sacrificing the structure and signal that make data valuable for training frontier AI agents. The work is high-impact: it directly gates what data can enter production training, evaluation, and synthetic data workflows.
What You'll Do
- Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information, designing transformations based on data type and downstream use case.
- Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods.
- Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation workflows.
- Create evaluation frameworks that measure privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts.
- Design systems that remain robust to new data sources, schema drift, unusual formats, and sensitive information embedded in unexpected fields.
- Collaborate with engineering, research, operations, and customers to translate privacy requirements into practical technical policies and safeguards.
What We're Looking For
- 2+ years of hands-on experience building production data or ML systems in Python.
- Proficiency in Python with a track record of building reliable, production-grade systems.
- Hands-on experience with PII detection, removal, or anonymization.
- Experience with information extraction, named-entity recognition, classification, or related methods for detecting sensitive or rare content.
- Proven ability to build end-to-end data processing pipelines without a fully prescribed roadmap.
- Strong experimental instincts: comfortable comparing approaches across recall, precision, latency, cost, and downstream data utility.
- Solid understanding of privacy transformation techniques: redaction, masking, pseudonymization, anonymization, and synthetic data generation.
- Experience designing systems that are robust to schema drift, unusual data formats, and edge cases.
- Familiarity with privacy-enhancing technologies such as differential privacy, k-anonymity, secure aggregation, or format-preserving encryption is a plus.
- Experience with low-latency or high-throughput ML inference and data-processing systems is a plus.
- Prior work with sensitive data in healthcare, finance, or security domains is a plus.
Location
On-site in San Francisco, California, USA. Visa sponsorship is available.