About this role
PATH current employees - please log in and apply Here
PATH is a global nonprofit dedicated to achieving health equity. With more than 40 years of experience forging multisector partnerships and with expertise in science, economics, technology, advocacy, and dozens of other specialties, PATH develops and scales up innovative solutions to the world’s most pressing health challenges.
About AI for Health
PATH’s Artificial Intelligence (AI) Initiative is our flagship technical team working across divisions to deliver best-in-class evaluation and implementation of AI in health, spanning model evaluation to product development and service innovation. We aim to use AI to address the world’s largest health inequities.
We’re looking for a research engineer to build and run evaluation pipelines that PATH will use to test whether different AI models and products can provide safe and effective mental health/wellbeing support for those most in need. This will include developing simulated users, grounded in data from South Africa, Kenya, Ghana, and the UK, and automated (LLM) judges to score simulated conversations against behaviours defined by local clinicians and lived experience experts. You will implement, maintain, and execute these pipelines in collaboration with leading AI evaluation labs. You’ll also contribute to tool development and support research across PATH’s AI portfolio, and, where time allows, lead research projects of your own.
If you’re a computer scientist and skilled engineer, with experience doing cutting-edge AI research, who wants your work to change how AI is built and used for (mental) health, then this is the role for you!
Responsibilities:
Build and Run Benchmarking Pipelines
- Contribute your expertise to deliver efficient and accurate AI evaluation design in global (mental) health.
- Run AI evaluations at scale: create and sample simulated users to hold single- and multi-session conversations with AI models and products, scoring transcripts with automated judges, and gathering results for analysis and publication.
- Work closely with our technical partners (e.g., Transluce, Digital Umuganda) to adapt and co-develop tooling for user simulation, LLM-judging, validation etc. to the countries, languages, models, and products we aim to evaluate.
- Keep our results traceable, reproducible, and (research) publication ready.
Develop Tools
- Help build and maintain tools across our work, such as tooling for user simulation or for integrating local expert input to improve evaluation accuracy.
- Write well-tested, documented code for open release, and prepare our tooling for long-term stewardship by partner organizations.
- Share knowledge and expertise with our partners and other stakeholders in the countries where we work so that they can learn how to use our tools to perform their own evaluations.
Support Research across the Portfolio
- Support PATH’s AI for Health team to deliver research by contributing ideas, methods, engineering support, and technical expertise.
- Where time allows, lead discrete research projects of your own, and publish the results.
- This is a full-time post based in London. You’ll report to PATH’s Deputy Director for AI for Health and work closely with our technical partners. This role includes occasional international travel (around 5% of time).
Required Skills and Experience
- BSc or MSc in computer science, software engineering, or a related technical field, or equivalent practical experience.
- Strong Python skills, with a record of writing well-tested, maintainable code (version control, code review, automated testing, continuous integration).
- Proven experience evaluating artificial intelligence models, e.g., an evaluation pipeline, or similar large-scale contributions to research.
- Experience with cloud infrastructure, containers, and data pipelines (e.g., AWS, GCP, or Azure; Docker).
- Ability to work independently with external technical partners, and to explain technical work clearly to clinical, research, and policy colleagues.
- Ability to handle sensitive data, including mental health content, with care.
Preferred Skills & Experience
- Understanding of the current AI safety literature, and an interest in topics relevant to alignment with positive health outcomes.
- Experience building or running healthcare-specific AI evaluations or benchmarks.
- Experience automating interaction with web or mobile apps (e.g., with Playwright for end-to-end testing).
- Experience with multilingual NLP.
- Experience contributing to or maintaining open-source software.
- Experience with frontend deployments.
We know great candidates won’t always meet every listed qualification. Research shows that some groups, on average, are more likely to self-select out if they feel they don’t meet all requirements. If you’re excited about the role and think you’d be a good fit, we encourage you to apply.
To be selected, you must have legal authorization to work in the United Kingdom.