Associate Director, Clinical AI Evaluation and Responsible Deployment

AstraZenecaBarcelona, CataloniaOn-siteFull-timeStaff, 8–12 yearsListed 40 minutes ago

Apply now

About this role

Associate Director, Clinical AI Evaluation and Responsible Deployment

About AISI

AI Science & Innovation (AISI) sits at the centre of AstraZeneca’s R&D AI transformation within Enterprise AI Unit (EAI). Our remit is to build, buy, and deliver the AI models and agents that change pipeline outcomes across discovery, translational science, biomarkers, and clinical development, ultimately improve patients’ lives.

We are building an end-to-end Enterprise AI engine that unites data foundations, AI models, platforms, and business-facing applications to accelerate results across the value chain. Success comes from reusing what already works, sharing ideas across teams, and scaling impact rather than building in isolation.

Role overview

AstraZeneca is building an AI capability for Clinical Development that will improve how trials are designed, conducted, monitored, and analysed. We are hiring an Associate Director, Clinical AI Evaluation and Responsible Deployment to create the evidence systems that determine whether clinical AI is useful, reliable, safe, and ready to scale.

The role sits in the Applied Clinical AI team and engages directly with clinical stakeholders, Engineering, IT, and data teams to shape priorities, requirements, and delivery in collaboraiton with AstraZeneca’s existing Engineering, Product and Clinical Solutions teams.

Clinical AI rarely has simple ground truth. Expert judgments can differ, source data can be incomplete, and the consequence of an error depends on where an output appears in the workflow. You will turn these realities into rigorous evaluation environments, release criteria, monitoring strategies, and validation-ready evidence. You will work directly with clinical stakeholders to define what “good” means and directly with scientists and engineers to implement it in code.

This is a hands-on technical leadership role. You will build evaluation harnesses, analyse model and workflow behaviour, design experiments, and help teams diagnose failures—not merely review documents after development is complete. You will also connect scientific evaluation with Quality, Regulatory, and operational expectations so that evidence is useful both to builders and decision-makers.

What you’ll do

- Define and implement the evaluation strategy for clinical AI models and agents across development, release, monitoring, and change control.

- Build evaluation harnesses, curated test sets, simulation environments, automated regression suites, and analysis pipelines in Python and related technologies.

- Translate expert judgment from CRAs, medical monitors, clinical scientists, statisticians, and operations leaders into task definitions, scoring rubrics, error taxonomies, and clinically meaningful acceptance thresholds.

- Evaluate complete workflows—not only model outputs—including retrieval quality, tool selection, orchestration, source fidelity, abstention, human hand-offs, latency, and downstream operational impact.

- Design approaches for noisy, sparse, delayed, or expert-dependent ground truth, including adjudication, inter-rater agreement, challenge sets, prospective studies, and post-deployment surveillance.

- Lead failure analysis and red-teaming for hallucination, unsupported claims, automation bias, data leakage, prompt injection, subgroup performance, and unsafe workflow behaviour.

- Establish traceability from intended use and user need through requirements, test evidence, and release decisions, including how model, prompt, tool, data, and workflow changes are assessed, monitored, and revalidated throughout the product lifecycle.

- Work with Quality, Regulatory, Clinical Operations, Privacy, Security, and R&D IT to align evaluation evidence with GCP, GxP, data-integrity, validation, and inspection-readiness expectations.

- Advise teams on when evidence supports progression from prototype to controlled pilot, broader deployment, or regulated use—and when it does not.

- Create reusable evaluation components and standards that can be adopted across agentic workflows for improving clinical operations and the wider Clinical Development AI portfolio.

- Communicate findings and residual risks clearly to technical, clinical, quality, and executive audiences; ensure uncertainty is visible rather than hidden behind aggregate metrics.

- Mentor scientists and engineers in rigorous experimentation, reproducibility, and responsible clinical AI development.

Essential for the role

- PhD in Machine Learning, Computer Science, Statistics, Biostatistics, Biomedical Informatics, Computational Biology, or a related quantitative discipline; or an MD, master’s degree, or equivalent experience with a strong computational record.

- 4 to 7 years of post-PhD (or equivalent) experience evaluating, validating, or deploying AI/ML systems in healthcare, life sciences, clinical research, or another safety- or evidence-critical environment.

- Current, hands-on coding ability in Python and SQL, including experience building automated evaluation pipelines, analysing large datasets, writing tests, and working in version-controlled environments.

- Strong grounding in experimental design, statistical inference, uncertainty, error analysis, and measurement reliability.

- Experience evaluating LLMs or agentic systems, including retrieval, tool use, structured outputs, multi-step workflows, robustness, and human-in-the-loop performance.

- Demonstrated ability to construct useful evaluation approaches when labels are noisy, expert opinions differ, or outcomes are delayed.

- Experience producing audit-ready evidence under GCP and GxP, including model-risk management, audit trails, and inspection readiness for regulated AI systems.

- Ability to work directly with clinical stakeholders to define intended use, harmful failure modes, decision thresholds, and acceptable human oversight.

- Deep understanding of trustworthy AI principles and lifecycle governance, and the ability to turn them into concrete release and monitoring criteria.

- Excellent communication skills and the judgment to explain technical evidence and residual risk to clinical, quality, regulatory, and executive audiences.

- Experience leading cross-functional technical work through influence, with the judgment to make evidence-based go or no-go recommendations under ambiguity.

Desirable for the role

- Direct experience with clinical trial conduct, medical monitoring, pharmacovigilance, or clinical data management in a pharmaceutical, biotech, CRO, or health-system environment.

- Experience conducting prospective, silent-mode, shadow-mode, or human-factors evaluations in live clinical or healthcare workflows.

- Familiarity with causal inference, calibration, subgroup analysis, weak supervision, active learning, or methods for learning with imperfect labels.

- Experience with LLM evaluation platforms, observability, MLOps/LLMOps, reproducible experimentation, and production monitoring.

- Peer-reviewed publications, standards work, open-source contributions, or regulator/industry-consortium engagement related to AI evaluation or medical AI.

What success looks like

- Every clinical AI release has clear intended use, measurable acceptance criteria, and traceable evidence.

- Clinical experts recognize the evaluation as representative of real work and real failure modes.

- Scientists and engineers can detect regressions quickly and diagnose why a system failed.

- Quality and regulatory partners are engaged early, with evidence generated by design rather than assembled after the fact.

- Evaluation assets are reused across products, studies, therapeutic areas, and deployment environments.

Office working requirements

When we put unexpected teams in the same room, we unleash bold thinking with the power to inspire life-changing medicines. In-person working gives us the platform we need to connect, work at pace, and challenge perceptions. That’s why we work, on average, a minimum of three days per week from the office. We balance this expectation with individual flexibility. Join us in our unique and ambitious world.

#EAI

Date Posted
08-oct-2026

Closing Date
19-oct-2026

AstraZeneca embraces diversity and equality of opportunity.  We are committed to building an inclusive and diverse team representing all backgrounds, with as wide a range of perspectives as possible, and harnessing industry-leading skills.  We believe that the more inclusive we are, the better our work will be.  We welcome and consider applications to join our team from all qualified candidates, regardless of their characteristics.  We comply with all applicable laws and regulations on non-discrimination in employment (and recruitment), as well as work authorization and employment eligibility verification requirements.