Master's Thesis: Human-Centred Evaluation of Explainable Reinforcement Learning

EricssonStockholm, StockholmOn-siteFull-timeSenior, 5–8 yearsListed 1 hour ago

Apply now

About this role

Join our Team

About this opportunity:

Reinforcement Learning (RL) is increasingly used for sequential decision-making in areas such as robotics, recommender systems, autonomous control, and telecommunications. In telecom, RL is being explored for radio resource allocation, traffic steering, energy saving, and network self-optimisation. However, RL policies are typically opaque, making it difficult for operators, engineers, and researchers to understand why an agent selected a particular action.

This is a serious obstacle to deployment in telecom, where trust, accountability, and the ability to diagnose misbehaviour are essential. Explainable Reinforcement Learning (XRL), a sub-field of Explainable AI (XAI), aims to make agent behaviour interpretable to humans.

This thesis will design and run a user study comparing Feature Importance (FI) explanations with Temporal Policy Decomposition (TPD), which explains actions through predicted future outcomes. The study will investigate whether outcome-based explanations are more useful to humans than feature-attribution explanations in an RL context.

The work corresponds to two students, 30 hp each, and can be organised into two subtracks. The students will collaborate on the user-study infrastructure and codebase. The location is Stockholm, Kista, and the preferred starting period is October 2026 to January 2027.

What you will do:

- Review XAI and XRL literature and identify appropriate metrics and evaluation protocols for explanation quality.
- Design a user-study protocol based on four conditions:

No explanation: participants see only the agent's actions.
- FI only: participants see feature-importance explanations.
- TPD only: participants see temporal-outcome explanations.
- TPD + FI: participants see both explanation types.

- Extend an existing web application to support the required XRL methods, the combined condition, and new measurements.
- Run a pilot study, refine the protocol, recruit participants, and conduct the main user study.
- Analyse the results statistically and evaluate the effectiveness of each explanation method individually and in combination.
- Write the thesis report and present the results to the research team.

The skills you bring:

- You are a Master's student in Computer Science, Human–Machine Interaction, Machine Learning, Data Science, or a related field.
- You have a foundation in machine learning, basic statistics, and data analysis.
- You have good programming skills in JavaScript and Python.
- You have good English proficiency and can communicate your findings clearly.