Research
My research is about building learning and decision-making systems whose behavior can be controlled, inspected and specified, with an emphasis on whether those properties remain meaningful under realistic evaluation conditions.
Controllable and coachable agents
High-performing agents are not necessarily useful agents. I study how policies can expose behavioral controls that are meaningful to a human while preserving the competence acquired through learning. This includes agents that can be instructed, shaped, or coached without retraining them for every desired behavior.
Current programme Coachable Agents for Interactive Gameplay. Project Lead, Sony AI — Game AI, since 2023.
This programme builds on earlier work on high-performance interactive agents, including GT Sophy.
Selected work- Nature 2022
Interpretable decision making
I work on decision-making systems whose internal computation can itself carry interpretable structure. I am particularly interested in explanations that are part of the learned model — prototypes, compositional representations, logical structure — rather than explanations generated after the decision by a second model.
Selected work- ICML 2026
This State Looks Like That: Self-Interpretable Reinforcement Learning Agents using Prototype Soft Actor-Critic
- AAAI 2026
CIP-Net: Continual Interpretable Prototype-based Network
- NeurIPS 2023
Towards a Fuller Understanding of Neurons with Clustered Compositional Explanations
- ECAI 2024
Transparent Explainable Logic Layers
Neuro-symbolic and non-Markovian reinforcement learning
Many sequential tasks cannot be specified by a reward attached only to the current state. I study representations that make temporal objectives, symbolic constraints and memory explicit, so that an agent can reason about what has happened and what still has to happen, rather than leaving that structure entirely to emerge from the learned state representation.
Selected work- ECAI 2024
Neural Reward Machines
- RLC 2026
Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization
Research areas Reinforcement learning · Deep reinforcement learning · Evaluation of machine learning systems · Neuro-symbolic artificial intelligence · Explainable artificial intelligence · Agentic systems · Robotics · Continual learning