My current research interests roughly lie on the general field of AI Safety, with a particular focus the intersection of interpretable/explainable AI, representation learning, and human-in-the-loop AI. More specifically, I am interested in (1) the design of methods that can construct explanations for a model’s predictions in terms of high-level “concepts” and (2) the broad applications that these methods may have in scenarios where experts can interact with the models at test time (e.g., model steering, monitoring, test-time feedback, concept interventions).
Below you can find a list of some of my publications, including their respective venues, papers, code, and presentations (when applicable). For a possibly more up-to-date list, however, please refer to my Google Scholar profile.
Publication browser
NeurIPS 2026 · Spotlight
TL;DR
Steering an LLM drags its attention off task-relevant tokens it needs, affecting its performance on long horizon tasks. SKOP steers an LLM while maintaining its attention patterns on a set of key task-relevant tokens.
InterpScience, NeurIPS 2026
TL;DR
We introduce an sparse autoencoder that learns feature-specific timescales in language model representations, enabling the discovery of persistent features.
ICML 2026 · Position paper
TL;DR
Leakage in concept representations is usually seen as a bug. In this paper I argue that well-conditioned leakage can be a feature, enabling high accuracy in incomplete tasks without sacrificing key properties of concept-based models.
ICML 2026 · Spotlight
TL;DR
Existing concept-based models rely on a single, fixed formula from concepts to task, limiting their flexibility across users with different needs and capabilities. M-CBE avoids this by learning a mixture of expert models with different functional forms, each of which can be selected by the user at inference time.
ICML 2026
TL;DR
We introduce a concept-based error slice discovery method that groups samples with shared concept prediction failures and identifies the keyword-concepts most responsible for each slice's failure-mode.
Compositional Learning Workshop, ICML 2026
TL;DR
We argue that symmetries are key to make interpretability something that can be evaluated and introduced into a model's design.
ICLR 2026
TL;DR
Concepts usually have a hierarchical structure. HiCEMs exploit this property to discover subconcepts in a pre-trained concept embedding model, which can then be introduced into the model without extra labels.
Principled Design for Trustworthy AI Workshop, ICLR 2026
TL;DR
We extend HiCEMs so that multiple levels of subconcept hierarchies can be learned and used for more fine-grained interventions.
NeurIPS 2025
TL;DR
We introduce a concept-based model that can explicitly defer predictions of intermediate concepts to human experts with different capabilities.
Preprint
TL;DR
We propose a definition of interpretability that is general, simple, and subsumes existing informal notions within the interpretable AI community.
ICML 2025
TL;DR
We show how, under distribution shifts, concept-based interventions may fail to improve model performance due to concept leakage. To solve this, we introduce a representation factorisation that prevents this effect, leading to more robust and reliable concept-based models.
ICML 2025
TL;DR
We introduce a preference-based objective for training concept-based models in the presence of concept label noise.
TMLR 2025
Also appeared at "XAI in Action: Past, Present, and Future Applications" at NeurIPS 2023.
TL;DR
We show how and why concept predictors often fail to properly capture the locality of a concept.
ICLR 2025
TL;DR
We introduce an end-to-end concept-based model that explicitly captures the causal reasoning process of its own inference.
Preprint
TL;DR
We argue that interpretability can be understood from the lens of inference equivariance. Under this view, concept-based interpretability emerges as a natural consequence to make checking equivariance tractable.
PhD thesis, University of Cambridge, 2025
TL;DR
My PhD thesis! Here I argue that previous concept-based models that support test-time feedback in the form of interventions come with several, often overlooked, limitations. I then propose a series of models and methods to overcome these limitations.
ECCV 2024 · top 15 of 8,500+ submissions · Oral, Best Paper Candidate
TL;DR
We argue that existing unsupervised debiasing methods still implicitly rely on group information for model selection. To solve this, we introduce an efficient hyperparameter-free bias mitigation approach called TAB.
ICML 2024
TL;DR
We show, empirically and theoretically, how most concept-based interpretable models learn concept representations that do not truly capture how concepts relate.
NeurIPS 2023 · Spotlight
TL;DR
We introduce a concept-based model that learns, end-to-end, to ask for help on intermediate concepts, significantly improving its performance with fewer human interventions.
AIES 2023
TL;DR
We argue that concept-based models trained to properly model annotator uncertainty are better suited for human-AI collaboration.
ICML 2023
Also appeared at ICML's Differentiable Almost Everything Workshop, 2023.
TL;DR
We introduce a differentiable neuro-symbolic model that learns, end-to-end, to make predictions using concept-based logic rules.
TMLR 2023
Also appeared at ICML's Workshop on Interpretable Machine Learning in Healthcare, 2023.
TL;DR
We introduce the first definition of a tabular concept, and propose a concept-based model to discover such concepts in a variety of tabular tasks.
AAAI 2023 · Oral
TL;DR
We introduce a new set of metrics that measure the purity of learnt concept representations.
XAI4Debugging Workshop, NeurIPS 2021 · Spotlight
TL;DR
We show that one may distill a deep neural network into an interpretable and tractable rule set in polynomial time.
IEEE Micro 2020
TL;DR
We discuss what microservices do to hardware, systems design, and the assumptions behind both.
ASPLOS 2019
TL;DR
We introduce DeathStarBench, an open suite of end-to-end microservice applications for benchmarking real systems.
No publications match those filters.