Junior Research Fellow in Computer Science

Mateo Espinosa Zarlenga

  • My research interests all roughly lie around topics in Artificial Intelligence (AI) safety. Concretely, I am interested in interpretability, representation learning, monitoring, and human-AI interaction.
  • I absolutely love the community aspect of Oxford’s tutorial and supervision system. Having the chance to engage in small-group discussions with students about interesting problems is a truly unique way to gain a deep understanding of a topic.
  • I am currently working, alongside close collaborators, on a textbook on AI interpretability, exploring the underlying principles for designing powerful AI models that we can truly understand and steer.
Image

Profile

I am a Junior Research Fellow (JRF) in Computer Science at Trinity College, Oxford. Prior to Trinity, I completed my PhD as a Gates Cambridge Scholar at the University of Cambridge, where I spent four lovely years exploring how to design AI models that can receive intermediate hints from experts during deployment. Before my PhD, I completed an MPhil in Computer Science at the University of Cambridge, worked for three and a half years at an AI startup in Silicon Valley, and earned an MEng and BA in Computer Science against the cold but beautiful backdrop of Cornell University’s campus.

Teaching

I have previously supervised undergraduate computer science classes (e.g., Discrete Mathematics, Introduction to Artificial Intelligence, Functional Programming). I am interested in continuing these efforts both for Trinity students and, more generally, for Oxford computer science students.

Beyond undergraduate teaching, I have been involved in teaching and designing a master’s-level course (Cambridge’s “Explainable AI” course) and have co-supervised and supervised several master’s-level research projects, several of which have subsequently resulted in publications. If you are interested in supervision for a research project, e.g., a research-focused undergraduate thesis or a masters project, please do not hesitate to reach out!

Research

Large Language Models (LLMs), and automated agents driven by them, are this decade's World Wide Web. They are becoming increasingly integrated into every aspect of our lives, whether we want that or not, and have revolutionised the field of Artificial Intelligence (AI). Nevertheless, the inability of LLMs to (1) indicate uncertainty on intermediate steps of their reasoning, (2) properly update their outputs or behaviours when intermediate steps are corrected or steered, and (3) generate thinking traces that faithfully reflect their reasoning process, all remain a significant barrier to their deployment in high-stakes environments. My research aims to enhance the reliability of LLMs by developing mechanisms that enable us to effectively incorporate human-driven corrections or steering instructions during inference, particularly when the model is heading towards an unwanted behaviour. As such, my work lies in the overlap of interpretability (we need to understand how models reason for us to be able to provide inference-time feedback to them), representation learning (we want LLMs to learn representations that align with notions we can reason about), and human-AI interaction (we want an LLM’s representations and reasoning to be something humans can manipulate).

Selected Publications

Please see my Google Scholar profile for an up-to-date list of publications (https://scholar.google.com/citations?user=4ikoEiMAAAAJ&hl=en) 

Espinosa Zarlenga, M. ‘In Defense of Information Leakage in Concept-based Models’. International Conference on Machine Learning. 2026.

Luo, H., Espinosa Zarlenga, M., Jamnik, M. ‘In Defense of Information Leakage in Concept-based Models’. To appear at Advances in neural information processing systems 39. 2026.

Espinosa Zarlenga, M., Dominici, G., Barbiero, P., Shams, Z., and Jamnik, M. ‘Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts.’ International Conference on Machine Learning. PMLR, 2025.

Espinosa Zarlenga, M., Collins, K.M., Dvijotham, K., Weller, A., Shams, Z. and Jamnik, M. ‘Learning to receive help: Intervention-aware concept embedding models.’ Advances in Neural Information Processing Systems 36 (2023): 37849-37875.

Espinosa Zarlenga, M., Barbiero, P., Ciravegna, G., Marra, G., Giannini, F., et al. ‘Concept embedding models: Beyond the accuracy-explainability trade-off.’ Advances in neural information processing systems 35 (2022): 21400-21413.

Subjects
Mateo Espinosa Zarlenga
mateo.espinosazarlenga@trinity.ox.ac.uk