Hi! I'm a postdoc at the University of Oxford, where I work on AI safety and interpretability. I'm interested in understanding how reliably we can monitor AI systems' behavior through models' internal representations, and in building better techniques for doing so. During my PhD, I worked on scalable methods to make models' computations easier for humans to interpret by design. I believe that interpretability may prove useful not only for mitigating the risks posed by AI, but also in positively steering models toward behaviors aligned with human values. I was a PhD student at Queen Mary University of London, during which i spent time as a visiting student at Oxford and visiting scholar at UW-Madison.
News
- [26.09] Serving as an Area Chair at ICLR'27
- [26.04] Recognized with a Gold Reviewer Award at ICML'26
- [25.10] Recognized as a Top Reviewer at NeurIPS'25
- [25.10] Started as a postdoc at the University of Oxford
- [25.05] Started as a visiting PhD student at the University of Oxford
- [25.04] Started as a Research Affiliate at AIGI Oxford
- [24.11] Recognized as a Top Reviewer at NeurIPS'24
- [24.09] Started at UW-Madison as an Honorary Associate
- [21.09] Started as a PhD student at QMUL
Selected publications
-
"Beyond Linear Probes: Dynamic Safety Monitoring for Language Models"
J. Oldfield, P. Torr, I. Patras, A. Bibi, F. Barez
ICLR, 2026
[pdf | code | project page] -
"Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders"
J. Oldfield, S. Im, S. Li, M. A. Nicolaou, I. Patras, G. G. Chrysos
NeurIPS, 2025
[pdf | code] -
"Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization"
J. Oldfield, M. Georgopoulos, G. G. Chrysos, C. Tzelepis, Y. Panagakis, M. A. Nicolaou, J. Deng, I. Patras
NeurIPS, 2024
[pdf | code | project page]
Experience
- [25.10] Postdoctoral Research Assistant (University of Oxford, Oxford)
- [25.05–25.09] Visiting Student (University of Oxford, Oxford)
- [25.04–25.09] Research Associate (AIGI Oxford, Oxford)
- [24.09–24.12] Honorary Associate (University of Wisconsin–Madison, Madison)
- [21.09–25.10] PhD Student (QMUL, London)
- [19.11–20.09] Research Intern (The Cyprus Institute, Nicosia)
Awards
- [26] Gold Reviewer Award: ICML 2026
- [25] Top Reviewer: NeurIPS 2025
- [24] Top Reviewer: NeurIPS 2024
Supervision
I'm very grateful to be co-supervising the following MSc students (jointly w/ Dr. Adel Bibi):- [25-26] Aniruddh Pramod (MSc CS)
- [25-26] Maryna Horbach (MSc CS)
- [25-26] Arshia Hemmat (MSc CS) jointly w/ Dr. Lin Li
Invited talks
- [24.06] Tensor Decompositions in Large Scale Deep Learning (Archimedes Research Unit, Athens)
Teaching
Teaching assistant on the following modules:
- [25–25] AI Safety and Alignment (Oxford, Michaelmas term)
- [24–24] Deep Learning and Computer Vision (QMUL, ECS795P)
- [21–22] Machine Learning (QMUL, ECS708)
- [21–21] Artificial Intelligence (QMUL, ECS629)