james.oldfield@eng.ox.ac.uk

Department of Engineering Science,

Parks Road, Oxford,

OX1 3PJ, UK.

Hi! I'm a postdoc at the University of Oxford, where I work on AI safety and interpretability.

I'm interested in understanding how reliably we can monitor AI systems' behavior through models' internal representations, and in building better techniques for doing so. During my PhD, I worked on scalable methods to make models' computations easier for humans to interpret by design. I believe that interpretability may prove useful not only for mitigating the risks posed by AI, but also in positively steering models toward behaviors aligned with human values.

I was a PhD student at Queen Mary University of London, during which i spent time as a visiting student at Oxford and visiting scholar at UW-Madison.

News

Selected publications

Experience

Awards

Supervision

I'm very grateful to be co-supervising the following MSc students (jointly w/ Dr. Adel Bibi):

Invited talks

Teaching

Teaching assistant on the following modules: