Louis Yiven Zhu

Oxford Internet InstituteAI Evaluation

00Who am I

I’m Louis, an MSc candidate at the Oxford Internet Institute. My research asks whether the benchmark scores used to judge AI systems can be trusted, and follows those scores into the things they decide: prices, wages, and rules.

ice hockeyalpine ski racingpiano

01Research

Benchmarks are the exams AI models sit, and an exam is only useful if it is fair, consistent, and testing the right thing. Much of this research applies the statistics of human testing — psychometricsThe statistical field behind standardised testing: how to tell whether a test is reliable, comparable, and measuring what it claims. — to the exams we set for machines, and then follows the scores out into prices, wages and rules.

  1. Q1Do benchmarks measure one ability or many?
  2. Q2Does a score depend on how the test was run?
  3. Q3What do the scores change in the real world?
under review

One Capability or Many?

Twelve benchmarks that claim to measure different skills mostly track a single underlying factor.

421 modelsHash-pinned frontier model configurations in the frozen twelve-benchmark battery. 74.5% one factorShare of common variance on the first factor, exploratory factor analysis. As reported in the manuscript.

pre-registered · with Marcos Barreto, LSE Statistics · aimed at ICLR 2027

under review

The Price of Intelligence

A price index for AI that holds capability constant — the way statistical agencies price computers.

782 modelsModels in the frozen 31-month price panel, each observation with source URL and archive timestamp. 87% hidden declineShare of the measured decline in the price of capability that standard matched-model methods cannot see, because it arrives through new models rather than price cuts on old ones.

pre-registered on OSF · aimed at FAccT 2027

in preparation

The Science of Evaluations

Cross-lab standards for what a field must show before an evaluation counts as measurement. I write the validity and evidentiary-standards sections.

EvalEval Coalition — Hugging Face, Edinburgh, EleutherAI · aimed at TMLR

published

From Advisor to Voting Teammate

When an AI joins a group decision, its institutional role matters more than its accuracy.

1.1M simulationsRuns of the agent-based model across authority structures and information conditions.

ACM CHI 2026, workshop on human–agent collaboration

to be submitted

The Unassembled Validity Argument

Six years of MMLU, the most-cited AI benchmark: the score depends on how the test is run, and the instability reaches the leaderboards built on it.

BSc dissertation · STS Best Dissertation Prize

working paper

Automation Risk and Wage Dynamics in the UK

Does being measured as automatable move your wage? UK risk scores linked to pay records, occupation by year.

SSRN · under review, UCL Journal of Economics

Also: adversarial CAPTCHAs that humans pass and AI agents fail, with UCL Computer Science and Holistic AI · a framework for when neural data should inform welfare policy arXiv

02Experience

University of Oxford UCL LSE MIT Hugging Face
2026–

Departmental Associate · UCL Science & Technology Studies

Designing a new module on AI, digital labour and the future of work, funded by UCL MAPS. Supervised by Dr Joanna Octavia

2026–

Summer Researcher · LSE Department of Statistics

AI measurement and the factor structure of frontier evaluation, with Prof Marcos Barreto

2026–

Researcher · EvalEval Coalition

Cross-lab evaluation standards with Hugging Face, Edinburgh and EleutherAI; writing the validity sections of the Science of Evaluations paper

2026

Teaching Assistant & Course Developer · UCL STS

Responsible Innovation in Practice and Governance of Emerging Technologies

2025–26

Student AI Researcher · Holistic AI

Adversarial robustness of LLM and vision–language agents

2025–26

Student Researcher · UCL Computer Science

Agent-based modelling of AI authority in group decisions, with Prof Maarten Speekenbrink — published at CHI 2026

2025

Research Contributor · Institute of Economic Affairs

UK graduate premium, contributing to Julian Jessop’s analysis of administrative earnings data

2025

Research Contributor · UCL Institute for Global Prosperity

AI & Youth, with Honorary Professor Noreena Hertz

03Talks & community

Talks & presentations

  • ACM CHI 2026, workshop on human–agent collaboration — paper presentation
  • UCL Centre for Responsible Innovation — invited talk on the MMLU validity work
  • Explore Econ 2025, UCL — mandatory AI risk disclosure as a regulatory mechanism

Roles

  • AI for Good, University of Oxford — Director of Engagement
  • Oxford Artificial Intelligence Society — Sponsorship Lead
  • Thinking About Thinking — Oxford Fellow
  • UCL Investment Society — Chairman, 2026/27

04Recognition

05Ask me anything

Ask Louis
ready

06Contact

For research, collaboration, or data and code requests: