Research

Five areas with one thread: working out when and why AI fails, with maths that predicts it and tools that check it. Each area page says what we’ve found and where it stops working.

  1. Hallucination and verification

    Why do models get facts wrong when the evidence is in front of them, and how can we catch it?

    In a pre-specified audit of 528 held-out questions, our answer-or-abstain gate kept hallucinations to 0.0–0.7% while abstaining on 20.6–27.9% of them (Chlon et al., ICML 2026).

    1 paper · Berry · 14 fellowship projects

  2. In-context learning

    What is a model actually computing when it learns from examples in a prompt?

    Across 92,160 positional edits on held-out prompts, our exact formula for attention predicted the direction of the change 95.4–96.5% of the time (Huang et al., 2026).

    2 papers · 6 fellowship projects

  3. Changing models without retraining

    Can a running model be adapted without retraining it?

    On Qwen2.5-7B, a 50,000-parameter controller came within a point of LoRA on GSM8K (71% against 72%), and lost 4.7 points on held-out code where LoRA lost 16.3 (ntkmirror).

    1 paper · ntkmirror · 7 fellowship projects

  4. Faster, cheaper AI

    How much of a large model’s memory can be dropped within a set error budget?

    Our add-on makes Triton, a widely used tool for writing fast GPU programs, work on NVIDIA’s GB10 (Blackwell) desktop hardware until official support arrives (triton-blackwell).

    triton-blackwell · 4 fellowship projects

  5. World models and science

    When do models trained on data learn the world’s symmetries, and when do they break them?

    Write the same robot trajectory as absolute joint targets instead of changes, and a world model’s retrieval degrades 2.6 to 13.4 times across three robot datasets (Karim and Chlon, 2026).

    1 paper · Mezzanine · 4 fellowship projects

Open problems

Questions we think matter and haven’t answered. Several have a fellowship project attached.

  • Will AI agents actually check their work?

    Our verifier catches unsupported claims well on benchmarks. In one coding-agent trial, the agent never used it unless it was required to. How should checking be built into agent workflows so it happens reliably, without slowing everything down?

  • When does more reasoning make answers worse?

    Longer chains of thought can increase competition between candidate answers at the final step. Can we predict, for a given problem, the reasoning length that gives the most reliable answer?

  • Prompt injection attacks come in families

    Real attacks vary in how they’re wrapped, where they appear and how they’re encoded. How can we give statistical guarantees about an AI system’s failure rate across a whole family of attack variants?

  • Independent checks for AI-written code

    A model can’t reliably check code it wrote itself, because it shares the blind spots that caused the bug. What’s the cheapest independent check that still catches most errors: running the code, a different model, or formal methods?

  • Which biological signals does a model really use?

    By hiding parts of a tumour’s multi-omics profile and asking a model to reconstruct them, we can see which relationships between data types it relies on. Which of those are reliable, and do they point to real biology?

  • Which facts does a model need to store?

    Early results suggest an AI system’s store of facts can be trimmed a lot for questions it can reliably work out again, but recall of other facts suffers. Where is the right line, and can we certify it?