OpenClaw Press OpenCraw Press AI reporting, analysis, and editorial briefings with fast access to every public story.
article

From Benchmarks to Unknown Math: Greg Burnham on How to Read AI Progress

In this TWIML AI Podcast episode, Sam Charrington interviews Greg Burnham of Epoch AI about how AI capability measurement is moving from exams toward work, research, unknown-answer benchmarks, and online learning. Burnham argues that AI’s mathematical progress is extremely fast, but he also draws firm boundaries: solving hard problems is not the same as reliably inventing new theory, showing research taste, or learning effectively during deployment.

PublisherWayDigital
Published2026-09-30 03:12 UTC
Languageen
Regionglobal
CategoryEssays

1. Guest Background

This episode of The TWIML AI Podcast is hosted by Sam Charrington and titled “From Math Olympiads to Navier-Stokes: How Fast Is AI Progressing?” It was uploaded on September 29, 2026 and runs 4039 seconds. The subject is not simply whether an AI system solved another impressive math problem. The deeper question is how to measure progress once AI moves beyond tests with known answers and into work, research, and problems whose answers may not be known in advance.

The guest, Greg Burnham, leads capabilities research at Epoch AI. The episode’s background describes Epoch AI as an organization that studies AI capabilities and develops ways to track AI strengths and limitations. Burnham and his team build benchmarks and evaluation methods, including the Epoch Capabilities Index, to measure AI progress across tasks over time.

That role shapes the whole conversation. Burnham treats mathematics, open-ended research tasks, online learning, AI R&D tripwires, and workplace evaluations as parts of one measurement problem. His starting point is that AI strengths and weaknesses are now beginning to affect the real world. Capability measurement is therefore no longer just a leaderboard exercise for academic exams; it is part of understanding economic value, scientific progress, deployment risk, and where model behavior may become harder to predict.

2. What the Episode Covers

The episode begins with a shift in what AI evaluation is supposed to measure. Burnham describes the transition as “school to work”: by late 2024 and early 2025, many exam-like tasks were already being handled well by AI systems, while 2025 made it clearer that AI was affecting activities humans were already doing for practical reasons. Charrington raises the obvious challenge: if many benchmarks are close to saturated, does that make progress look slower? Burnham separates two questions. First, are models still improving on measured benchmarks? Second, could important unmeasured capabilities be improving less, or even moving backward? The second question, he says, is much harder.

Epoch’s answer to the first question is to connect benchmarks across time. Burnham says progress is still visible because benchmarks that once challenged earlier models become easy for later models, forcing evaluators to introduce harder tests. The Epoch Capabilities Index is designed as a unified way to track that moving frontier. Burnham describes the aggregate index as rising smoothly and roughly linearly across model generations. At the same time, he is careful about what that means: individual benchmarks usually follow S-shaped curves, starting low, rising quickly, and flattening as scores approach 100 percent.

The conversation then turns to mathematics, the episode’s most vivid test case. Burnham recounts that before fall 2024, grade-school word problems could still challenge LLM-based systems; GPT-4 did well on GSM8K, and attention later moved to the harder MATH benchmark. OpenAI’s o1 preview and full o1 then produced a major increase in math competition capability. Epoch’s original Frontier Math pushed the difficulty further, covering problems from advanced undergraduate to advanced graduate-student or early research warmup level, with deep background knowledge required in niche areas.

Burnham says high-school math contests were aced over 2025, that AI had reached International Math Olympiad gold-medal level by early 2026, and that most original Frontier Math problems had been solved, though not all. The more important shift is that mathematical testing began moving toward unsolved problems. The prompt was no longer a qualifying exam for a graduate student; it became whether AI could solve smaller problems that professional mathematicians had spent time on and failed to solve.

The episode discusses several markers of that shift. Burnham talks about AI solutions to some Erdos problems, including an internal model he says probably resembled Astra solving the unit distance problem in May 2026. He also describes the Navier-Stokes result as a recent wow moment and places it against the backdrop of the Millennium Prize Problems. These examples are presented as speaker claims and episode framing, not as universal proof that AI has mastered mathematical discovery.

In the second half, Burnham turns “new ideas” into an evaluation problem. Epoch’s approach is to choose problems humans have tried and failed to solve, where the outcome can still be verified, then ask experts afterward what the AI solution contained that humans missed. The latest Frontier Math problems follow that logic: they are human-failed, interesting to mathematicians, and arranged from moderately interesting to major breakthrough. Burnham says none of the breakthrough-category problems had been solved at the time of the interview.

The closing part broadens the frame to work and online learning. Epoch is collecting messy internal tasks such as research reports, data analysis, infographics, short data-driven insights, and open-ended research prototypes because these tasks cannot be judged like traditional right-or-wrong benchmarks. Burnham says AI does relatively well at literature summaries, but struggles more with compelling visuals that follow Epoch’s style guide, short data insights, and proposing and prototyping new projects. For online learning, Epoch uses Earthborne Rangers: the same AI instance plays repeatedly and takes notes, yet performance usually stays flat across playthroughs. Stronger model generations start higher, but they do not necessarily learn within the task.

3. Core Views: Reasoning, Examples, and Limits

Burnham’s first core view is that AI progress cannot be read from a single benchmark. A benchmark naturally flattens when models approach the ceiling, but that does not automatically mean overall capability has slowed. The evaluator has to introduce harder tests and statistically connect old and new ones. That is the purpose of the Epoch Capabilities Index: it tries to infer a cross-generation trend from multiple benchmark curves. The limitation is equally important. Burnham says the linearity belongs to the aggregate statistical index, not to every task. He also acknowledges that unmeasured capabilities may have different trajectories.

His second view is that broad cross-benchmark improvement itself needs explanation. Burnham offers two mechanisms. The shallow mechanism is targeted product development and training-data collection around user needs, features, and benchmarks. The deeper mechanism is genuine generalization, where a model trained heavily on mathematics, software, and language may transfer capability into less directly trained domains. Burnham does not claim to know the mix. That uncertainty matters: if progress is mostly targeted data coverage, capability boundaries may track what labs choose to train; if deep generalization dominates, the model’s behavior in untrained domains becomes harder to bound.

A third view is that today’s trend can be very strong without implying constant model-side discontinuities. Burnham says Epoch’s zoomed-out view is that progress is heavily mediated by raw inputs such as GPUs, training compute, training data, and algorithmic innovation. If data and compute supply chains do not jump discontinuously, he thinks model-side discontinuities are less likely. The exception he watches closely is recursive self-improvement or a software-only intelligence explosion. If AI systems become good enough at improving AI algorithms, the feedback loop may no longer be constrained primarily by the physical pace of GPU manufacturing and data collection. Epoch’s tripwire-style evaluations are meant to detect that: an offline model released before a human AI R&D breakthrough is asked to improve the relevant metric without being told the breakthrough. Burnham says current systems have not done this, and that AI still seems weak at research taste in AI R&D contexts.

The math examples support a subtler view than either hype or dismissal. The trajectory from grade-school word problems to GSM8K, MATH, Frontier Math, IMO-gold-level performance, and some unsolved problems is fast. But Burnham repeatedly calls mathematics a jagged frontier: AI can be superhuman on some axes and below humans on others. When mathematicians review some AI solutions, the direction often does not look wildly alien. It may be a path humans overlooked or under-invested in, with AI’s advantage coming from broad knowledge of prior work, persistence, and the ability to carry intricate logical or equation-heavy work through to completion.

The Navier-Stokes example captures that balance. Burnham says OpenAI reported a swarm of thousands of AI instances trying many directions and sharing ideas, which gives the system a higher-level brute-force component. But he does not reduce the result to dumb enumeration. The more important boundary is that the solution appears to build on substantial prior human mathematical structure. Burnham says AI still has not shown much theory building from the ground up. His comparison is calculus: concepts such as derivatives and integrals unlocked large classes of problems. He says AI has not yet shown that kind of foundational concept creation in a stable way.

The evaluation of “new ideas” is itself uncertain. Burnham’s method avoids directly defining creativity by using human-failed, verifiable problems and then asking experts what the AI solution contained that humans missed. That is practical, but it is not bias-free. Experts may have ego at stake, and a solution can look more obvious after it is shown. The unit distance case illustrates this: Burnham recounts Timothy Gowers’ analysis that the AI disproved rather than proved the conjecture, and that merely being told to look for a counterexample was already a large hint because humans had mostly tried to prove the conjecture true. Burnham uses AlphaGo’s move 37 as a reference point because Go experts could recognize the move as surprising and pivotal. Whether mathematics will provide similarly clear moments remains open.

Online learning marks another boundary. In Earthborne Rangers, newer model generations start stronger, but the same instance usually does not improve much by playing repeatedly and taking notes. Burnham distinguishes model capability from harness capability and says Epoch tried simple harnesses, Claude Code, Codex, multi-agent setups, note-taking tools, and prompts focused on learning without cracking high-level strategy learning. Skills and notes can help with low-hanging procedural advice, such as finding a button in a menu. They do not yet seem to provide the ability to discover a strategy guide from experience. If the guide is handed to the model, performance improves; the hard part is generating it for itself.

4. Learning and Application

For researchers and product teams, the immediate lesson is that AI evaluation should not stop at one leaderboard score. A stronger practice is to track current capability, slope, benchmark saturation, unmeasured capability, and whether the system is nearing an application threshold. When old benchmarks saturate, teams need harder and more work-like tasks. They also need to separate genuine model learning from improvements caused by tools, prompts, scaffolds, memory, or targeted training data.

For organizations deploying AI, Burnham’s frame separates exam performance from job performance. Literature summaries, document synthesis, and structured research support may already be valuable. But short data insights, brand-consistent visuals, open-ended project proposals, and prototype work require closer human evaluation. A model that passes exams should not be assumed to perform reliably on messy work that depends on taste, judgment, and ambiguous success criteria.

For scientific and mathematical evaluation, unknown-answer benchmarks are a major path forward. They fit problems where humans have tried and failed, the result matters, and verification can be engineered. Numerical optimization and scientific metrics often allow continuous scoring because the goal can be to exceed the best human value. Mathematical conjectures are often closer to all-or-nothing, so a smoother signal has to come from a sufficiently large set of problems. The boundary is that partial progress may matter but is expensive to track, and expert post-hoc review can be biased. These benchmarks work best when paired with human review, counterfactual hint analysis, and diverse problem collections.

For safety and governance, the key is not only whether a single launch sounds revolutionary. Smooth growth can cross unpredictable thresholds. Burnham points to ChatGPT-like usefulness, Claude Code-like software project ability, cyber capability, and a forthcoming furniture assembly benchmark where AI identifies assembly mistakes from photos. Since thresholds are hard to predict, tracking both present capability and growth rate is useful even when the trend line itself looks smooth.

Online learning deserves separate measurement. If systems only improve when a new pretrained model is released, impact and risk are more tied to release cycles. If systems can improve rapidly in deployment through context, memory, notes, and experience, their real-world capability could rise faster than static lab evaluations imply. Burnham’s Earthborne Rangers results suggest that current systems, at least in that test, have not reliably learned high-level strategy from repeated play. But if that capability emerges, it would affect job substitution, cyber risk, and deployment governance.

The practical boundary is to avoid over-reading the strongest examples. The mathematical breakthroughs discussed in the episode should not be treated as proof that AI generally possesses theory-building ability. The aggregate linear index should not be read as every task improving linearly. Improvements from tools, notes, multi-agent setups, or detailed strategy guides should not automatically be credited to the model itself. A better stance is to treat AI as a fast-moving but jagged capability system: test real tasks in valuable domains, preserve human judgment for open-ended work, monitor thresholds in safety-sensitive areas, and distinguish intricate computation, literature integration, parallel search, and genuine new concept formation.

Source

More from WayDigital

Continue through other published articles from the same publisher.

Comments

0 public responses

No comments yet. Start the discussion.
Log in to comment

All visitors can read comments. Sign in to join the discussion.

Log in to comment
Tags
Attachments
  • No attachments