When AI Starts Researching AI: Beyond Faster Experimentation
Liu Ziming’s argument is not simply to make AI run more experiments, but to turn observation, hypotheses, interventions, and outcomes into a learnable research process.
When AI Starts Researching AI: Beyond Faster Experimentation

Model companies add GPUs to training clusters. Startups add agents to code repositories. Both moves increase speed. But Liu Ziming, speaking on Zhang Xiaojun Podcast, asks a question one step earlier: if AI can only run more trials, is it doing research—or simply spending the experimental budget faster?
That distinction matters. Much of today’s AI-for-AI story treats research as an execution problem: write code, tune parameters, run experiments, inspect results, and repeat. That can be useful, but it resembles an indefatigable lab assistant. Liu is after a different capability: forming better judgments before an experiment starts—what is worth trying, and why it might work. The shift is from a more diligent AI researcher to a more intelligent one.
Model design is not yet a mature science
Liu’s path moved from AI for physics to the physics of AI. The first applies machine learning to scientific problems. The second borrows physics’ appetite for compression, explanation, and unification to understand why models behave as they do—and to use that understanding to design better ones.
His analogy is astronomy. First came Tycho Brahe’s observations, then Kepler’s empirical regularities, and only later Newton’s unifying theory. AI may not have reached its Newtonian moment. It may not even have surveyed the whole sky: the field may have converged too early around Transformers. Language is already an extraordinarily compressed artifact of human civilization; models can become powerful by consuming it. Vision, robotics, and world models deal with data that has been far less pre-abstracted. In those domains, architecture and the ability to form abstractions may matter much more.
That is why “the next Transformer” should not be treated only as an inspired guess. If the relationships among architectures, training dynamics, and data conditions can be recorded, compared, and predicted, model design can become accumulated knowledge rather than an artisanal practice.
Turn the research process into training data
Coding agents have advanced quickly not merely because models can write code. Environments such as GitHub supply structured artifacts, visible changes, and comparatively clear feedback. Research has no equivalent default corpus. Papers usually preserve the final answer, while the observations, abandoned paths, failed attempts, and revised hypotheses that shaped the next step disappear.
Liu describes an OPHIS framework for structuring research: Observation, Problem, Hypothesis, Intervention, and Speedup. The point is not to force science into a rigid form. It is to retain a reusable causal chain: what was observed, what problem it implied, what hypothesis followed, which intervention was made, and whether a measurable improvement resulted.
This needs a careful boundary. OPHIS is a framework described in the conversation, not an industry standard with broad validation behind it. Its value is that it pulls “research taste” and “intuition” back toward an inspectable process. Researchers may not be able to explain every good intuition at once. But repeated recording of decisions and experimental loops can begin to create material a system can learn from.
A model of models
The interview’s most concrete technical proposal is a “meta-model”: a model about models. Given an architecture, dataset, optimizer, and related conditions, it would predict a training curve or experimental outcome. It would not replace real training. It would rank many candidate ideas before expensive training, leaving physical experiments to validate the most promising few.
Liu connects the idea to a personal exercise: for dozens of days, he randomly selected datasets and models, predicted training curves before each experiment, and then used the outcome to correct his judgment. He found that the prediction skill could begin to transfer across modalities. The conjecture is that experienced researchers hold an unexternalized world model of AI training. A meta-model would try to externalize and scale that predictive capability through many pairs of models and training outcomes.
The constraint is equally clear. The data does not simply exist online. It must be created: train many different small models, preserve their dynamics, and deliberately cover unusual architectures and conditions. The goal is not necessarily cheaper than all training. It is to reduce the cost of blind exploration in model R&D.
Why new labs are appearing now
The discussion of neo-lab fundraising offers a better explanation than “capital chases hype.” Some problems fit neither the slow cadence of a conventional university lab nor a product roadmap that is already clear. World models, automated research, and mechanistic interpretability occupy that early territory. They need research freedom, but they also need an early prototype capable of becoming a product.
Liu describes a stage transition: explore and converge as a lab for roughly six to twelve months; once a route has a shape, operate more like a company, with products, resources, and commercialization. It is not a universal startup template. It does expose the central tension in research-led companies: productizing too early can lock a team into an easy-to-sell but unimportant feature; researching indefinitely can exhaust both the organization and its funding.
Interpretability is more than naming neurons
Mechanistic interpretability runs through the interview, but Liu does not equate it with telling a neat story about every neuron. Explanations at the individual-neuron level may not be robust across random seeds or model instances. A more useful question is which training conditions produce which behaviors, which tricks work in which regimes, and which structures retain their capabilities as conditions shift.
That is a more physics-like ambition. Explanation need not end with microscopic naming. It can also be a predictive map of conditions: not merely “this trick works,” but “where it works, what constrains it, and why it is worth testing again.”
The product is not a paper generator
The conversation ends with a more distant product vision: a training autopilot, or “vibe training.” A user states an objective and a budget; the system chooses an approach, designs a model, trains it, deploys it, and returns a usable result. The wager is that model training will not remain the exclusive domain of a few large companies. More teams, industries, and perhaps individuals could eventually obtain private models tuned to concrete problems.
No one can responsibly announce that future in advance. But the interview clarifies an overlooked fork. Automated research is not only the path of making agents run more. Another path makes research itself observable, comparable, and accumulable. One expands execution bandwidth. The other tries to improve the quality of judgment. The consequential moment may come when those two paths finally connect.
Original episode and notes
- Original episode: Zhang Xiaojun Podcast #149 — A firsthand look at the China–US neo-lab capital rush, AI for AI, mechanistic interpretability, and Max Tegmark
- This article is based on the episode transcript and summary. Research directions, timelines, and market views discussed in the episode are the guest’s perspectives, not independently verified conclusions.
More from WayDigital
Continue through other published articles from the same publisher.
Comments
0 public responses
All visitors can read comments. Sign in to join the discussion.
Log in to comment