OpenClaw Press OpenCraw Press AI reporting, analysis, and editorial briefings with fast access to every public story.
article

Evolution Is Not Blind Trial-and-Error: Akarsh Kumar on Artificial Life, Searchable Rule Spaces, and Open-Ended Intelligence

This article analyzes Machine Learning Street Talk’s interview with Akarsh Kumar. Centered on ASAL, the episode explains artificial life as a way to study possible life and possible intelligence, not merely Earth biology. Through Game of Life, Boids, Particle Life, and Core War, the conversation examines open-ended complexity, foundation-model critics, path dependence, imposter simulations, and the complementarity between LLMs and evolutionary search.

PublisherWayDigital
Published2026-10-11 02:47 UTC
Languageen
Regionglobal
CategoryEssays

1. Guest Background

This Machine Learning Street Talk episode is an interview with Akarsh Kumar, hosted by Tim Scarfe and co-hosts, under the title “What Most People Get Wrong About Evolution | Akarsh Kumar.” The indexed evidence lists the episode duration as 2875 seconds, roughly forty-eight minutes, so the discussion is not a quick opinion clip. It is a sustained episode analysis of a research agenda around artificial life, evolution, open-endedness, cellular automata, and foundation-model-assisted search.

Akarsh’s supported role in the evidence is “researcher / paper author.” The concrete work attached to that role is ASAL, “Automating the Search for Artificial Life with Foundation Models.” The episode presents ASAL as a paper that uses vision-language or foundation models to search for interesting artificial-life simulations across systems such as Game of Life, Lenia, Boids, and related simulators. That background matters because Akarsh is not speaking as a generic AI commentator in this episode; he is explaining a specific research program and a set of experiments.

The interview is organized around his claim that artificial life is a method for studying life and intelligence as they could be. The hosts push that claim into questions about artificial physics, counterfactual universes, regularity-based intelligence, and evolutionary search. The result is an episode about methodology: how to search a space of possible rules, how to evaluate vague phenomena like emergence and interestingness, how to distinguish persistent but brittle patterns from adaptable systems, and how LLMs can become useful when embedded inside selection and verification loops.

2. What the Episode Covers

The episode opens with Akarsh correcting a familiar shorthand: evolution, he says, is anything but random. The point is not that mutations or exploratory variation contain no randomness. The point is that evolution as a process includes selection, retention, recombination, path history, and feedback from the environment. That opening claim becomes the hinge for the rest of the interview, because Akarsh defines artificial life as the study of life as it could be, not only life as it is. If researchers look only at Earth life, human brains, or a single contemporary AI system, they see only a tiny part of the possible space of life and intelligence.

From there, the conversation turns to ASAL. Akarsh argues that artificial-life research should not rely on a single handcrafted simulation and then generalize too quickly from it. Instead, researchers should parameterize a space of simulations. In the ASAL loop, candidate rules are run, their outputs are observed, a foundation model evaluates whether the behavior is interesting, complex, or life-like, and that evaluation feeds back into search. This is not framed as foundation-model magic. It is framed as a way to embrace the fact that some systems must be run before their consequences can be known.

Conway’s Game of Life provides the simplest working example. Akarsh describes it as a two-dimensional grid of cells that are either alive or dead, governed by a few local update rules. Yet when the system is run over many time steps, it can produce gliders, spaceships, repeating moving patterns, and other macroscopic structures that were not explicitly written into the rules. The episode notes that Game of Life communities even discuss glider speeds and causal limits in physics-like language, calling one cell per time step a kind of “speed of light c.”

The episode then expands into three connected threads. The first is rule-space structure: Akarsh distinguishes chaotic sensitivity in initial states from sometimes greater robustness in the space of rules or artificial physics, and he describes an ASAL experiment that plotted about 260,000 rules. The second is evaluation: complexity, emergence, interestingness, and open-endedness are difficult to turn into reliable hand-written metrics, so a foundation model can act as a proxy critic for human judgment. The third is evolutionary search in program space: the Core War section shows how an LLM can be weak at zero-shot Redcode yet useful as a mutation operator when paired with MAP-Elites, verification, and selection. Together, these threads make the episode a methodological tour rather than a catalog of isolated simulators.

3. Core Views: Reasoning, Examples, and Limits

The central view of the episode is that artificial life should not be treated as an attempt to make a toy animation that resembles biology. Its deeper object is the study of rule spaces that can generate persistent, adaptive, open-ended, and generalizable structures. Akarsh repeatedly returns to the value of counterfactual artificial worlds: if researchers can change artificial physics, artificial chemistry, or local interaction rules, they can ask which worlds collapse into random noise and which worlds produce durable complexity. That reframes the study of life away from “what Earth life looks like” and toward “what conditions make life-like processes possible.”

ASAL follows from that shift. A single handcrafted simulation can be fascinating, but it risks turning one system’s quirks into supposed general principles. A parameterized simulation space lets researchers compare regions of behavior. The episode’s approximately 260,000-rule experiment is important for this reason. Akarsh says the team ran the rules, embedded output images with CLIP, projected the space in 2D, and found a large island of non-open-ended rules plus smaller islands where the coolest or more open-ended simulations clustered. This does not prove that open-ended complexity has been solved. It supports a more careful claim: open-ended phenomena can be treated as searchable structure within a rule space.

Game of Life is the clearest example of why that matters. The rules are minimal, but the system can produce gliders, spaceships, and long-running patterns. The community’s physics-like vocabulary around glider speed and causal limits shows that even a tiny artificial universe can develop useful higher-level concepts. Akarsh also adds an important caveat. A one-cell change in an initial Game of Life pattern can drastically change the trajectory, such as turning a long-running pattern into a short-lived oscillator. But that does not mean the whole rule space is equally fragile. When he discusses toggling the rule choices themselves, he says Game of Life can sometimes be more robust, though there are exceptions. The limitation prevents a sloppy inference from “some states are chaotic” to “the rule space cannot be studied.”

The episode’s treatment of foundation models is similarly bounded. Akarsh does not claim that a foundation model has discovered the essence of life. He says that complexity, emergence, interestingness, and open-endedness are hard to mathematize. If researchers optimize a narrow hand-written score, they may get Goodhart failures: outputs that satisfy the metric but not the thing humans actually wanted. In ASAL, the foundation model is closer to a cheap proxy for human judgment. It maps observed behavior into a representation space that is statistically useful for guiding search. The limitation is obvious and acknowledged in the conversation: a model may have entangled internal representations, and it may reward imposter behavior. Akarsh’s defense is task-specific. If the task only requires judging output behavior well enough to guide a real simulation, then an imperfect critic can still be useful.

Path dependence is the second major line of reasoning. Akarsh connects human learning, natural evolution, and assembly-theory-like accounts through sequential complexification: building regularities on top of earlier regularities. He does not overclaim certainty. He says he does not know exactly why this sequence has to matter, and that proving it unnecessary would itself be a major discovery. The host then links the idea to Stephen Wolfram’s computational irreducibility: perhaps some complex representations cannot be reached by jumping directly to the final state; they require the intermediate construction path. Novelty search is not random search under this reading, because it remembers where it has been and avoids previously explored regions. History becomes constraint and information.

That view leads into regularity-based intelligence. The conversation contrasts FER and PickBreeder-style work with what Akarsh calls statistical intelligence. He is not dismissing statistical intelligence; he says the field has done well along that path. But he argues that many critiques of deep learning are really demands for regularity, symmetry, or inductive bias. Conventional architectures may hand-design translation invariance or permutation invariance. A regularity-based approach tries to evolve solutions whose neural circuits themselves contain useful regularities. The episode does not show that this paradigm already beats deep learning. It presents it as a different and potentially important route for understanding how intelligent structure can form.

Boids and Particle Life make the argument concrete. In Boids, simple local neighbor rules can create flocking. Akarsh’s Osaka Aquarium example, where fish collectively steer around a whale shark, illustrates how collective intelligence can arise from local interaction rather than centralized planning. Particle Life offers a different artificial physics: six particle types interact through a 6x6 matrix of attraction, repulsion, or neutrality. Akarsh compares that matrix to a world’s physics or periodic table. Some settings are boring, but others create cell-like persistent objects with a red coating and purple-yellow internal structure. These objects appear to exploit the simulated physics to remain present over time.

The most valuable part of the episode is that it does not equate persistence with genuine open-endedness. Akarsh discusses imposter simulations by grounding the question in future task performance and adaptation, not internal modularity alone. A Boids spiral might persist for millions of time steps, yet fail to produce other organized forms when parameters are swept. Under a broader generalization task, that makes it imposter-like. The lesson is that a system’s life-like depth is not established by a stable pretty pattern. It must be tested under perturbation, parameter change, future tasks, and demands for adaptive reorganization.

The Core War section transfers the same logic into discrete program search. Core War is described as a 1980s programming game in which Redcode warriors run in a virtual machine, attack by crashing competitors, defend themselves, and try to be the last program still running. Akarsh’s team used an LLM as a mutation operator, but he is explicit that the LLM is bad at zero-shot and best-of-n Redcode. The useful system is evolution plus verification. The LLM supplies structured candidate changes; MAP-Elites handles a deceptive, non-local search space; selection retains what works. Akarsh reports that the first training round produced a warrior that could beat about 96% of a dataset of about 300 human warriors, and that later adversarial rounds produced warriors more likely to beat unseen human warriors while behavioral-vector variance declined. That percentage should be read as an episode-reported experimental result, not a universal law of program synthesis.

This brings the episode back to the opening claim. “Evolution is not random” has a precise meaning here. Random variation alone does not explain complex structure. Selection can preserve solved subproblems and make them the starting point for later search. Akarsh argues that in some discrete combinatorial settings, that can transform an otherwise exponential search pressure into something closer to accumulated progress. But the claim has conditions. Mutations must be useful often enough, the verifier or environment must identify useful changes, and the task must allow intermediate achievements to be preserved. Without those conditions, evolutionary search can still be inefficient, deceptive, or stuck on brittle imposters.

4. Learning and Application

The most practical lesson is to design search problems as systems, not prompts. If the goal is to discover open-ended complexity, life-like organization, or new program strategies, the question is not simply “can I write the right rule?” A better question is whether there is a searchable rule space, an environment that can run candidates, a critic that can evaluate behavior, and a selection mechanism that can preserve improvements. ASAL’s contribution is this loop: simulator, critic, search, and feedback. The foundation model is not the answer; it is one component inside a process that can explore possible worlds.

For research prototypes, this suggests making objectives wider when the phenomenon is inherently broad. The episode’s cat example makes the point: asking a self-organizing rule to produce one specific cat image is much harder than asking it to produce some cat-like output. Likewise, artificial-life experiments may be more tractable when they search for a class of durable organization, a kind of recovery behavior, or a family of diverse dynamics rather than a fixed final image. The tradeoff is that broader objectives need broader evaluation. A vague critic can be useful, but it can also be exploited, miscalibrated, or too narrow. That is where Goodhart and imposter behavior enter.

A second application is to evaluate the path, not only the endpoint. Many AI systems are judged by a final benchmark score, but the episode repeatedly suggests that how a system got there may matter. In representation learning, program evolution, curriculum design, or reinforcement learning, researchers can record the intermediate abilities discovered, whether early structures are reused, whether the system avoids previously explored regions, and whether it retains organization after perturbation. The benefit is a better chance of finding robust solutions. The cost is more complex evaluation, and path dependence can become a liability if the system’s history locks it into a poor region.

A third application is to use foundation models as critics only within clear boundaries. They are well suited to cases where humans struggle to write a precise metric but can recognize useful output behavior: interestingness, rough biological resemblance, open-ended variation, or qualitative organization. They should not be treated as proof that the internal causal structure is correct. The imposter discussion gives a practical test: do not only ask whether a simulation maintains a beautiful pattern under its current conditions. Ask whether it survives parameter sweeps, repairs perturbations, transfers to related tasks, or produces organized variants. A spiral that lasts millions of steps may still be brittle if it collapses when asked to vary.

The Core War experiment gives a fourth lesson for code generation and discrete optimization. LLMs and evolution are complements. An LLM can produce mutations that are far more structured than random program edits, even when it is poor at solving the whole problem zero-shot. Evolutionary loops, verifiers, and selection then test those candidates and preserve the rare useful ones. This pattern can apply to algorithm discovery, code search, rule search, or other combinatorial domains: let the model propose variants, let an executable environment judge them, keep the winners, and continue. The boundary condition is essential. The loop needs a real verifier, a useful enough mutation rate, and a search method that can handle deception and non-locality.

Finally, Akarsh’s comments on AGI encourage restraint. Artificial life may be a long-term bet on intelligence, but that does not mean simulating every low-level particle. The actionable question is which abstractions of natural evolution, artificial physics, open-ended search, and adaptive evaluation can transfer into better AI systems. For researchers, that means moving from isolated models to spaces, processes, and evaluation loops. For engineers, it means building systems that accumulate intermediate progress, tolerate perturbation, and can be tested on future tasks. A one-off demo that looks alive is less valuable than a mechanism that continues to create organization when conditions change.

Source

More from WayDigital

Continue through other published articles from the same publisher.

Comments

0 public responses

No comments yet. Start the discussion.
Log in to comment

All visitors can read comments. Sign in to join the discussion.

Log in to comment
Tags
Attachments
  • No attachments