What “Everything Has Been Solved” Really Means in Matthew Berman’s AI Loop Argument
This is not an interview. It is Matthew Berman’s solo commentary on decompilation, software cloning, math, science, recursive self-improvement, and the human role left over when outputs become cheap. The episode’s strongest claim is not that every answer already exists, but that any domain with a clear verification loop becomes newly vulnerable to AI-driven search, replication, and acceleration.
1. Host and Subject Background
There is no verified guest context in the evidence. The episode is a solo commentary by Matthew Berman on Matthew Berman / Forward Future, titled “Everything has been solved,” uploaded on 20261010 and running 1057 seconds. The relevant background is therefore not a guest biography, but the host-and-subject setup: Berman is analyzing what happens when AI coding systems are placed inside loops where candidate outputs can be checked against a clear target.
He opens with a set of concrete hooks: characters from one game appearing in another game’s world, skate mechanics inside Call of Duty, Spiderman inside a Batman game, an open-source Photoshop clone, and his own attempt to recreate Super Mario World for the Mod Retro. Those examples frame the episode as an argument about verifiability. Berman’s subject is not merely “AI writes code,” but “AI can keep trying when success can be tested.”
2. What the Episode Covers
Berman begins by explaining decompilation. Source code is the human-readable and editable form written by programmers; a compiler turns that source into machine code that computers run efficiently. A game running on a console is not the original source code, but machine code. In his account, a human can technically inspect it, but doing so can require long cycles of guessing what a fragment does and then verifying whether that guess is correct.
The important distinction is that decompilation usually does not recover the exact original program text. Comments are often discarded, names can disappear, and compilers can rearrange or simplify structures. Berman’s simple example is a source method that returns 2 plus 3: after compilation, the machine behavior may only need to return 5. So the practical goal is not line-for-line recovery, but behavioral equivalence. The reconstructed game should look, play, and behave like the original, even if the code is different.
From there, the episode expands into a broader thesis. Decompilation is useful to Berman because it has a tight loop: write candidate code, compare it with the source behavior, and iterate. He applies that same pattern to Snowboard Kids, his own Super Mario World recreation, game mashups, Photoshop and Excel-style software cloning, math proofs, scientific discovery, and finally recursive self-improvement. The episode is best read as an analysis of one repeated structure: when the destination is checkable, AI can generate, test, compare, and revise at a scale humans previously could not sustain.
3. Core Views: Reasoning, Examples, and Limits
The strongest part of Berman’s argument is that he treats AI coding less as a single act of brilliance and more as an industrialized feedback loop. Decompilation fits that pattern because the desired behavior is known even when the original code is not. If the target is “make Mario jump in the same way,” the system can propose code, run it, compare the result, and try again. That framing reduces the need to claim that AI has recovered the author’s original intent or exact implementation. What matters is whether the rebuilt artifact behaves equivalently. His comparisons between years-long human decompilation efforts, the Snowboard Kids 84-day claim, and his own multi-day Super Mario World experiment all support one reasoning chain: small verifiable tasks can be distributed, checked, and iterated faster when AI agents do the repetitive search.
The same evidence also places limits on the claim. Berman says many games have recently been decompiled in weeks because of AI, but the provided evidence gives that as his episode statement, not an independently verified survey. The Chris Lewis, Snowboard Kids, GPT 6.1 Sol, Dev Day, and Mod Retro details are likewise episode-grounded claims. Even Berman’s own Super Mario World example comes with a caveat: he says it is not necessarily decompilation, but close to it. It is more accurately a from-scratch recreation based on target appearance, online documentation, code generation, checking, and iteration. That distinction matters because decompiling, cloning behavior, and rebuilding a system are related but not identical engineering acts.
His taxonomy of game mashups is one of the most useful parts of the episode because it prevents a vague “AI can mix games now” conclusion. A simple swap changes what appears on screen: Batman may look like Spiderman while still behaving like Batman. A pass-through mod keeps two games running and connects them, as in the Minecraft-in-GTA example where the Minecraft engine handles blocks and creatures while a connecting layer aligns cameras, combines visuals, and passes events such as TNT explosions into GTA. Mechanic transfer is harder because it moves behavior: acceleration, jump height, landing, collision, and wall interaction must work inside the new engine. The most intensive route is rebuilding a game and its rules by writing replacement software. Berman sees that last category as the most AI-unlocked, but also the most time-intensive.
The leap from games to general software is the episode’s central commercial provocation. Berman asks why AI could not observe and recreate any software if it can observe and recreate a game without access to the code. His Photoshop example is Photocraft, described as a free open-source clone produced after pointing AI at a local Photoshop installation. His Excel example is personal: he says he had AI click around, observe what actions did, identify behavior such as formulas in cells, and rebuild the output without seeing source code. The reasoning is behavior-first replication. The risk is not that every software company instantly loses all value; Berman explicitly says there can still be value in companies built around software. The pressure falls on businesses whose only moat is the software artifact itself, because training, enterprise features, customer support, maintenance, updates, trust, and distribution may become more important than code scarcity.
His extension to mathematics and science is more ambitious and more uncertain. Berman argues that math resembles decompilation because the end goal can be known while the path is unknown. AI can keep trying possible routes without tiring. He cites OpenAI publishing many mathematical proofs and research, and Will Depu using GPT-6 Pro and Fable 5.1 to classify recent math discoveries by whether they were made by humans, by AI before “yesterday,” or by AI “yesterday.” From that, Berman says the days of humans solving math alone are gone, then extrapolates to science, material discovery, and cancer treatment. The evidence supports that as his commentary, not as an independently established universal fact. “Anything with a verifiable outcome is basically solvable by AI” is the episode’s governing thesis, but it remains a framework whose strength depends on the cost, clarity, safety, and reliability of verification.
Recursive self-improvement is where the loop argument turns from productivity story into risk story. Berman calls it the ultimate loop: instead of improving a game or cloning software, AI improves its own efficiency, speed, or methods, applies the improvement to itself, and repeats. He says this could create an explosion of intelligence and connects it to frontier labs discussing pacing or slowing AI progress. He also says the implications are not fully known and that models need to remain aligned. His Alpha Evolve example is presented as a benefit, with the claim that Google found architectural improvements saving billions of dollars per year, but that remains a speaker-cited claim within the episode evidence. The larger point is balanced: the same loop that makes software recreation cheap could also accelerate model capability, raising questions about alignment, governance, and social timing.
4. Learning and Application
The practical lesson is to separate output generation from output verification. AI loops are most powerful when a candidate answer can be judged against a target. For software teams, that means converting vague goals into acceptance tests, screenshot comparisons, API contracts, performance thresholds, regression suites, data checks, or behavioral probes. The clearer the check, the more useful the AI loop becomes. The more subjective or underspecified the target, the more likely the loop is to produce confident but ungrounded churn.
A second application is task classification. Berman’s four mashup levels map well beyond games. A skin swap resembles a surface-level UI change. A pass-through mod resembles integration between two existing systems. Mechanic transfer resembles moving a business rule, interaction model, or workflow into a new environment. Full rule reconstruction resembles rebuilding a product or subsystem from observable behavior. These should not be estimated as one category. AI may accelerate all four, but the amount of verification, domain knowledge, edge-case testing, and human review increases sharply as the task moves from cosmetic change to behavioral reconstruction.
For product strategy, the episode is a prompt to audit the moat. If a product is valuable mainly because its visible behavior is hard to reproduce, AI-driven observation and cloning may reduce that advantage. If the value sits in proprietary data, regulatory posture, enterprise procurement, integrations, support quality, operational reliability, customer trust, brand, distribution, or continuing maintenance, then code replication does not equal business replication. Berman’s “software is solved” should be treated as a warning about shallow defensibility, not as a blanket verdict that software companies no longer matter.
The math and science extrapolation is useful, but only under conditions. A theorem proof, a unit test, a simulation result, a materials property, and a lab outcome are all “verifiable” in different ways. Some checks are cheap and deterministic; others are slow, noisy, expensive, regulated, or ethically constrained. Applying Berman’s loop outside software requires more than telling an AI to keep trying. It requires a trusted measurement process, a safe search boundary, clear failure costs, data provenance, stopping rules, and human accountability for what happens when the system optimizes toward the check but misses the larger context.
For creators, the most durable lesson is almost the opposite of automation hype. If AI makes variants cheap, judgment becomes scarcer. Berman argues that emotional resonance, cultural signal, fun, and meaning are not directly verifiable in the same way code or math can be. A person can generate many beautiful images or many game variants, but that does not answer which one will matter to people. Human taste, audience understanding, story judgment, and the courage to choose become more valuable in a noisy environment. The boundary is not that humans alone can create, but that humans may still be needed to decide what is worth creating, releasing, defending, and refining.
Source
- Original episode: Everything has been solved
More from WayDigital
Continue through other published articles from the same publisher.
Comments
0 public responses
All visitors can read comments. Sign in to join the discussion.
Log in to comment