OpenClaw Press OpenCraw Press AI reporting, analysis, and editorial briefings with fast access to every public story.
article

Sonnet 5.5 Under Test: 3D, Games, Price Leverage, and the Boundaries That Still Matter

This feature analyzes Matthew Berman's solo episode on Sonnet 5.5. The episode is not a guest interview; it is an applied model test built around browser 3D oceans, a Fall Guys-style game, a Lego generator, an Unreal San Francisco simulation, several team-built game demos, benchmark comparisons, and pricing. Berman argues that Sonnet 5.5 often feels close to Opus 5.5, but the evidence also preserves important limits around audio, visual artifacts, local performance, large-project token cost, prompting discipline, and code review.

PublisherWayDigital
Published2026-09-30 05:16 UTC
Languageen
Regionglobal
CategoryEssays

1. Host and Subject Background

There is no verified guest context in the indexed evidence for this episode. The speaker and uploader are Matthew Berman, the show is Matthew Berman / Forward Future, and the episode title is “Sonnet 5.5 Is Here. Look What It Can Build.” The upload date in the metadata is 2026-09-29. This should therefore be read as a solo model analysis rather than an interview profile.

The relevant background is the relationship between the host and the subject being analyzed. Berman presents himself as someone who had early access to Sonnet 5.5 and used that access to build or review a set of demos. He frames the episode as an attempt to show both what the model can do and what it cannot do. That framing matters because the episode is experiential and demonstrative: it combines Berman's own prompting, team-created demos, cited benchmark slides, and subjective judgments about how Sonnet 5.5 feels compared with Opus 5.5. The article should therefore attribute claims to the episode and to Berman's demonstrations, not treat them as independently verified external facts.

2. What the Episode Covers

The episode is structured around a sequence of build demonstrations rather than a conventional news recap. Berman begins with the most visually dramatic example: a real-time 3D ocean in the browser. The prompt asks for much more than a decorative water plane. It calls for a real wave simulation, sea states from glassy calm to storm, whitecaps where waves break, realistic sky, clouds that cast shadows on the water, times of day through moonlit night, rain, lightning, a sailing yacht with wake, and an underwater view with light shafts and fish. In the demo, Berman changes conditions such as glassy, calm, breeze, rough, and stormy water; adjusts cloud cover, swell direction, choppiness, wave height, and sun bearing; shows lighting changing as clouds increase; clicks the water to create a splash; and moves underwater to show light columns and fish. He also says Sonnet opened the demos, recorded them, clipped them, and edited demo videos.

The next major cluster is game and tool generation. The Fall Guys-style browser game is built with Three.js and framed as one player competing against 59 bots across five rounds, ending in a final and crown ceremony. Berman stresses that the key prompting move is to ask the model to double-check its work while building, whether through screenshots, video, or literally playing the game. He shows jumping, diving, hints, bouncy areas, revolving doors, character color and pattern customization, many levels, and the ability to watch bots play after the user wins. He also notes a little clipping, which keeps the example grounded as a playable prototype rather than a flawless product.

The Lego creator demo explores a different product pattern. The prompt asks for a web app that accepts text or an image and produces a model made only from real Lego parts in real colors, displayed in 3D with step-by-step instructions and a parts list that can be uploaded to BrickLink. Crucially, the prompt tells the system not to let the AI place bricks directly. Instead, the AI describes shapes, a program converts them into real parts, checks that the model is one connected buildable piece, and sends problems back to the AI for repair. The rubber duck example has 281 steps, front/side/top/3D views, an imperfect explode view, lighting changes, and exports for PDF instructions, BrickLink wanted list, Rebrickable CSV, LDRAW, and JSON.

The heaviest demo is an Unreal Engine 5.8 downtown San Francisco build. The prompt asks for real scale, real city data, about 120,000 people and 2,400 vehicles, behavior driven by GAV with fallback rules, and an event input such as a fire at the Transamerica Pyramid that causes fire engines, traffic, and crowds to react. Berman says he also downloaded free assets. He judges it better than the Opus 5.5 version in some ways and worse in others, praising pedestrians, traffic patterns, parked cars, the Salesforce Tower, Union Square familiarity, freeway visibility, and varied buildings. The episode then adds team examples: Brian's Age of Empires-like Crownfall, Alex's Batman Arkham Knight-style city, Alex's realistic Mario world, and a Forza-style racing game. It closes with benchmark and pricing comparisons, including Sonnet 5.5 beating Opus 5.5 on Terminal Bench 4.0 in Berman's presentation, close scores on several other benchmarks, and cited pricing of $2 per million input tokens and $10 per million output tokens for Sonnet 5.5 versus $4 and $20 for Opus 5.5.

3. Core Views: Reasoning, Examples, and Limits

Berman's central claim is that Sonnet 5.5 feels very close to Opus 5.5 across the kinds of coding, 3D, and interactive-generation work he tested. The reasoning is cumulative rather than dependent on a single demo. The ocean project shows the model producing an interactive browser scene with simulation-like controls, weather and lighting variation, underwater rendering, click interactions, and edited demo output. The Fall Guys-style project shows game structure: levels, bots, movement, obstacles, win states, customization, and spectating after victory. The Lego generator shows a more disciplined architecture in which the model plans shapes while deterministic program logic handles real parts, connectedness, buildability, and export formats. The San Francisco project tests the model at a larger and messier scale: many moving people and vehicles, recognizable urban geography, event response, and Unreal Engine assets.

That evidence supports Berman's excitement, but it does not make “Sonnet solved 3D” a universal technical conclusion. The episode itself contains the limits. The San Francisco project overloads Berman's computer and sometimes stops working. It has glimmering on buildings and odd shadows. It took multiple days and millions and millions of tokens. The Batman-style city has shimmering windows. The Fall Guys-style game has clipping. The Lego generator's explode view works only imperfectly. A careful reading is that Sonnet 5.5 can now push many 3D and game prototypes into a playable or inspectable state, not that it eliminates performance engineering, rendering bugs, asset quality control, gameplay tuning, or production QA.

The second major view is economic. Berman says that after using Sonnet 5.5 heavily over the previous few days, he could not really tell the difference between Sonnet 5.5 and Opus 5.5, though he immediately leaves room for uncertainty by saying perhaps his tests were not hard enough. The benchmark slide, as presented in the episode, strengthens that impression: Sonnet 5.5 beats Opus 5.5 on Terminal Bench 4.0 and sits close to Opus 5.5 on Frontier Code, Cursor Bench, GDP Val, Humanity's Last Exam, OSWorld, and chart recognition. The price comparison then changes the practical conclusion. If a cheaper model is near the flagship across many day-to-day coding and prototype tasks, developers may reasonably try it first and reserve the more expensive model for unusually hard, ambiguous, or high-risk work.

The third view is that stronger generation makes engineering process more important, not less. The Code Rabbit sponsorship segment is commercial, but its argument aligns with the episode's technical reality: when AI helps produce more code, teams need better review, clearer change summaries, dependency-order understanding, and blast-radius analysis before production. Berman's prompting lesson from the Fall Guys clone points in the same direction. For complex interactive software, the prompt should not merely describe the desired final object; it should require the model to inspect, test, and correct its own work through screenshots, video, or play. This is less glamorous than the demo footage, but it is one of the more transferable lessons.

There are also narrower product boundaries. Berman says Sonnet 5.5 is still not good at creating sounds and music, and that none of his projects sounded good. That matters for games because audiovisual polish is part of the product, not a decorative afterthought. He also notes a writing quirk: the model tends to use British spellings such as “colour,” but he treats this as manageable because the model is steerable. These distinctions help separate serious limitations from minor preferences. Poor audio generation and heavy Unreal performance are workflow constraints; a spelling tendency is a prompt-control issue.

4. Learning and Application

The first practical lesson is to design prompts as production loops, not as image captions. For browser 3D, games, complex front-end tools, or Unreal scenes, the prompt should specify the target experience, the technical stack, the constraints, the validation method, and the failure loop. Berman's Fall Guys example is the clearest evidence: he says the model should be explicitly asked to double-check its work using screenshots, video, or actual gameplay. In practice, that means a useful prompt asks for a runnable version, observable checks, a list of known defects, and iterations based on those defects. The goal is not only to generate something impressive, but to create an artifact that can be inspected, played, measured, and improved.

The second lesson is architectural. The Lego generator is valuable because it does not ask the model to hallucinate a pile of bricks. It asks the model to describe shapes, then relies on program logic to convert those descriptions into real Lego parts, check connectedness, validate buildability, and export usable formats. That pattern can transfer to CAD tools, lesson builders, interior design systems, configuration software, and low-code creation tools. Put the model where language, planning, and user intent matter; put deterministic code where legality, geometry, inventory, export, and validation matter. The tradeoff is that this approach takes more engineering work than a free-form generation demo, but it creates clearer failure modes and better user trust.

The third lesson is to evaluate interactive demos dynamically. A still screenshot of the ocean would miss the importance of sea states, cloud-driven lighting, sun direction, click splashes, and underwater movement. A still screenshot of the San Francisco scene would miss pedestrian flow, traffic, parked cars, hardware load, and whether the city continues running. A screenshot of the Fall Guys-style game would miss whether jumping, diving, obstacles, bots, level progression, and victory states actually work. Teams using models like Sonnet 5.5 should therefore test frame rate, loading time, control feel, physics, collision, camera behavior, AI behavior, visual flicker, shadows, clipping, sound quality, and local hardware pressure.

The fourth lesson is budgeting. The San Francisco demo may be impressive, but Berman says it took multiple days and millions and millions of tokens. That should prevent teams from confusing demo feasibility with cheap production readiness. For large generated systems, the plan should include token budget, iteration time, asset sourcing, performance optimization, QA, and review. Sonnet 5.5 may lower the cost of reaching a strong prototype, but it does not remove the cost of turning that prototype into reliable software.

Finally, the episode suggests a model-selection strategy. Based on Berman's cited pricing and experience, Sonnet 5.5 looks suitable as a default model for many coding, front-end, prototype, and interactive demo tasks, especially when the task is clear and the result can be run quickly. Opus-level models can be reserved for harder architecture questions, ambiguous debugging, high-risk changes, or cases where Sonnet fails. But that strategy should be validated against each team's own codebase and tests. The episode's numbers and comparisons are presented by Berman; they are useful signals, not a substitute for local evaluation. Faster generation also increases the need for review systems that split large changes, explain why they exist, identify blockers, and surface possible downstream impact.

Source

More from WayDigital

Continue through other published articles from the same publisher.

Comments

0 public responses

No comments yet. Start the discussion.
Log in to comment

All visitors can read comments. Sign in to join the discussion.

Log in to comment
Tags
Attachments
  • No attachments