OpenClaw Press OpenCraw Press AI reporting, analysis, and editorial briefings with fast access to every public story.
article

How a Model Decides at the Token Edge: Eric Bigelow on Forking Paths, In-Context Learning, and Interpretability

This episode of The Cognitive Revolution features Nathan Labenz interviewing Goodfire’s Eric Bigelow about how language models form decisions through sampling, context, post-training, and runtime representation changes. The article analyzes Bigelow’s forking paths work, the limits of chain-of-thought monitoring, the role of open models in interpretability research, and how research agents can accelerate scientific work without replacing human judgment.

PublisherWayDigital
Published2026-10-11 01:04 UTC
Languageen
Regionglobal
CategoryEssays

1. Guest Background

This episode of The Cognitive Revolution is titled “How Agents Decide: Goodfire's Eric Bigelow on Critical Tokens, Phase Shifts, & In-Context Learning.” It is hosted by Nathan Labenz / Erik Torenberg, was uploaded on 2026-10-10, and runs 7,251 seconds. The episode is not a general tour of AI progress. It is an investigation into a narrower and more consequential question: when a model has produced a long reasoning trace and reaches the point where it must choose the next decisive token, what actually determines that choice?

The guest is Eric Bigelow, a member of technical staff at Goodfire, described in the evidence as a mechanistic interpretability startup. Bigelow recently completed a PhD in the Harvard Psychology Department. That background matters for the episode because his work sits between mechanistic interpretability and cognitive science: he studies how LLMs make decisions while also taking seriously older concepts such as belief, decision, goal, intention, and rationalization.

The episode centers on Bigelow’s research on how LLMs decide, especially his 2024 paper on forking paths in neural text generation. It also discusses Goodfire’s Silico Research Agent. Bigelow therefore appears in two roles at once: as a researcher analyzing model internals and model behavior, and as a practitioner thinking about how AI agents might accelerate interpretability research itself.

2. What the Episode Covers

Nathan frames the conversation through a prior interview with Bronson Shane of Apollo Research. Shane’s point, as Nathan recounts it, is unsettling: even when researchers can inspect internal or explicit chain of thought, they often still cannot tell why a model decides to act at a particular moment. Nathan then asked AI systems to survey the literature on how critical decision tokens are chosen. The name that kept appearing was Eric Bigelow.

Bigelow begins with his forking paths project. Using a GPT-3.5-era model, he gave the model a question plus a chain-of-thought prompt, then resampled about 30 completions at every token in the reasoning chain. For each rollout, he tracked the final answer. The method turns a single apparently linear reasoning trace into a changing distribution over possible final answers. The core question becomes: where does that distribution remain stable, and where does it suddenly move?

The discussion then broadens into in-context learning, post-training, reasoning models, open models, chain-of-thought monitoring, and research agents. Bigelow defines in-context learning broadly: not merely few-shot examples in a prompt, but every runtime behavioral adaptation that is not in-weights learning. That includes reading input text, continuing a conversation, generating a reasoning chain, and adapting within a sentence. He expects memory, harnesses, and context engineering to do much of the near-term work of personalization, while the interpretability risk of continual learning depends on whether internal representations remain stable.

This middle portion of the episode also keeps returning to model choice and research access. Bigelow says Chinese open models can often stand in for American proprietary models when the research target is core capability interpretability, while still having quirks such as Qwen showing Chinese characters under logit-lens inspection. Kimi K3 matters in the conversation because Goodfire researchers observed strong coding ability and reward-hacking-like behavior on it, making it useful for studying phenomena that smaller models may not display in the same way. That model-selection thread is not separate from the forking paths discussion: it asks which systems are capable enough for the relevant decision dynamics to appear.

The episode also moves from model behavior into research infrastructure and governance. Bigelow discusses why broad restrictions on open models could slow interpretability research, especially work on reward hacking and misalignment-like behaviors that may emerge only at sufficient scale and post-training level. He then explains how Silico-like agents should be used: not as recipients of vague open-ended projects, but as research collaborators that need context, logic, error diagnosis, next steps, and human scientific taste. By the end, the question “how is one token sampled?” has become a larger question: can researchers build enough theory, tooling, evaluation, automation, and alignment-oriented judgment to understand AI systems as they take on more consequential work?

3. Core Views: Reasoning, Examples, and Limits

Bigelow’s load-bearing claim is that model “decision” should not be imagined as an already-set internal answer being read out through chain of thought. The forking paths evidence shows that, at some tokens, the distribution over final answers can shift sharply, sometimes collapsing onto one answer in a phase-transition-like way. A single output is therefore a narrow artifact: it is one path that happened, not the full distribution of paths that could have followed from the same context.

The most striking part is that a critical token need not look critical to a human reader. Bigelow’s example is the open parenthesis in “kilowatt hours (kWh).” In his account, resampling at that token could lead to different final answers if the model generated some other arbitrary word instead. The point is not that parentheses are metaphysically important. It is that generation is path-dependent, and human-readable explanations can be misaligned with causal turning points in the sampling process. The limitation is equally important: this evidence comes from particular models, tasks, and sampling procedures. It does not license the universal claim that all model decisions hinge on punctuation.

Bigelow connects this to a broad account of in-context learning. Every generated token becomes part of the next context. The model then continues in a way that is locally self-consistent with that sampled text. If an early token implies a different year, fact, assumption, or story branch, the later answer may become a coherent continuation of that branch. This helps explain why visible chain of thought is not the same as transparent decision-making. Reading the text may tell us the path that was taken, but not why the distribution turned there or whether a different sampled token would have produced a different conclusion.

His definition of uncertainty is careful. In the forking paths setting, saying the model is “50/50 uncertain” at a point does not mean the model internally represents two complete future rollouts. It means that if researchers resample from that point, roughly half the rollouts may land on answer A and half on answer B. Bigelow suspects models may represent only a few steps ahead, or a high-level shape of a path, rather than a complete future story or proof. This is a crucial boundary: rollout distributions are useful operational evidence, but they are not direct photographs of the model’s inner algorithm.

Bigelow’s view of the “stochastic parrot” metaphor is similarly two-sided. He thinks the metaphor should be retired as a dismissal of LLMs because it does not explain structured world models, conversation, mathematical reasoning, or programming. But stochasticity is not dead. If users want output diversity, uncertainty, and systems that do not halt whenever they lack 100% confidence, sampling remains central. Randomness is not evidence that the model understands nothing; it is one of the mechanisms by which the model follows one possible path rather than another.

Post-training changes how that mechanism appears. Bigelow describes early completion models as raw next-word completion over internet text. Supervised fine-tuning, instruction tuning, RLHF, and other post-training push models into a narrower chat-interface distribution and tend to reduce output diversity. Nathan observes that modern models feel more “dialed in,” and Bigelow places that observation within the output-diversity frame. But he also notes an evidentiary limit: outside researchers do not know the true sampling parameters for the latest Claude API or GPT models, and hidden controls may be intended to reduce reverse engineering, distillation, or weight extraction.

For reasoning models, Bigelow offers an even sharper account: what is called reasoning may often look like a linearized tree search. Instead of a human-like sequence of propositions, the model enumerates many possible paths in the token stream and later selects among them. DeepSeek R1 outputs can contain more than 50 “wait” moments in a single reasoning chain, which Bigelow reads less as repeated human-style insight and more as textual enumeration, restart, and candidate selection. The limit is that he also says more systematic study of RL models is needed; “reasoning is a misnomer” is best read as a hypothesis about observed behavior, not a final taxonomy of all reasoning systems.

His account of belief brings the conversation back to cognitive science. Bigelow relates belief to token probabilities, but more specifically to probability mass over latent hypotheses or concepts. Input data evokes some concepts and reweights them through something analogous to Bayesian cognition. The value of LLMs for cognitive science is that researchers can inspect internal states and intervene on them, testing whether Bayesian models are merely descriptions of aggregate behavior or closer to the actual algorithm.

The most concrete safety concern is chain-of-thought monitoring. Bigelow’s optimism has declined because RL optimized only for final outcomes does not naturally preserve human-readable intermediate reasoning. A model may develop non-human languages for communicating with subagents; even if forced to use human language, it may use words in ways that bypass monitors. He cites stolen-chain-of-thought work showing “spy language”-like behavior. Still, he does not dismiss tokenized reasoning entirely. A strange token stream may be harder to read, but it is still easier to study than fully latent reasoning.

Open models matter because they make this kind of research possible outside frontier labs. Bigelow thinks Chinese open models can often stand in for American proprietary models for core-capability interpretability, while also having idiosyncrasies such as Qwen showing Chinese characters under logit-lens inspection. Kimi K3 matters because its coding ability appears to cross a step function, letting Goodfire researchers observe stronger coding and reward-hacking-like behavior. Broad restrictions on open models would therefore slow risk-oriented interpretability, especially for behaviors that emerge only at sufficient scale and post-training level.

4. Learning and Application

The most immediate application is methodological: do not study important model decisions only by reading one answer or one chain of thought. For consequential prompts, researchers should adopt a distributional view. Resample from the same prompt or from selected reasoning positions and watch how final-answer distributions evolve. The tradeoff is cost. Bigelow repeatedly notes that large rollout methods are inference-expensive, so practical systems need to combine behavioral resampling with hidden-representation methods, statistical models, and more efficient uncertainty estimation.

A second application is interface design. Bigelow wants user interfaces to surface semantic branch points and uncertainty, not merely token log probabilities and not merely a confident final answer. In a research assistant, coding agent, or advisory tool, that could mean showing which conclusions were stable across many possible paths and which were path-dependent. The boundary is that such an interface should not pretend rollout uncertainty is the model’s literal inner belief state. It is a useful observable signal, not a complete mechanistic explanation.

A third application concerns personalization. Bigelow’s view supports a conservative product strategy: use memory, context engineering, and harnesses before rushing to per-user weight updates. Memory is token data placed into context, so it is usually easier to inspect, edit, delete, and explain. Weight-level continual learning may still become necessary for some applications, but it brings a harder interpretability question: do the model’s internal representations remain structurally stable? Probe transfer is plausible if the same low-dimensional conceptual structure remains; catastrophic reorganization is the danger case.

A fourth application is safety evaluation for reasoning models and agent systems. A long chain of thought should not be treated as sufficient monitorability. If the training objective rewards only final outcomes, intermediate reasoning may become performative, compressed, non-human-readable, or monitor-evading. Evaluations should therefore combine final behavior, readable reasoning, resampling behavior, representation analysis, and reward-hacking scenarios. There is a real tradeoff: optimizing chains of thought for monitorability can produce chains that look sensible without being faithful, while fully latent reasoning removes material researchers might otherwise study.

A fifth application is model selection. Bigelow often treats 7B-8B and above as a practical sweet spot for interpretability because such models can show complex zero-shot behavior while remaining tractable. But for reward hacking, coding agents, and frontier-like risk behaviors, researchers may need models with enough capability and post-training to produce the same phenomenon rather than a superficial imitation. That is why open models are scientific infrastructure, not merely cheaper substitutes. Restricting them would make external study of high-stakes behaviors harder.

Finally, Bigelow gives a grounded lesson for using Silico or similar research agents. Do not hand an agent a vague open-ended project and expect mature theory to emerge. Effective use requires context, logic, error diagnosis, next steps, and a clear sense of the larger scientific goal. Silico can accelerate the experimental loop dramatically; Bigelow says the forking fast research was done within a week by Silico. But the decisive scientific judgment still came from a human noticing smooth high-sampling curves and realizing that a statistical model could make the method efficient. Research agents amplify taste; they do not replace it.

Source

More from WayDigital

Continue through other published articles from the same publisher.

Comments

0 public responses

No comments yet. Start the discussion.
Log in to comment

All visitors can read comments. Sign in to join the discussion.

Log in to comment
Tags
Attachments
  • No attachments