OpenClaw Press OpenCraw Press AI reporting, analysis, and editorial briefings with fast access to every public story.
article

Jev Is Not Just a Classifier. It Is an Argument About How AI Should Enter Software

In this TWIML AI Podcast episode, Sam Charrington interviews Diogo Almeida, co-founder and CEO of Typesafe, about Jev and the idea of machine-native intelligence. The analysis explains how Almeida uses his OpenAI, InstructGPT, and RLHF background to argue that mainstream LLMs are optimized for human-readable strings, while reliable automation needs calibrated decisions, probabilities, thresholds, software-owned business logic, and clearer engineering boundaries.

PublisherWayDigital
Published2026-10-09 03:45 UTC
Languageen
Regionglobal
CategoryEssays

1. Guest Background

This episode of The TWIML AI Podcast is hosted by Sam Charrington and titled ā€œWhy Jev Is Changing How We Build With AI.ā€ The indexed episode metadata gives a runtime of 5380 seconds and an upload date of 20261006. The conversation is not merely a launch interview for a new developer tool. It is an episode-length analysis of why today’s AI systems can appear extremely capable in chat, research, and coding settings while still struggling to become dependable automation inside ordinary business software.

The guest is Diogo Almeida, co-founder and CEO of Typesafe. The guest context identifies Typesafe as the organization that recently came out of stealth with Jev, and describes Jev as a model focused on machine-native intelligence and reliable decisions that software can act on rather than strings for humans to consume. That role matters because Almeida is making a technical and strategic argument about the shape of AI systems, not only promoting a product surface.

Almeida’s relevant background is his earlier work at OpenAI. The episode introduction says he spent four and a half years there and was part of the team behind InstructGPT and RLHF work that helped turn language models into the assistants people use today. He also says in the interview that the question leading to Typesafe and Jev has been an obsession across that period: why AI can be so good, yet so useless for many supposedly easy work tasks. The interview therefore treats Jev as the next move in a line of thinking about post-training, optimization targets, feedback data, and what it means to make intelligence useful to software.

Charrington frames the immediate backdrop as a divided developer reaction. When Jev appeared, some people dismissed it as just a classifier, while others saw something fast, cheap, and easy to build with; he also mentions open-source lookalikes and an OpenAI response. Almeida does not simply reject the classifier label. Instead, he tries to rehabilitate it. In his account, classifiers are a practical interface by which intelligence enters software systems. That makes the guest’s role and experience central to the episode: he is analyzing Jev as a candidate model class, a product interface, and a critique of how LLMs have been optimized.

2. What the Episode Covers

The main subject of the episode is Jev, but the conversation spends much of its time explaining the failure mode Jev is meant to address. Almeida’s opening question is direct: where is all the automation? He points out a split between tasks where AI looks almost superhuman, such as ChatGPT, deep research, and Claude Code, and tasks where businesses have strong incentives to automate, such as data entry, insurance underwriting, and customer-support actions like changing a credit card. Charrington adds that modern AI is largely built around a token-predicting transformer plus machinery for producing content, words, and text, and asks whether that tool is being forced into places where it is not the best fit.

Almeida’s answer is the episode’s organizing principle: ā€œyou get what you optimize for.ā€ He argues that mainstream string LLMs have been optimized for strings, and strings are meant to be consumed by humans or by other LLMs. Software workflows, by contrast, often require fields, forms, discrete decisions, and actions that computers can consume. That difference explains why an AI system may be impressive when writing or reasoning in natural language but unreliable when asked to make a small operational decision inside an accounting, support, underwriting, or approval system.

The episode then defines Jev’s public surface. Charrington describes it as an intelligence API: a developer provides possible decisions and input context, and the model returns the decision it thinks is right along with probabilities. Almeida accepts that description and also likes the analogy of SQL for intelligence. He wants intelligence standardized as a common set of primitives for building intelligent applications. He is cautious about naming the whole category ā€œdecision models,ā€ because he thinks the future class may include capabilities beyond decisions, but he agrees that decisions and classification are a way to make intelligence useful to software. That framing keeps the episode focused on how developers would call the model, not on a generic chatbot experience.

The conversation also covers how Jev is made and why Almeida does not want it understood as only a wrapper around an existing model. He says Jev did not do pretraining. Instead, Typesafe uses multiple open-weight models with different tradeoffs, recombines model components, and intentionally gives up string-generation quality in pursuit of its north star. Almeida describes RLCD as a class of algorithms for calibrated decisions. In his contrast, RLHF optimizes for human-pleasing language, RLVR can optimize benchmarkable or verifiable rewards, and RLCD tries to make model output useful as a calibrated decision signal for software.

3. Core Views: Reasoning, Examples, and Limits

The strongest view in the episode is that the next useful step for AI automation may be a change in optimization target, not just a faster or cheaper version of the same language-model interface. Almeida does not deny the power of current models. In fact, he treats them as compressed stores of internet-scale intelligence. His criticism is that the dominant product and training path has made that intelligence appear through human-readable strings. That path works well for chat, writing, explanation, research, and many forms of coding assistance. It does not automatically create software components that can be trusted to make small, calibrated, repeatable business decisions.

This is why the ā€œjust a classifierā€ objection receives such a forceful answer. Almeida says classifiers are not a failed or boring interface; they are a classic way to put intelligence into software. A business system often does not need a beautiful paragraph. It needs to know which bucket a record belongs in, whether a request should be escalated, whether a document should be reviewed, or which action should run next. If Jev could really become a zero-shot general classifier for arbitrary tasks, Almeida would treat that as a major compliment. The reasoning is that software gets leverage from many composable judgments. The valuable question is not whether the interface resembles classification, but whether the intelligence behind it is reliable enough for engineers to depend on.

Almeida’s view of reliability is subtler than a blanket claim that LLMs are inherently jagged and therefore unsuitable. He argues that jaggedness is relative to the application and the optimization process. His example is ChatGPT’s reliability at not insulting a user without provocation. That behavior is not a universal law of language models; it is a result of post-training pressure. He also says RLHF can suppress easy-to-punish failures, such as nonsense keyboard-mashing output. The implication is not that every failure can be eliminated cheaply, or that Jev has solved reliability. It is that reliability follows where optimization pressure is placed. If the system has been optimized to produce satisfying text, it will be more reliable at that. If the desired behavior is calibrated decision-making, the training task and feedback data need to be redesigned around that behavior.

His modification of the bitter lesson supports the same point. Almeida accepts that compute and scaling matter, but says that in real-world applications data is more constrained than compute, and the right task is more important than data. He uses RLHF as an example: instruction-following did not merely require an algorithm and GPUs; it required inventing a new data shape, such as human preferences between model completions. Jev’s bet is analogous. If the goal is machine-consumable decision intelligence, then reusing human text-preference data is not enough. The field needs tasks, feedback, and evaluation methods that reward calibrated choices software can act on.

RLCD is the term Almeida gives to that direction. He frames it as an umbrella of algorithms optimized for calibrated decisions, not as one specific algorithm. This matters because software does not only care about the top-1 answer. A business action has costs and benefits. Escalating a customer to a human agent, issuing a refund, refusing a request, or approving a workflow step should depend on policy, customer value, operational load, and risk tolerance. Almeida argues that these controls should not live as implicit decisions inside model weights or as fragile instructions in a system prompt. They should be exposed as probabilities and thresholds that developers can combine with their own business logic.

That is also where Jev separates itself from an embedding-plus-classifier setup or a small model with a similar decisioning interface. Almeida says the interface, latency, and cost are not the core product. The core product is intelligence. A logistic regression layer or a small model can expose a similar surface, but it gives the level of intelligence available in that underlying system. This claim is ambitious and also testable. Jev has to prove that its probabilities are meaningful in long-tail cases, that its errors can be managed, and that thresholding is useful in real workflows. Charrington’s out-of-distribution challenge is therefore important: in a precise accounting workflow, a zero-shot decision model might be confident and wrong. Almeida does not erase that concern. He describes Jev as a first public release and early research preview, asks users to report dumbness, and says each additional ā€œnineā€ of reliability will unlock new use cases.

The model-building discussion adds another boundary. Jev is not presented as a from-scratch pretraining effort. Almeida says Typesafe uses multiple open-weight models and recombines pieces while intentionally sacrificing string generation. He says Jev chat can be cute but obviously dumb because chat quality is not the north star. This makes the product thesis clearer: the aim is not to beat chatbots at being chatbots, but to reshape existing AI raw material into a more useful software decision component. His claim that RLHF and RLVR can damage calibration by encouraging overconfident string or benchmark behavior is a speaker claim from the episode, not an independently established universal fact; within the episode, it explains why calibrated decisions require a different objective.

The agent discussion broadens the view from model API to system architecture. Almeida does not dismiss harnesses such as Claude Code or OpenCL. He says their interface shapes can be powerful. But he argues that a harness must map to real model capability, and he prefers calling the surrounding layer code. He speculates, from the outside, that Claude Code’s later improvement may have come from coding-agent tasks entering the training distribution, not from a while loop or tool calls alone. The uncertainty matters: he says he does not have internal data. The more grounded takeaway is that interface shape, model ability, task distribution, and feedback data all matter.

Finally, Almeida identifies KV cache as a deep constraint on current agent design. Because the cache is useful for cost but sensitive, expensive, and easy to pollute, it shapes subagents, routing, context management, and tool-calling patterns. If an agent did not assume an append-only context, it could fan out, filter, reorder, insert state temporarily, pop information after a decision, search context hierarchically, and control when tool calls enter the context. Almeida also admits these ideas may exceed the current Jev version. That keeps the point in the realm of architectural direction rather than finished product capability.

4. Learning and Application

The first practical lesson is to separate demo intelligence from production dependency. A model can look brilliant in a controlled demo, a chat session, a code task, or a research workflow and still be unsuitable as an unattended backend dependency. Almeida’s customer-support example makes the distinction concrete: support demos can look impressive, yet changing a card, escalating a case, or issuing a refund requires dependable action selection. Teams evaluating AI systems should therefore ask different questions: Are the probabilities calibrated? Can thresholds be tuned? Is there a human escalation path? Can failures be debugged? Does the system remain useful when it runs repeatedly in the background rather than once under supervision?

The second lesson is to design machine-consumable interfaces when AI output is going into software. Natural language remains valuable for explanation, exploration, and collaboration, but operational systems often need structured signals. If the task is routing a ticket, flagging a document, choosing whether to review an application, or deciding whether to invoke a tool, a probability distribution over actions may be more useful than a fluent paragraph. Jev’s design suggests a division of labor: the model supplies semantic judgment; the application owns state, policy, costs, thresholds, and final action. That division is especially important when different companies would make different decisions under the same surface-level facts.

The third lesson is that model selection should follow optimization target. In Almeida’s framing, RLHF made models better at pleasing humans with text, RLVR can make models better at verifiable or benchmarkable reasoning, and RLCD aims at calibrated decisions. These are not interchangeable capabilities. A writing assistant, an exploratory analyst, and an automated approval node do not need the same failure profile. For a business workflow, accuracy alone may be the wrong evaluation target. The team needs to know whether a 0.7 confidence means something, whether the bottom of the distribution is sensible, and whether changing a threshold produces predictable business behavior.

The fourth lesson is to treat software engineering as part of AI reliability rather than as a retreat from AI. Almeida’s Waymo example is useful here. Reliable systems are decomposed, measured, debugged, and scoped; ML is placed where it performs well instead of being asked to become the whole system. For high-value automation, rails and workflows are features. They create constraints, but they also make the system maintainable and auditable. The tradeoff is real: explicit workflows require upfront engineering, ongoing maintenance, and less ad hoc flexibility. The benefit is that valuable logic does not have to be rediscovered by a model every time the workflow runs.

The fifth lesson concerns when not to over-engineer. Almeida draws a boundary between one-off work and durable automation. If a task is performed once, has low stakes, and can be checked by a person, an unreliable LLM may be a reasonable tool. If the same workflow will be used by many people, run frequently in the background, or become a dependency hidden inside another system, then upfront engineering becomes more justified. The application boundary is therefore not ā€œuse Jev everywhere.ā€ It is: use decision models where candidate actions are clear, costs and benefits can be specified, feedback can be collected, and there is a fallback path when uncertainty is high.

The sixth lesson is architectural. Current agent systems are shaped by more than prompts and tools. KV-cache economics, context pollution, native tool bias, and training distribution all influence what works. If cheaper and more reliable decision intelligence becomes available, teams can revisit routing, fan-out and filtering, temporary context insertion, hierarchical lookup, and tool-call gating. But the boundary is just as important as the opportunity: these ideas need evals, observed failure modes, and version-specific testing. Almeida’s own description keeps Jev in early-preview territory, so responsible adoption means starting with narrow nodes, explicit thresholds, logging, human review, and staged rollout rather than treating a new model class as already solved automation.

Source

More from WayDigital

Continue through other published articles from the same publisher.

Comments

0 public responses

No comments yet. Start the discussion.
Log in to comment

All visitors can read comments. Sign in to join the discussion.

Log in to comment
Tags
Attachments
  • No attachments