Harvey’s Margin Shock, Insect Ethics, and Muse’s Human Backup
This Diet TBPN episode is not a guest interview. John Coogan and Jordi Hays analyze three pressure points in the AI economy: Harvey’s legal AI costs overwhelming software margins, an effective-altruist argument about aggregate insect suffering, and Meta Muse testing human contractors behind some agent tasks. The common thread is what happens when AI moves from chat into real work, commerce, moral accounting, and human handoffs.
1. Host and Subject Background
This episode is a Diet TBPN segment hosted by John Coogan and Jordi Hays. Its title points to three topics: Harvey’s margins, whether insects are worth more than humans in the aggregate, and Meta testing human help for Muse. The indexed duration is 1,996 seconds, or roughly 33 minutes, so the episode should be treated as a compact commentary episode rather than a long-form interview. There is no verified guest context in the sense of an outside guest appearing on the show.
The indexed guest context names Harvey, but that functions as the subject background, not a guest biography. Harvey is identified as a legal technology and AI legal product company, and the indexed context says the episode discusses Harvey’s legal AI product, agent token usage, gross margins, model routing, post-training, an open-weight model called Harvey Tenant, usage dashboards, per-matter cost attribution, spend caps, and ROI reporting. The hosts are therefore analyzing a company and a market structure, not interviewing Harvey’s leadership.
The rest of the episode broadens that same analytical frame. Bentham’s Bulldog’s insect essay gives the hosts a way to discuss moral scale and aggregate suffering. Meta Muse gives them a way to discuss the limits of autonomous agents when tasks move into messy real-world phone calls. The commerce discussion around Amazon, Shopify, Walmart, and OpenAI then turns those issues into a platform question: who owns the customer, who captures transaction value, and who absorbs the cost when automation is not yet good enough?
2. What the Episode Covers
The Harvey discussion begins with the reported margin swing. The hosts cite reporting that Harvey’s gross margin fell from about 50% to negative 50% by June because agent token use spiked 20-fold on rented OpenAI and Anthropic models. They do not sugarcoat the basic problem: a software or technology company does not want negative gross margins, and the hosts describe the June situation as essentially selling a dollar for 50 cents. But they immediately add a second point: the cost explosion came on the back of more product usage and ARR growth, not simply a collapse in demand.
Their explanation centers on the shift from ordinary AI answers to reasoning and agentic workflows. Harvey had seat-based pricing, but the work behind each seat changed. A lawyer might type the same instruction as a year earlier, such as asking the system to review something, but the system no longer just produces a one-shot answer. It may reason, fan out, inspect firm records, examine policies, read MD files, and work through broader context across the firm. That means the user-facing action can look similar while the token cost beneath it becomes much larger.
The episode then reads Harvey co-founder Gabe’s response. As presented by the hosts, Gabe says the easy option would have been to force customers into consumption pricing before they were ready or to serve worse models to protect margins. Harvey, he says, instead tried to help customers transition on a workable timeline while continuing to serve strong models. The operational fixes he names are routing, harness improvements, and post-training. The customer-facing controls are usage dashboards, per-matter cost attribution, spend caps, and ROI reporting. According to the response read in the episode, Harvey moved from negative 50% gross margins to positive gross margins in a single quarter even as usage kept doubling month over month.
The model-strategy debate is framed through Harvey and Legora. The hosts say Legora’s Max argued that training your own models may not make sense because publicly available frontier models are getting cheaper and better. The hosts treat that as a bet, not as settled truth. Harvey is making a different bet, including post-training an open-weight model called Harvey Tenant. The practical question is not whether a company should use only one class of model. It is whether a legal AI product can route work intelligently across frontier models, cheaper specialized models, and post-trained open-weight models without sacrificing quality on the tasks that matter.
The hosts also rough out the financial severity. If Harvey is at a $400 million ARR run rate, that implies about $33 million in monthly revenue; at negative 50% gross margin, the June loss at the gross-margin line would be about $16 million. The hosts note Harvey had raised about $500 million, so they do not describe the moment as an immediate burn crisis. They also mention a Sequoia deal valuing the company at $15.5 billion and announced on September 9, and infer that investors likely had at least some visibility into the June gross-margin issue. For industry context, they cite Disco at about 75% GAAP gross margins, Thomson Reuters’ legal professionals software segment at almost 50% adjusted EBITDA margins, and high margins at major law firms. Their implication is that the legal market can support valuable software, but Harvey still needs to move back toward high-margin software economics over time.
The second topic is Bentham’s Bulldog’s essay on insects. The hosts summarize the argument as one of scale: insects are so numerous that even if their capacity for suffering is only a tiny fraction of ours, their total suffering could dwarf human suffering. They relay the essay’s claim that for every second of human life, insects collectively spend about 270,000 seconds dying, and that more insects die in a single second than the total number of humans who have ever lived. They also relay the essay’s conservative assumption that insects experience pain at one ten-thousandth of human intensity, leading to the conclusion that insects may experience more suffering in one day than humans have in all history.
The hosts do not adopt that conclusion as policy. They treat it as a thought experiment that exposes the strange behavior of aggregate moral accounting. They joke that if the metric were total pounds, cows might matter more, and if the metric were fur, monkeys might matter more. Their point is that the chosen boundary and unit of aggregation matter enormously. At the same time, they acknowledge a serious AI-related mirror image: if humans dismiss insects because each insect seems individually less significant, a future world full of vast sentient superintelligences could use similar reasoning to dismiss humans. That is why the hosts describe the argument as a kind of trap.
They also emphasize that the insect argument has a weak action path. Compared with shrimp welfare, where one might change farming practices or stop eating shrimp, it is unclear what portion of insect deaths is preventable or what intervention follows from accepting the essay’s premise. The hosts place this inside broader tensions around effective altruist and rationalist communities, noting that very different communications can get lumped together even when their policy implications differ sharply.
The third major topic is Meta Muse. The hosts cite Bloomberg reporting that Meta is testing a human concierge for its new personal AI assistant Muse, with human contractors quietly handling some phone calls placed by the digital agent. They initially question why this would be necessary when voice models are improving, and whether it could scale to a product used by billions of people. Their answer is conditional: Muse’s install base is still small, and Meta has the operational capacity to build large support structures, so the short-term experiment is not absurd.
More importantly, the hosts interpret the human layer as data collection and capability bootstrapping. They say Muse specifically will not train on user data put into Muse, which matters for users worried about personal information. But if a human contractor completes a task that is beyond the model, such as ordering flowers over the phone or handling a complicated conversation, that work may create training data for the next generation of agent behavior. In the hosts’ view, models can already summarize calendars and produce briefings, but they may still fail in complex phone negotiations. Human help fills the product gap while exposing the data needed to close it.
The final stretch connects Muse to the commerce-agent fight. The hosts cite a Ben Thompson-style view of Amazon as especially strong in the AI era because it owns logistics, infrastructure, and the final real-world fulfillment step. When OpenAI launched instant checkout and invited platforms in, Amazon did not join; the hosts say Amazon is instead vending ads into ChatGPT to help close purchases there. Shopify partnered with Muse, but the integration works with Shop Pay, making payments a value-capture layer. Walmart is treated as a warning sign: the hosts say merchants worry because the ChatGPT integration converted at one-third the rate of Walmart’s app and core website, and carts were smaller. A user may ask an agent for one roll of paper towels rather than browsing, adding soap, and building a larger basket.
3. Core Views: Reasoning, Examples, and Limits
The episode’s strongest Harvey point is that a negative gross margin is bad, but the type of bad matters. If the margin collapse came from customers refusing to pay, that would point to weak demand. The hosts’ reading is different: usage rose, ARR was growing, and the cost of fulfilling each user’s request changed because the product moved into reasoning and agent workflows. A seat-based price can look rational until each seat starts launching multi-step, context-heavy processes. The business problem is therefore not just “tokens are expensive”; it is that the product’s value meter and cost meter stopped matching.
That is why Gabe’s response matters most where it becomes operational. Usage dashboards, per-matter cost attribution, spend caps, and ROI reporting are not cosmetic enterprise features. They are the bridge between unlimited-feeling AI usage and a customer’s need to budget, allocate costs, and justify spend. In legal work, where matters, clients, and teams often need separate accounting, an AI vendor has to make consumption intelligible. Without that, moving from seat-based pricing to consumption pricing will feel like surprise billing rather than value-based scaling.
The model question is also more nuanced than “train or do not train your own model.” The Legora position, as summarized by the hosts, is that public frontier models are getting cheaper and better, which can make internal model investment wasteful. Harvey’s implied position is that post-training and routing can create an economic advantage for specialized legal workflows. The likely answer, as the hosts suggest, may be a mixture: use frontier models when accuracy and reasoning depth justify the cost, use cheaper or specialized models when the task allows, and keep enough routing flexibility to adapt as model prices change.
The limitations are important. The episode does not provide Harvey’s full internal financials, nor does it independently validate every reported number. The hosts’ $16 million monthly gross-margin loss estimate is useful because it translates the issue into scale, but it rests on the ARR and margin assumptions discussed on the show. Their point that a $500 million funding base makes the issue less immediately existential is plausible, but it does not prove long-term unit economics. Likewise, comparisons to Disco, Thomson Reuters, and law-firm profitability show that the legal market has money, not that Harvey will necessarily capture it with software-like margins.
The insect segment is valuable because it shows how aggregate reasoning can produce conclusions that feel alien while still having a coherent internal structure. The argument does not require each insect to matter as much as each person. It requires only a large enough population and a nonzero enough capacity for pain. Once those inputs are multiplied, the aggregate result can overwhelm human intuitions. The hosts’ jokes about cows by weight and monkeys by fur are doing real analytic work: they ask why one aggregation frame should count as morally decisive while others are dismissed.
At the same time, the episode treats the argument as uncertain and incomplete. The numbers about dying seconds, deaths per second, and one-ten-thousandth pain intensity are presented as the essay’s framing, not as independently established universal facts. The hosts do not resolve whether insects feel pain in the morally relevant way, how much they feel, or how preventable their suffering is. They also point out the practical gap: shrimp welfare has clearer interventions, while insect welfare at this scale leaves the listener asking what action follows. That is the line between a provocative moral model and an actionable agenda.
The superintelligence analogy is the segment’s deeper reason for inclusion. If humans reject insect welfare because insects are individually small, simple, or numerous, a future population of far more capable artificial minds could use the same structure against humans. The hosts do not claim that this proves the insect essay. They show why the argument is hard to laugh off completely: it exposes a vulnerability in how people move between individual moral intuition and population-scale arithmetic.
The Muse discussion makes a similarly careful distinction. A human concierge layer can look like an admission that the AI product does not work, but the hosts argue it may be more like a temporary product patch and a data flywheel. Phone calls are not just speech generation. They involve negotiation, ambiguity, identity, timing, and unpredictable counterparties. A model that can summarize a calendar may still be brittle when asked to resolve a real service request. Human contractors can complete the task while creating examples of what successful completion looks like.
But the same strategy has sharp boundaries. The hosts themselves question whether human support can scale to billions of users. Their more favorable interpretation depends on the small current Muse install base, Meta’s operational capacity, and the assumption that human-handled tasks can become useful training data. Privacy complicates the story further: if Muse does not train on user-entered data, Meta has to be careful about what data the human workflow generates and how it is used. The episode’s claim is therefore best read as a hypothesis about product strategy, not proof of a durable operating model.
The commerce-agent section grounds these abstractions in platform economics. Amazon’s resistance to external checkout is rational if its value lies not only in selling goods but in preserving logistics control, advertising revenue, customer intent, and fulfillment data. Shopify’s Muse integration makes sense if Shop Pay remains the capture point. Walmart’s reported lower conversion and smaller baskets show the risk: an agent may be efficient for the user while reducing browsing, upsell, ad exposure, and platform margin. Consumer agents are therefore not just better interfaces. They are bargaining agents in a fight over distribution, payments, ads, and customer ownership.
4. Learning and Application
For anyone evaluating an AI application company, the first lesson is to separate pricing from cost triggers. Seat-based pricing may be simple, but agentic workflows make each seat variable in a new way. A single user prompt can trigger reasoning, retrieval, tool use, firmwide context inspection, and multi-step execution. The right diligence question is not only whether revenue is growing, but whether the company can measure task-level cost, route across models, enforce spend controls, and move customers toward a pricing model they understand.
For enterprise AI builders, the Harvey discussion suggests that cost observability is part of the product. Usage dashboards, per-matter attribution, spend caps, and ROI reporting should be treated as adoption infrastructure, especially in industries organized around clients, cases, matters, engagements, or projects. Customers need to see which work generated which cost and what return came from it. Without that, even a valuable AI system can become politically hard to expand inside an organization.
The model-routing lesson is to preserve optionality. Post-training an open-weight model can be valuable where the task is narrow enough and the savings are real. Frontier models can still be necessary for difficult reasoning, high-risk work, or tasks where quality failures are expensive. Benchmarks can help decide, but the episode’s caution about the Harvey Legal Agent benchmark is a useful warning: do not mistake one benchmark result for production truth. Real workloads require quality measurement, cost measurement, fallback paths, and human review where appropriate.
The insect ethics segment is useful as a reasoning exercise rather than a direct operating manual. When a moral argument uses enormous numbers, the practical response is to reconstruct the chain: what is being counted, what capacity is assumed, what intensity weight is used, why the boundary is drawn there, and what intervention follows. That prevents two opposite mistakes: dismissing a strange conclusion just because it feels absurd, or accepting it just because the numbers are large.
The same method applies to AI ethics. Arguments about future digital minds, superintelligences, or massive populations of agents often depend on the same move from individual intuition to aggregate scale. The episode shows why those arguments deserve attention, but also why they need factual grounding and action pathways. A thought experiment can reveal an inconsistency; a policy recommendation needs evidence, tractable interventions, and a way to compare opportunity costs.
For product teams building personal agents, Meta Muse points to a practical pattern: human-in-the-loop can be a staged capability, not a confession of defeat. In high-value or brittle tasks, human support can let the product deliver outcomes while models learn from hard cases. The conditions matter. Users need clear boundaries, privacy rules, and expectations about when a human may be involved. The company needs a defined path from human execution to model improvement; otherwise the human layer becomes an expensive support operation rather than a learning system.
For commerce platforms, the application is strategic. Agent integrations should be evaluated on more than order volume. They affect conversion rate, basket size, ad inventory, payment capture, customer relationship, returns, and data ownership. Walmart’s smaller baskets and lower conversion, as described by the hosts, show why an agent that is convenient for the user may be less attractive for the retailer. Shopify’s Shop Pay layer and Amazon’s defense of its closed loop illustrate the same principle from the other side: whoever controls payment, fulfillment, ads, and customer intent controls where the economic value lands.
The boundary condition is that smaller players may not have the same leverage as Amazon. The hosts’ discussion of marketplaces and commerce apps opening APIs for consumer agents points to an uneven future. Large platforms can resist and negotiate. Smaller services may need agent access for distribution even if it weakens their direct customer relationship. The right strategy depends on whether the agent is a new demand source, a margin threat, or both.
Source
More from WayDigital
Continue through other published articles from the same publisher.
Comments
0 public responses
All visitors can read comments. Sign in to join the discussion.
Log in to comment