Why Enterprise AI Agents Stall Before Production: Manoj Saxena on Runtime Control, Trust Posture, and the Agent Workforce
Once an AI agent can sign into databases, call tools, and act for a business, accuracy is no longer the only trust question. The harder question is who can constrain the action before it happens. Eye on AI host Craig S. Smith and TrustWise founder and CEO Manoj Saxena discuss that production gap through AI control towers, guardian and Genesis agents, modular shields, the semantic action layer, cost, regulatory evidence, and the limits of the company’s own claims.
1. Guest Background
Manoj Saxena comes to enterprise AI from a specific commercialization problem. In the Eye on AI conversation, he introduces himself as the founder and CEO of TrustWise and calls it his fifth startup. Earlier, he helped commercialize IBM Watson. He recalls that after IBM’s board asked him to turn the Jeopardy-playing system into a deployable product, he spent about three years shrinking it from the size of a large bedroom to a pizza-box-level system.
That history explains his emphasis throughout the interview. A striking model demonstration is not the finish line. The harder work is turning the system into something an enterprise can deploy, approve, pay for, audit, and operate over time. The conversation runs just under 62 minutes. Its title uses the claim that “95% of AI Agent Projects Fail to Reach Production,” but the percentage comes from an MIT report Saxena cites in the episode. It is part of the episode’s framing, not an independently established measure for every industry, period, and agent project.
Saxena says the idea for TrustWise formed after ChatGPT, when he watched the industry race toward larger and more capable models without building an equivalent protective structure. His analogy is a larger nuclear core without a dome around it. TrustWise positions Harmony AI as that missing control layer: an “AI control tower” for generative and agentic systems, rather than another foundation model or application interface.
2. What the Episode Covers
The conversation starts by narrowing the meaning of trust. For Saxena, the question is no longer only whether a model produces an accurate answer. It is whether the AI does what the person or business intended. He points to three changes: AI has moved from generating emails, images, and answers to taking actions; enterprise architectures are moving from single-model workflows to multi-agent, multi-model systems; and policy can no longer be enforced only before deployment. It has to be applied continuously across tool calls, actions, and outputs.
That is why he describes agents as a digital workforce. A task may run for minutes, hours, or days. If a company introduces thousands of digital workers, he argues, it needs infrastructure analogous to HR and finance: onboarding, policy handbooks, performance review, budgets, kill switches, and audit evidence. In this framing, governance is not a document approved before launch. It is an operating system for digital labor.
TrustWise divides the control-tower agents into two classes. Guardian agents handle onboarding, supervision, alignment, and behavioral control. Saxena emphasizes human-in-the-loop design and deterministic outcomes because enterprises cannot manually onboard and test very large agent populations. Genesis agents are aimed at higher-value exploration. Later in the interview, he describes them as systems that search for unknown unknowns and generate hypotheses. He is also clear that TrustWise’s present focus is still mostly on guardian agents.
Modular AI shields are capability packs attached to a guardian agent, not independent agents. Saxena groups them into safety and security, compliance and regulation, and cost and carbon management. TrustWise calls the combined discipline trust posture management. That is the company’s category claim: it does not want to be only a security layer, only governance, or only observability. It wants to trade off safety, compliance, efficiency, and evidence while an agent is acting.
The product shape follows that idea. Saxena says TrustWise is primarily a collection of APIs and CLIs, not another dashboard, though it provides a reference UX for human operators. That interface supports four jobs: onboarding agents; assessing them through simulation-like tests across different jobs and action paths; monitoring production drift; and generating provable evidence for audit and learning. Underneath sits the semantic action layer, described as a time-series semantic graph of models, data, policies, constraints, and context. It is meant both to steer behavior and to reconstruct what happened, while also producing tables, reports, and natural-language answers.
Ownership is spread across the enterprise. Saxena says agent risk has reached boards and CEOs. Production approval may involve the CIO, head of AI, responsible-AI leaders, risk, compliance, security, and finance, while day-to-day operation spans business and IT, risk and compliance, and internal audit. He cites an MIT report saying that 95% of agent projects fail to move from pilot to production, then argues that the bottleneck is confidence among those functions rather than the difficulty of assembling an agent. That causal explanation is his view, but it identifies the episode’s central concern: approval and continuing control are harder than a demo.
TrustWise began with banking and insurance, then expanded into healthcare, retail, and consumer products. Saxena says Yum Brands is exploring runtime control for voice ordering across Pizza Hut, KFC, and Taco Bell. He also says Hitachi is an investor and is looking at TrustWise as an OEM control layer across 650 business units. The scope matters: these are Saxena’s descriptions of customer exploration and partnership direction, not independent evidence that every cited deployment is complete.
Technically, the control tower sits above agent orchestration and below the experience layer. It connects through a gateway, probe, or proxy and can run in live, sidecar, batch, or pre-production simulation mode. Saxena says it is agnostic to agents, frameworks, clouds, and models. He mentions Azure, AWS, GCP, and Dell on-prem deployments in partnership with Nvidia, along with agent sources such as LangGraph, ServiceNow, Microsoft Copilot, and Claude. The first product release was called Optimize AI and focused mainly on generative AI. The current name, Harmony AI, reflects the goal of coordinating multiple agents from multiple vendors; the broader category remains the AI control tower.
3. Core Views: Reasoning, Examples, and Limits
Saxena’s clearest conceptual claim is that AI is not an application but an actor. Traditional enterprise software waits for input and follows relatively stable rules. An agent may sign into websites, access databases, call tools, and take a different action tomorrow as data, patterns, and context change. Governance therefore shifts from asking whether an application is compliant to asking which paths an actor is allowed to take. His emerging enterprise stack runs from compute and data to model and agent orchestration, then runtime control, with user or machine experience on top.
Runtime governance differs from declarative governance at the moment of decision. Saxena defines runtime control as a judgment made before the action, not a policy written weeks earlier or an audit log examined hours later. His UK FCA Consumer Duty example makes the distinction concrete. An 85-year-old applicant who recently lost a spouse and is in financial distress may require different tone, clarity, helpfulness, approval, refusal, or escalation behavior from a 24-year-old applicant who has just started a job. Before the agent proceeds, the system has to decide whether the customer is distressed, whether the action and tone are allowed, whether a human must approve, and whether uncertainty should trigger refusal or escalation.
His criticism of existing enterprise tools follows that timing problem. Security is usually outside-in defense and may not address an agent as an internal behavioral risk. Governance can define policy without steering behavior at runtime. Observability can show what happened without changing it. This argument is not equally strong for every AI system. A low-risk drafting assistant may not justify a full control tower. Once an agent has permissions, calls tools over time, and affects banking, insurance, healthcare, claims, or voice-ordering outcomes, after-the-fact logs are a poor substitute for control before action.
TrustWise also places itself in the role of a vendor-agnostic meta control plane. Saxena’s reasoning is that large companies rarely standardize on one model, cloud, orchestrator, or agent vendor. Each platform may have a control plane, yet the enterprise still needs control and evidence across the estate. The logic fits the fragmented reality of enterprise AI procurement. The limitation is equally plain: the episode offers company descriptions, customer examples, and deployment claims, not an independent comparison of cross-platform reliability.
Saxena supplies a set of runtime performance details. He says the platform uses 12 small, high-precision, low-latency models and 15 components he calls “scales and orchestrators.” Depending on the workload, he puts evaluation and steering between roughly 300 milliseconds and 10 seconds. The numbers show how TrustWise says it approaches real-time control, but the interview does not disclose test methods, workload distributions, or error rates. They cannot establish how the system performs across enterprise environments.
Cost is not a side issue in this account; it is part of trust. Saxena says one input can trigger 20 to 50 actions and consume 20x to 40x the tokens used by a generative-AI system two years earlier. He also says TrustWise case studies have shown up to an 83% cost reduction while improving safety by 40% and latency by 60%. Those percentages are company-claimed case-study results, not independent industry benchmarks. The sounder conclusion is narrower: when agents loop, retrieve, retry, call other systems, and collaborate with other agents, cost, latency, and safety become coupled runtime variables.
The regulatory product claims are similarly specific and similarly attributable. Saxena says TrustWise ships more than 1,100 controls across 17 regulatory and risk frameworks, including NIST AI RMF, the EU AI Act, OWASP, HIPAA, SR 11-7, and UK FCA Consumer Duty. He says the libraries are updated every 30 days and that customers and partners can create custom policies. This supports the company’s claim that it translates requirements into executable runtime policies rather than merely collecting logs. The episode does not independently test the completeness, update quality, or regulatory acceptance of those controls.
The semantic action layer is the most important and least externally verifiable part of the architecture. Saxena argues that knowledge graphs describe relationships, context graphs add situation, and world models capture broader regularities, but none is enough to govern what an agent may do at a particular moment. TrustWise’s layer is said to predefine allowable action paths and permission pathways tied to an alignment framework. The autonomous-driving analogy is useful: instead of recomputing the whole map every few meters, the system follows a modeled set of valid routes and checks the next move. Saxena also calls the implementation core IP, so the episode establishes the intended concept and purpose, not technical sufficiency.
The alignment framework has six levels: global rights, national law, industry rules, company values, business-unit and workload context, and customer SLA or engagement requirements. Saxena argues that a one-time prompt is too weak for autonomous loops that set up experiments, act, and review results over hours or days. Continuous alignment matters more as prompting becomes less central. He also says foundation models contain an imperfect and largely Western snapshot of human values, while companies are already starting with corporate values as a practical layer. Runtime control is therefore not purely an engineering problem; it encodes organizational choices, regional differences, and delegated authority.
The discussion extends from UX to AX, or agent experience. Saxena predicts that within three years, 90% of AI control-tower users will be agents rather than humans. Craig S. Smith raises a drug-discovery example in which one AI writes documents for people and another AI reads them; the middle layer might eventually give way to direct structured-data exchange. Saxena connects that idea to multi-agent, multi-enterprise workflows. It is a forecast, not a completed transition. Regulation, accountability, audit formats, and institutional trust will determine whether natural-language intermediaries can actually disappear.
Genesis agents reveal the exploratory side of the thesis. Saxena describes them as hypothesis-generating systems for finding unknown unknowns in revenue leakage, fraud, and related domains. He argues that hallucination can be a feature as well as a bug when it is controlled and validated, then supplies his own boundary: “maybe only two out of ten is right.” That places Genesis agents in discovery and analytical augmentation, not automatic truth or execution. He goes further and argues that vertical AGI or autonomous finance depends on enterprise context, data, and policy, so model providers cannot solve enterprise alignment by themselves.
Saxena calls the emerging category cyber trust rather than cybersecurity. He relays estimates of roughly $2 billion for the current year and $80 billion by 2030, and expects governance, security, and observability vendors to enter. Like the 95% production figure, the performance percentages, and the three-year 90% forecast, these market numbers belong in the category of speaker-cited estimates and company thesis. The interview does not independently establish them as market facts.
4. Learning and Application
Enterprise teams can turn the episode into a sharper production question: not “can the agent demo?” but “what is it allowed to do?” A drafting assistant can carry a lighter control burden. A system that signs into sites, accesses databases, calls tools, affects customer handling, or acts continuously for the business needs explicit allowed paths, forbidden paths, human approval points, refusal conditions, and escalation rules. Saxena’s digital-workforce analogy is ultimately a job, permission, policy, and accountability design.
That design becomes a production-readiness checklist. What is the agent’s job boundary? Which policy handbook governs it? Who approves release? How is it simulated before production? Which drift signals are monitored? When does a kill switch fire, and who takes over after abnormal cost, unauthorized tool use, or missing evidence? An audit team should be able to reconstruct a consequential action years later: which models, data, policies, and context were used; which tools were called; where a human intervened; and what business result followed. TrustWise’s reference UX and semantic action layer are organized around those four jobs of onboarding, assessment, drift monitoring, and evidence reconstruction.
The risk inventory also needs three AI surfaces. Saxena calls them the AI a company builds, the AI it buys, and the AI that may be used to exploit existing mainframes, chatbots, databases, and models: build, buy, and protect. In practice, that means governance cannot stop with internal development teams. It has to include SaaS agents, purchased capabilities, and the exposure of existing assets. Boards, CIOs, AI leaders, risk, compliance, security, finance, business teams, and internal audit also need a shared view rather than separate logs.
The integration mode should match risk and timing. A control layer can intervene live, observe in sidecar mode, review data in batches, or run only in pre-production simulation. Enterprises should choose among them based on consequence, regulation, and acceptable latency rather than applying the heaviest control everywhere. In a cross-vendor environment, a unified evidence chain may be more urgent than immediately standardizing every model: who started the task, what the agent accessed, which tools it called, which policies constrained it, where a person intervened, and what business action resulted. That requirement exists even if the company never buys TrustWise.
Cost governance should begin before the first production bill. Saxena’s process starts by classifying the system with the Responsible AI Institute’s TrustX framework. It then simulates the workload and tunes chunk size, number of chunks, models, and endpoints before launch. After deployment, it looks for drift, over-control, and token waste. An enterprise can turn that into its own loop: classify by task and risk, stress-test real failure paths and tool-call chains, then put cost, latency, policy deviation, human escalation, and business outcome in the same operating view. This avoids both failure modes—too little control for risk approval, or so much control that the agent becomes slow, expensive, and unusable.
Labor comparisons need the same granularity. Saxena expects some agents’ annual operating cost to equal or exceed an employee’s annual cost. His proposed levers include different worker classes, dynamic model routing and selection, analytics that expose waste, and eventually agents that adjust their own model choices. He also says that last capability is not mature yet. A business case should therefore avoid one average token price for every agent. It should model task complexity, autonomy duration, tool-call volume, model mix, and human review separately.
Guardian and Genesis capabilities belong under different rules. Guardian functions can sit in the production gate: permissions, policy, cost, tone, approval paths, and auditability. Genesis functions fit exploratory work such as revenue-leakage analysis, fraud leads, and financial hypothesis generation, where outputs enter a validation process before action. “Beneficial hallucination” has value only when evidence review, causal testing, compliance judgment, and accountable business approval remain in place.
The cited 95% figure is not a benchmark that an enterprise should paste into a business plan. It is more useful as a prompt for internal measurement. What share of pilots reaches production? How many actions does a task actually trigger? Which loops cause token cost to run away? How much latency does runtime policy add? Does unified evidence measurably reduce human review and audit effort? An AI control tower becomes production infrastructure only when those answers hold inside the company’s own systems.
Source
More from WayDigital
Continue through other published articles from the same publisher.
Comments
0 public responses
All visitors can read comments. Sign in to join the discussion.
Log in to comment