← All Reports
Executive Summary
Here is the basic problem this whole industry exists to solve. A big company buys access to a frontier model — GPT-5, Claude, whatever — and discovers that a brilliant model sitting next to its messy data does approximately nothing. Someone has to wire the model into the company's actual systems, teach it the company's actual mess, point it at real work, and make it safe enough to trust. The companies that do that wiring are the "deployment layer," and they range from two-year-old startups to Accenture. This report maps them by strategy rather than one-by-one, scores 22 of them on a single matrix, and pokes hard at the one architectural argument the whole category is secretly having. The house rule throughout: every real claim gets a confidence label, and when the evidence is thin, we just say so — because in this category the honest answer is often "nobody actually knows," and pretending otherwise is how you lose money.
- The big architectural argument is real, but smaller than everyone thinks. The category loves to sort itself into "model the whole enterprise first" (Palantir) vs. "just automate one workflow first" (Distyl). Except Distyl now quietly ships a permanent, company-wide data layer of its own ("Context Mesh"), and Palantir now sells 1–5-day workshops that stand up a single workflow fast. Both sides are walking to the same place from opposite doors. The difference is increasingly build-order, not architecture.
- In the entire category, exactly ONE success story fully checks out. We went looking for cases where a public company's own audited filing both states a number and names the vendor that produced it. We found one: AIG → Palantir + Anthropic, disclosed in AIG's own annual report and investor day. Every other "proof point" either confirms the number but scrubs the vendor's name from the filing (GE/Palantir, T-Mobile/Distyl), or is just the vendor's own marketing. The fact that the count is one, and that a customer had to volunteer it, is the finding — this is a category that hides its own evidence.
- That viral “95% of AI pilots fail” number is basically vibes. Trace it back and it's a preliminary deck: 153 survey responses, self-reported, where “success” meant an executive said something nice about the tool. Not nothing, but not a statistic you'd bet on. The numbers that actually hold up are duller and more useful — ISG finds roughly 31% of prioritized use cases reach production and about 1 in 4 hit their ROI target, and Ramp's card-spend data shows enterprises quietly ~4×'d their AI spending year over year. People vote with budgets, not surveys.
- The consulting giants aren't getting killed — they're the store. The tempting story is that Accenture and McKinsey get disrupted into dust by AI-native startups. The evidence says the opposite: all six of them chose to partner with or buy the deployment layer rather than build it, and the money is real — Accenture booked $5.9B of GenAI work / $2.7B of advanced-AI revenue, IBM a >$12.5B cumulative book. They're the distribution channel, not the roadkill.
- Everyone, independently, built the same one part: a verification layer — the machinery that checks the AI's work before anyone acts on it. This is the thing that turns “the model probably got it right” into “you can bill for this,” and it's the closest thing the category has to a universal law.
Bottom line: the deployment layer is real, big, and growing fast — enterprises roughly tripled their generative-AI spend to about $37B in 2025 — but “what works” is far less provable than the sales decks suggest. The edges that look durable are unglamorous: the plumbing that lets an agent safely write back into systems of record, the years of integration wiring that's annoying to rebuild, and that verification layer. The sexier moats — outcome-based pricing, armies of embedded engineers — are narrower than the hype. And the one question every investor actually wants answered — does the ontology approach compound over time, or does the workflow approach just win by getting to production faster? — cannot be answered from public data today. That gap is the single most important thing to go diligence in this category.
1 · The Deployment Layer in One Screen
What these companies are, and why the category is opaque
Start with the thing everyone gets wrong: a frontier model is not a product. You cannot buy Claude, point it at your company, and get a result — any more than you can buy a brilliant new hire and get value on day one before they know where anything is. Someone has to plug the model into your systems, teach it your specific mess, aim it at real work, check that it did the work right, and figure out how to charge for it. The companies that do that — the “deployment layer” — are where an enterprise's AI budget actually turns into anything useful. That's the whole job.
Every player in this report, from a two-year-old startup to Accenture, runs the same core loop:
The core loop everyone runs
plug in the model → teach it your data → point agents at the work → deliver a result → bill for it. Every company here runs this same loop. What makes them different is which step they obsess over: some build a permanent map of the whole enterprise, some just automate one task at a time, some fly engineers out to sit with the customer, some bet the business on getting paid per result — and, crucially, each one has to build the part that double-checks the AI's work before it's safe enough to charge for.
Now, this category is genuinely hard to analyze, for three reasons — and those three reasons are why this report is written the way it is:
- Almost every “what works” claim is marketing. The case studies are written by the people selling the thing, and the most impressive numbers always seem to belong to private customers who never have to back them up. Convenient.
- Nobody publishes prices. Pricing, margins, unit economics — essentially all of it is inferred, never disclosed. Wherever we're guessing, we tell you we're guessing.
- The good evidence gets scrubbed. This is the strange one: even when a public company confirms in its own audited filing that AI saved it real money, the filing almost never names the vendor that did it (Section 7 is the whole autopsy). The proof exists; the fingerprints get wiped.
How to read the confidence labels
So here's the one habit that runs through the entire report: every real number gets a confidence label, so you always know how much weight it can bear. There are four:
Confirmed disclosed in a primary source (filing, official docs, company statement).
Reported credible secondary source, not independently verifiable.
Estimated / Inferred our analytical inference; no disclosure exists.
Consensus-but-Unproven widely repeated, weakly evidenced.
And when the only source is the vendor's own material, we call it Confirmed-as-claimed — meaning the company definitely said it, but nobody independent checked it. “They said so” and “it's true” are different things, and that difference is the whole game here.
2 · The Master Matrix
Twenty-two companies, five strategic axes
This is the map of the whole report — everything after it is just the walkthrough. Every company is scored on the five choices that define a deployment strategy. A note on who's here: OpenAI and Anthropic are now on the map, not just next to it. They used to be pure model suppliers — you rented the brains, someone else did the deploying. But in 2026 both stood up real deployment arms (OpenAI literally launched an "OpenAI Deployment Company" and bought a 150-person forward-deployed team; Anthropic built its "Applied AI" org and a $100M partner network of embedded engineers). So the arms dealers are now also fighting the war — which is a big enough deal that we score them and unpack it in the callout below. Scale, Mercor, and Turing are still left out on purpose; they're in the data-and-evaluation business next door, not the deployment business.
The five axes — and how to read the matrix
The big table below packs a lot into each cell, so here's the key. Each company is scored on five choices; these are the terms you'll see in the cells and what they actually mean:
| Axis | The strategic choice it captures |
| 1 · Models & weights | Whose brains does it use? Closed frontier = renting GPT/Claude/Gemini; open-weight = a model you can host yourself; agnostic router = swaps between them per task; owns models = trained its own. And whether it fine-tunes (custom-trains on a domain) or just uses models as-is. |
| 2 · Ontology ↔ Workflow | Does it model the whole enterprise once as shared, durable structure (ontology-first), or attack one workflow at a time (workflow-first)? The load-bearing axis (Section 3). |
| 3 · Orchestration | The machinery that runs the AI. Ingestion/semantic layer = getting the company's data in and organized; planner/executor = one model breaks the job into steps, others do them; verifier = the part that checks the AI's work before it counts; HITL (human-in-the-loop) = a person approves before anything ships; memory/state = what it remembers between runs. |
| 4 · Delivery | How you buy it. FDE (forward-deployed engineers) = the vendor sends engineers to build inside your company; product-led/self-serve = you just sign up and use it. Many land as services, expand into a platform. |
| 5 · Commercial | How you're billed. Outcome-based = pay per result (e.g. per resolved ticket); seat/subscription = per user; consumption = per usage (tokens/calls). And who's contractually on the hook for the result. |
| Company | 1 · Models & weights | 2 · Ontology↔Workflow | 3 · Orchestration | 4 · Delivery | 5 · Commercial |
| Palantir | Model-agnostic: open-weight (Llama/Mistral) first-class + BYOM self-host (vLLM/Ollama) for on-prem/air-gapped; frontier via catalog | Ontology-first (pole) | Ontology semantic layer + AIP orchestrator + verifier + HITL | FDE → platform/partner | License + consumption (value-framed) |
| Distyl | Multi-model: closed frontier (OpenAI+Anthropic via Azure) for hardest reasoning + open-weight Nemotron via NVIDIA NIM for always-on/long-running agents | Workflow-first → converging | Routines task-graph + Context Mesh + verifier | FDE-embed (purest) | Outcome-framed services → platform |
| Northslope | Gemini Enterprise + AIP-brokered frontier | Ontology-first (rented from Palantir) | AIP substrate + client app layer + HITL | FDE-embed (purest) | Services/project + substrate passthrough |
| Claritype | Frontier-orchestrated (unspecified) | Ontology / semantic-first (ex-Palantir DNA) | Semantic model + NL→governed-query + inspectable reasoning | Product-led self-serve | Seat / SaaS (not disclosed) |
| Sierra | Multi-model “constellation” (15+): frontier + open-weight + proprietary, routed per subtask | Workflow-first | Self-supervising agent + verifier + HITL | Product-led + SE | Pure pay-per-resolution[57] |
| Cognition (Devin) | Closed frontier (Anthropic-heavy) | Workflow-first | Planner/executor + self-test verifier + PR HITL | Product-led | Subscription + usage (ACU) |
| Happy Robot | 6-model pipeline: swappable LLM + proprietary voice models | Workflow-first (rule-guarded) | SIP/K8s voice stack + rule-verifier + human QA + HITL | Product-led + enterprise onboarding | Bundled platform + usage |
| Factory | Model-agnostic per-task router (Claude/GPT/Gemini) | Workflow-first | Orchestrator + role-scoped Droids + Review/Test verifier + HITL | Product-led, no seat minimums | Subscription / consumption |
| Decagon | Multi-model frontier | Workflow-first (AOPs) | Runtime + knowledge grounding + QA verifier + HITL | Product-led + SE | Per-resolution option (most pick per-conversation) |
| Harvey | Frontier + legal fine-tunes; ships own open-weight model (“Tenet”) | Workflow-first | Retrieval + task agents + citation verifier + heavy lawyer HITL | FDE program (services → platform) | Seat / subscription |
| Cresta | 20+ models: frontier + fine-tuned open-weight (Mistral base) + Ocean-1 CC foundation model | Workflow-first | Real-time transcription + agent runtime + QA verifier + HITL | Product-led + SE | Per-seat + usage ("outcome-LED") |
| Glean | Multi-model (Bedrock/Vertex/Azure OpenAI) | Ontology-leaning (Enterprise Graph) | Connectors + graph semantic layer + RAG + Agents + citation verifier | Product-led SaaS | Per-seat subscription |
| Cohere | Owns own models (Command/Embed/Rerank); open weights (Apache-2.0); private/air-gapped | Workflow-first (platform) | North RAG + agent orchestration + verifier, in-VPC | Model/platform vendor (no FDE) | Consumption (tokens) + platform license |
| 8090 | Frontier-orchestrated (agnostic) | Workflow-first (PDLC) | Software Factory pipeline + engineer-verifier | Services-led via EY (SI channel) | Project/services (outcome-framed, undisclosed) |
| OpenAI (lab arm) | Owns own models; single-vendor, not single-modality — deploys OpenAI models only, but that now spans closed frontier + own Apache-2.0 open-weight gpt-oss (on-prem-capable) | Workflow-first (use-case-led) | Own models + FDE build + evals + HITL | FDE-embed ("Deployment Company"; Tomoro ~150 FDEs) | Consumption (tokens) + enterprise + services |
| Anthropic (lab arm) | Owns own models (Claude), closed frontier only; ships no open weights — purest single-model deployco | Workflow-first (use-case-led) | Own models + Applied AI build + evals + HITL | FDE-embed ("Applied AI" + $100M partner network) | Consumption (tokens) + enterprise + services |
| Accenture | Multi-model; Palantir + OpenAI aligned + open-weight via AI Refinery/NVIDIA (custom Llama + Llama Nemotron NIM); on-prem/sovereign | Ontology-via-partner (Palantir group) | Accelerators + FDE model (absorbed via M&A) | Bodyshop → asset/accelerator | Labor leverage + accelerators |
| McKinsey (QuantumBlack) | OpenAI + Anthropic (ecosystem) | Advisory (n/a) | Lilli / Horizon / Kedro internal assets | Advisory-led, high-touch | Fees (judgment + seats) |
| Deloitte | Anthropic (Claude to ~470K); Palantir EOS; open-weight Llama Nemotron (NVIDIA stack) powers Zora AI, on-prem-capable | Ontology-via-partner (Palantir EOS) | Zora AI agentic + Omnia + Global Agentic Network | Bodyshop → agentic product | Labor leverage + emerging product |
| EY | Frontier via 8090 + Microsoft | Workflow (via 8090 PDLC) | Rents 8090's Software Factory | Bodyshop renting a factory | Fees + rented delivery |
| BCG X | OpenAI Elite + Anthropic | Workflow (bespoke build) | Build-and-scale org + ARTKIT red-teaming | Consulting-led build org | Fees (AI ≈ 25% of rev) |
| IBM Consulting | Owns Granite (Apache-2.0 open weights) + Llama/Mistral via watsonx; integrates OpenAI/Anthropic; on-prem/cloud | Workflow (watsonx-anchored) | watsonx.ai/.governance/Orchestrate + consulting | Services welded to owned stack | Software + consulting signings |
All matrix placements are evidence-backed in Sections 3–8. Cell confidence ranges from Confirmed (Palantir/Distyl architecture, Cohere model ownership, Factory routing, disclosed revenues) to Estimated (most pricing/delivery cells for private companies). Sources §.
The labs are now players, not just suppliers
Here's the wrinkle worth staring at. For two years the deal was simple: OpenAI and Anthropic sold the models, and a whole ecosystem of startups did the messy work of deploying them inside enterprises. The labs were the arms dealers; the deploycos were the soldiers. In 2026 that stopped being true.
OpenAI launched an actual “Deployment Company” — a services org with forward-deployed engineers — and bought
Tomoro (~150 FDEs) to staff it overnight.
[63] Reported Anthropic built “Applied AI,” its own embedded-engineer function, plus a
$100M partner network of Applied AI engineers and technical architects.
[64] Reported So the supplier is now walking straight into its customers' business. Why? Because the deployment layer is where the margin and the stickiness live — the model is a commodity you rent, but the deployment is a relationship you own. The strategic read for an investor: every deployco in this report now competes with its own most important vendor, and that vendor has infinite model access and a direct line to the same CIOs. It's the same move Distyl made going
up into ontology, run in reverse — the labs collapsing
down into deployment. The one hedge: the lab arms are single-
vendor by design (Anthropic's Applied AI deploys Claude exclusively and ships no open weights — the purest single-model case; OpenAI's Deployment Company deploys OpenAI models only, though since gpt-oss its portfolio now spans closed frontier
and its own Apache-2.0 open weights), while the independents stay model-agnostic and — as the matrix now reflects — most already run open-weight models (Nemotron, Llama, Mistral, Granite) somewhere in the production pipeline, especially for regulated/on-prem/air-gapped work. So “we're not locked to one lab's models” becomes a live selling point it wasn't a year ago.
The deployment layer, mapped on the two questions that decide the valuation
Two axes and a color, each a distinct question.
Left–right (X): what do you build first? Model the whole enterprise (ontology, right) or attack one workflow (workflow, left).
Up–down (Y): do you own your substrate or rent it? “Substrate” = the model + semantic layer + platform your customers build on. Own it (top) and you get platform economics and a platform multiple; rent someone else's (bottom) and you get app-layer margins plus dependency risk. This is the report's single clearest value-capture fork — which is why it's the vertical axis of the flagship chart, not a footnote.
Dot color: data gravity. ● Rose = capture (to work, your data moves into the vendor's platform — lock-in). ● Teal = sovereignty (it acts where your data already lives). ● Grey = mixed.
Placements are our analytical read Estimated-Inferred from public architecture evidence, scored on those single questions — not disclosed numbers.
Three things to notice. Palantir sits top-right and rose: it owns its substrate (platform economics) but it's a capture business — Foundry pulls your data into Foundry. That capture/lock-in is the whole rub, and the color now says so out loud while the height still credits the platform ownership. Northslope sits bottom-right: same ontology ambition as Palantir, but it rents Palantir's substrate — app-layer economics and dependency risk in one dot. The labs (OpenAI, Anthropic) sit top-left and rose: they own the deepest substrate there is (the model itself) but are workflow-led and capture-heavy — you send your data to them. The map's punchline: almost everything is rose. Genuine data sovereignty (teal — Cohere, Claritype) is rare, and that scarcity is itself a differentiator.
Prior art: market maps of this category usually plot vendors by stack layer (infra / orchestration / apps) or by category logos — see the [65] market-map index. The closest published framings to ours are a16z's argument that data/context gravity outweighs raw model quality[66] and vendor analyses organized around trust, flexibility, and lock-in.[67] We didn't find an existing map that crosses substrate ownership against ontology-vs-workflow and overlays data gravity for this specific deployco cohort — so treat this as our own synthesis, informed by those, not a restatement of a standard chart.
3 · The Load-Bearing Section
Ontology-first vs. workflow-first — and why the fork is quietly collapsing
This is the central architectural fork of the category and the deepest section of the report. Two opposite bets, same go-to-market motion: model the whole enterprise first (Palantir), or attack one workflow first (Distyl). The investor's prior was that these two might not actually be that different. We were mandated to disprove the fork before asserting it — and the disprove-first hunt landed the report's most important finding.
The money finding
The clean “whole-enterprise vs. one-workflow” split is now mostly a difference in
what you build first, not what you build. Distyl, supposedly the poster child for the one-workflow approach, has quietly gone and built a permanent, governed, company-wide data layer it calls
Context Mesh — and describes it in language almost identical to Palantir's famous moat.
[3] Meanwhile Palantir now sells
1–5-day workshops that stand up a single working workflow fast
[7] — which was supposed to be Distyl's whole edge. They started at opposite ends and walked toward each other. What's left between them is real but narrow: how deep and how safely each one can
write changes back into the systems that actually run the business, plus the years of unglamorous integration wiring that's a pain to rebuild.
Palantir (the ontology pole), mapped at architecture depth
Palantir's differentiation is architectural, not merely go-to-market. Foundry's Ontology is a largely deterministic, governed model of an enterprise — Palantir frames it as representing the enterprise's decisions, not merely its data.[1] It has two first-class halves:
- Semantic elements — object types, objects, properties, and typed link types (Palantir insists links be "semantically meaningful," not shared foreign keys). Interfaces provide polymorphism so "a single workflow covers all implementing types."[1] Confirmed
- Kinetic elements — Action types (governed writes committed as a single transaction) and Functions (business logic, computation, model calls, ontology edits).[2] Confirmed
Here's how the AI actually touches all this, in plain terms. It never gets handed the company's raw database. Instead it works through five guardrails: (1) only the relevant facts get pulled into its view on each message, never the whole schema; (2) it can walk the connections between objects — Order → Shipment → Carrier → hub — which ordinary keyword search simply can't do; (3) it can run the company's actual business logic; (4) when it wants to change something, that change goes through a controlled, optionally human-approved “Action” rather than a raw database write; and (5) the whole thing is governed, so the agent only ever acts through the sanctioned boundary.[5] Palantir's own phrase for the philosophy: “determinism where possible, autonomy where necessary, constraint everywhere.” The key point for an investor: the valuable part is not the model — Palantir just rents whichever model is best — it's this governed layer that turns the reckless version (“hand the AI your database password”) into the safe version (“let the AI act only through a controlled interface”).[5] Estimated-Inferred
Distyl (the workflow pole), and its 2026 convergence
In June our verified map of Distyl was the pure workflow-first archetype: a Standard Operating Procedure auto-decomposed into "Routines" — discrete, auditable tasks with logged inputs/outputs/tool-calls — a thin, local semantic layer per workflow.[16] That is still true at the execution layer. But Distyl's own 2026 technology pages describe a four-layer platform whose foundation is now a persistent semantic layer Confirmed-as-claimed:
- Context Mesh — "an AI-native data foundation. It reads from your existing data systems… and writes back what every agent decides. Your data stays where it is. No migration required." "When an agent makes a decision, that output flows back into the mesh, semantically indexed… Each new use case inherits everything Context Mesh already knows." Agents are "first-class citizens" with "explicit, auditable permissions."[3]
- Context Views — explicitly "the semantic layer… a foundation of human-understandable objects: business entities, relationships, and constraints… your agents can reason from." SME "Directed Review" edits become "a reusable update to the schema itself."[4]
Distyl's own NVIDIA-integration post describes moving "past the concept of the flat, static retrieval index" to build "a living, traversable graph… encompass[ing] episodic memory, tacit and explicit knowledge, policies, workflows, and domain logic," where agents "write back state that persists long after the session ends."[6] A third-party analysis states it plainly: "Distyl is applying [Palantir's ontology] thesis natively to LLM agents."[17] Reported
Head-to-head: same destination, different construction method
| Dimension | Palantir Foundry/AIP (ontology-first) | Distyl Distillery 2026 (workflow-first → converging) |
| Persistent semantic layer? | YES — Ontology (objects/props/links) | YES — Context Mesh + Context Views |
| Typed entities + relationships | Object types / link types (hard schema) | Nodes/edges: "typed relationships" Rep |
| Governed write-back to the layer | Action types (transactional, audited) | "writes back what every agent decides" as-claimed |
| Multi-hop traversal by agents | Object Query Tool link traversal | "living, traversable graph," multi-hop as-claimed |
| Data stays in place | Virtual tables enable in-place | "No migration required" as-claimed |
| How the model is BUILT | Engineers design schema up front (deterministic) | SMEs + agents accrete it; "Directed Review" edits schema |
| Compounding mechanism | Interfaces + canonical objects → workflows inherit | "Each use case inherits everything Context Mesh knows" |
| Maturity of the layer | ~decade of production hardening | New (2026 re-architecture); unproven at Palantir depth |
Disprove-first: the null hypothesis is strongly supported
Evidence Distyl is moving toward Palantir: Context Mesh is a write-back knowledge graph; Context Views is an explicit "semantic layer"; the ex-Palantir founders (Arjun Prakash + Derek Ho) make the ontology DNA intentional.[3][4] Evidence Palantir is moving toward Distyl: the AIP Bootcamp promises "zero to use case in 5 days or less," with "at least one production-ready workflow by day five" — Palantir attacking the fast per-workflow motion.[7] Confirmed-as-claimed
Verdict on the fork
The fork survives only as (1) a sequencing difference (build-model-first vs. ship-first-then-accrete) and (2) a schema-determinism difference (a designed, strongly-typed object model where agents cannot invent objects, vs. a mesh partly accreted from agent decisions). The genuine, surviving daylight is the depth of the deterministic action/write-back layer into systems of record — Palantir's atomic, staged, governed writes into ERP/WMS/edge over a decade of connectors are far deeper than anything Distyl publicly documents (Distyl writes back into its own mesh, not documented ERP-grade write-back). That is a head-start moat, not a categorically different architecture. As an architectural claim, "ontology-first vs. workflow-first" no longer describes two different architectures — it describes two on-ramps to the same highway. That itself is the headline.
The compounding-vs-ceiling stress-test
The thesis under test: workflow-first converts to production faster but ceilings into 40 disconnected point-solutions; ontology-first is slower-to-value but compounds because every new workflow inherits the object model.
- For compounding: Palantir documents the mechanism (interfaces + canonical objects → workflow reuse).[8] And the strongest single data point in this leg: Distyl's own pivot is revealed-preference evidence the ceiling is real. A workflow-first company at a ~$1.8B valuation chose to spend engineering building Context Mesh — whose entire pitch is escaping "add a second agent, then a tenth, and each operates from its own blind spot… No compounding. Just isolated automation, multiplied." Distyl is describing the exact ceiling the thesis predicts, and building to escape it.[3]
- Against / complicating: "Slower-to-value" for ontology-first is now weakly false — a first workflow ships fast on either pole (AIP Bootcamp, ≤5 days). And Palantir itself warns compounding only holds if teams converge on shared concepts; duplicated object types can accrete redundancy instead of value.[8] Ontology-first does not automatically compound — it compounds if governed well.
The core diligence gap — stated plainly
The "ceilings out into disconnected point-solutions" outcome is not observable in public data for Distyl or any workflow-first peer. No public deployment census shows point-solution sprawl vs. compounding; no audited before/after exists. The stress-test is architecturally sound and corroborated by Distyl's revealed behavior, but empirically unproven. Production-conversion rates are unknown for both poles. Any investment thesis that rests on "ontology-first compounds and workflow-first ceilings" is, today, resting on architecture and inference — not measured outcomes. Consensus-but-Unproven
4 · Models & Weights
Open vs. closed, by use case — and the one company that owns its models
The pattern here is refreshingly clear for once: 13 of the 14 independents just use the big frontier models as-is — they don't train their own. The deployment layer sits on top of the model, not inside it. So the only real choices are which models, how easily you can swap between them, and whether you bother fine-tuning one for a specific domain.
- Model-agnostic routing is becoming table stakes. Factory explicitly routes Claude/GPT/Gemini per task to avoid lock-in and control cost[24] Confirmed; Glean routes across Bedrock/Vertex/Azure OpenAI[28]; Palantir and Sierra broker multiple providers.
- Closed frontier dominates the reasoning core; open-weight appears where the environment demands it — regulated, on-prem, air-gapped, or cost-sensitive. Cohere's entire pitch is private deployment ("GPU in a closet")[29]; Distyl runs open Nemotron via NVIDIA NIM alongside closed frontier for its deterministic pipeline.[6]
- Cohere is the lone model owner — and it ships open weights. It trains its own Command/Embed/Rerank family and releases open weights under Apache-2.0 for self-host, selling private, in-VPC, “GPU-in-a-closet” deployment — a fundamentally different business (model vendor, not deployment layer) that happens to sit adjacent.[29] Confirmed
- Domain fine-tuning is the exception, concentrated where accuracy is legally load-bearing: Harvey layers legal case-law fine-tunes and has shipped its own open-weight post-trained model (“Tenet”)[26], Cresta runs contact-center models fine-tuned on client conversation data[27], and Happy Robot [60] builds proprietary in-house voice/end-of-turn models for latency.[22] Reported The pattern is instructive: owned/open weights appear in exactly two places — regulated/air-gapped (Cohere) and domain adaptation where accuracy is legally load-bearing (Harvey’s Tenet, Cresta) — never in the generic orchestration core.
The single most-cited enterprise-spend datapoint of 2026: Anthropic leads enterprise LLM spend at 40%, having "unseated OpenAI" (Menlo Ventures).[30] Reported — a market-model estimate from a VC with portfolio exposure, directionally strong but not audited. Distyl, Cognition, and Harvey all skew Anthropic-heavy in their orchestration, consistent with that leadership.
So what
Since nearly everyone is renting the same few models, the model itself is not the moat — picking one is now basically an automated routing decision, like choosing the cheapest flight. Whatever defensibility exists lives in the stuff around the model: how you feed it the company's data, how you orchestrate it, how you check its work, and how you deliver. Which is exactly why the Section 3 argument about ontology matters more than “which model do you use.”
5 · Orchestration Architecture
Where designs converge — and the one component everyone builds
Strip the marketing away and the orchestration stacks rhyme: ingestion/connectors → a semantic or context layer → a planner/executor split (an orchestrator model directing task models) → tool and action calls → a verification layer → human-in-the-loop checkpoints → memory/state. The interesting question is where they converge vs. genuinely differ.
The universal building block: the verification layer
Category-wide convergence
Everyone, separately, built a “verifier” — the part that checks the AI's work — and it's what makes the output sellable in the first place. The flavors differ by industry: answer-quality checks for customer service (Sierra, Decagon), code review and test bots for engineering (Factory), citation-grounding for law (Harvey), hard rule-guards for freight quotes (Happy Robot), show-your-work reasoning for analytics (Claritype), and the governed “Action” system for Palantir. But it's the same idea every time, and it's the load-bearing commercial part of the whole category: without it, a customer has no reason to trust — or pay for — what the agent produced.
The academic literature corroborates why this layer is non-negotiable: single-run benchmark success does not predict production reliability, and agent failures are frequently silent — hallucination, misinterpreted instructions, and reasoning drift that propagate without throwing any runtime exception (MAS-FIRE; ReliabilityBench; IBM's CUGA deployment paper).[34][33][35] Reported A verifier is the only thing standing between "the demo worked" and "the deployment silently corrupted a workflow."
Where designs genuinely differ
- The semantic layer — persistent enterprise ontology (Palantir, Distyl-2026, Claritype [56]) vs. a knowledge-graph substrate that grounds retrieval (Glean's Enterprise Graph) vs. thin per-workflow context (CX/coding agents). This is the Section-3 fork.
- Determinism placement — Happy Robot puts deterministic rule-guards on high-stakes outputs (a freight rate, a booking number) so the LLM cannot bypass validation[22]; Palantir puts determinism in the Action transaction; Distyl runs a deterministic deep-research pipeline where "the code lays the tracks, the model reasons over results."[6]
- Substrate ownership — own it (Palantir, Distyl), rent it (Northslope [59] rents Palantir's AIP/Foundry), or skip it (pure per-workflow agents).
The planner/executor pattern (an orchestrator model decomposing work and directing task models) is now near-universal in the agentic players — Factory's coordinator assigning role-scoped Droids[25], Cognition's long-horizon task decomposition[18], Palantir's AIP Logic. It is converging into a commodity pattern; the differentiation has moved to the context layer beneath it and the verification layer above it.
6 · Delivery & Commercial Models
FDEs are dying at the edges; outcome pricing is rare and CX-shaped
Delivery: forward-deployed engineering survives only where the substrate is complex
Forward-deployed engineering — the Palantir-invented playbook that started this category — is expensive, and across the 14 independents only four still really run it: Distyl, Northslope, Harvey (literally described as “a services business with software margins”), and 8090 (delivered through EY).[26] Reported Everywhere the job is cleaner and more self-contained — writing code (Factory, Cognition), search (Glean), support tickets (Decagon, Sierra), voice (Happy Robot) — companies skipped the embedded-engineer model and just shipped a product you can buy yourself. Factory sells off the shelf with no seat minimums[24]; Glean lands on seats and grows into agents.[28]
The pattern
Whether you need embedded engineers isn't about philosophy — it's about how messy the job is. Rewiring a whole enterprise across siloed, filthy data needs humans on-site. Resolving a support ticket or shipping a code change doesn't; you can package that as software. And as the models get better, more jobs move from the first bucket to the second, so the embedded-engineer model keeps retreating to the genuinely hard, cross-system builds. That, by the way, is also why the consulting giants (Section 8) are a channel and not roadkill — they already own the relationships for exactly the big, complex, high-touch deployments where embedding still wins.
Commercial: outcome pricing is rarer than the narrative claims
“We only get paid when you get a result” is the category's favorite line. In practice, genuine pay-for-outcome billing is rare, and almost entirely a customer-service thing. Only Sierra truly bills per-resolution. Decagon [58] offers it as an option but says most customers still pick per-conversation because it's predictable. Cresta talks “outcome-led” but actually invoices per-seat plus usage (its own rate card: 125K chats/yr for $150K).[27] Confirmed (rate card) Coding, legal, and search bill the old-fashioned way — seats and usage.
The reason is simple once you see it: to bill per outcome, you need an outcome you can count — a “resolved ticket.” That clean unit exists in customer service and almost nowhere else. You can't invoice “per transformed supply chain.” So wherever you hear “own the outcome” outside of CX (Distyl, 8090), it's marketing language wrapped around what is really a services contract, not a meter that ticks per result. Estimated-Inferred — no rate cards exist; the pricing structure is our inference.
Where "own the outcome" becomes a liability
Contractually owning the outcome transfers risk to the vendor. In CX with a countable resolution and a verifier, that risk is priceable. Outside CX — where the "outcome" is a multi-quarter transformation with confounded attribution — owning it means owning blame for results the vendor only partly controls. The companies loudest about outcomes (Sierra, Decagon) are precisely those where the outcome is measurable; the ones doing complex builds quietly price on services. That asymmetry is a tell about where outcome pricing actually works.
7 · What Actually Works — The Evidence
The hard data: mostly, it does not exist — and that is the finding
Everyone wants the same thing here: hard proof of what actually converts a pilot into production and moves the P&L — not another testimonial. We looked three ways: at what the big deployment aggregators show across thousands of projects, at the customer-side filings behind specific wins, and at the academic literature. The honest headline: the gold-standard proof barely exists in public, and the statistic everyone quotes is the flimsiest one in the pile. But the aggregate data is real and it does say something — so start there.
7.0 — What the aggregators actually show, at scale
Rather than cherry-pick a few flattering case studies, start with the databases that have already collected thousands of real deployments and asked “what shipped, and what held up?” Three are worth knowing, and they agree with each other:
- ZenML's LLMOps database — ~1,200 production deployments catalogued through 2025 — is the most useful because it's specifically about what survives contact with production. Its recurring lesson is almost blunt: simple, narrow workflows beat flashy general-purpose agents, and the systems that work pair retrieval with real evaluation, monitoring, and a human in the loop rather than clever prompting. In their words, “context engineering over prompt engineering” — getting the right data in front of the model matters more than the model.[61] Reported This is the same conclusion the rest of this report reached from the architecture side, arrived at independently from 1,200 shipped systems.
- Evidently AI's database — ~800 production ML/LLM case studies across 150+ companies — is the cleanest “what's actually running” catalog; it deliberately excludes demos and pilots.[55] Reported Useful for texture (which patterns recur, in which industries), useless as a success rate — it only contains the winners.
- AIUseCaseHub — ~3,814 public deployments across ~3,241 companies, filterable by maturity (Production / Scaled Production) and “evidence strength.”[62] Reported Its own fine print is the tell: it warns that its evidence signals describe “what the source supports, not an independent audit that a system is still operational.” Even the aggregator built to track this stuff won't vouch that the deployments still work.
The catch with all of them
These databases are enormous and genuinely useful for one question — what does a shipped system look like? — and useless for the question everyone actually asks — what's the success rate? They're built from published deployments, which means they're a list of survivors with no graveyard next to it: a big numerator and no denominator. And crucially, they're indexed by model (Copilot, Bedrock, Claude), almost never by deployco — so they tell you AI is being used, not whose deployment approach won. To get to “who actually moved the number,” you still have to do the filing-by-filing dig below. Which is where it gets bleak.
Where AI actually ships — deployments cluster in a handful of jobs
Pool the aggregator catalogs and one thing jumps out: production AI isn't spread evenly across the enterprise — it piles up in a few well-defined jobs (customer support, writing code, search/knowledge) and thins fast after that. This is the “narrow workflows win” thesis as a picture: the cleaner and more countable the task, the more real deployments exist. Bars show the rough share of catalogued production deployments by function, pooled and normalized across the ZenML and Evidently corpora.[61][55] Estimated-Illustrative — aggregators use different taxonomies; treat as directional shape, not precise percentages.
The enterprise-AI funnel: the money is huge, the throughput is not
The single hardest fact about this category, in one chart. Enterprises are spending enormous money — ~$37B on GenAI in 2025, roughly triple the prior year.[48] But that money narrows sharply as it moves toward results: the widely-quoted (and weak) MIT/NANDA line puts pilot success in the single digits[31], ISG's survey work puts production conversion around ~31% with roughly 1-in-4 reaching clear ROI.[32] Different studies, different denominators — so read the shape, not the exact widths: lots of spend in, not much proven value out. Estimated-Illustrative — stages come from different sources and are not a single measured cohort.
7.1 — Named contract outcomes: exactly one clean triangulation (AIG), and it took the customer choosing to disclose it
So the aggregators can't tell you who won. Can anyone? The gold standard would be a public company saying, in its own audited filing, “this specific vendor produced this specific result.” That beats any case study, because the customer has to answer to auditors. We took the deployco case studies that named public customers and went looking for the number in the customer's own annual reports, 10-Ks, investor days, and earnings calls. Nine worked examples. Here's how many survived:
Grading vocabulary: Confirmed-Triangulated customer's own disclosure names the deployco AND a quantified outcome. Partial customer corroborates the outcome magnitude but not the vendor name (or vice-versa). Vendor-Claimed deployco-published only, no customer-side corroboration. A separate class, Logo-Confirmed / Outcome-Unattributable, is where a customer confirms the relationship (logo wall, press release) but no number attaches to it.
| Enterprise | Deployco | Claimed result | Grade | Customer-side check |
| AIG (NYSE:AIG) | Palantir + Anthropic | Lexington 370K+ submissions (~+26% YoY); submit-to-bind ~+35–40%; underwriting cycle weeks→<1 day; 100% of private/non-profit submissions reviewed without adding underwriters; target $4B+ new premium by 2030 | Confirmed-Triangulated | AIG's own 2025 annual report / investor day / earnings disclose the outcomes and name Palantir + Anthropic as the partners[40] |
| GE Aerospace (NYSE:GE) | Palantir (AIP) | "AIP to help increase engine production by 26%" | Partial | 10-K corroborates deliveries +26%, LEAP +28%, names AI tools — but 0 EDGAR hits for "Palantir"[37] |
| T-Mobile (NASDAQ:TMUS) | Distyl | ~$3B AI savings by 2027; ~60% chatbot containment | Partial | T-Mobile's own Q4-2025 CMD calls corroborate the magnitude but never name Distyl — credits a partner “with OpenAI”[40] |
| Assoc. Materials (private) | Palantir | OTIF 40%→90% in 9 months | Vendor-Claimed | Private — no filing to triangulate[38] |
| Goldman Sachs (NYSE:GS) | Cognition (Devin) | "Hundreds of Devins, scaling to thousands" | Vendor-Claimed | Exec press quote, no P&L number; 0 EDGAR hits for "Devin"[39] |
| Cigna (NYSE:CI) | Sierra | Auth time −80%, live in 8 weeks | Vendor-Claimed | Cigna discloses $200M AI savings but never names Sierra; different metric[41] |
| Rippling (private) | Decagon | Deflection 38%→>50% | Vendor-Claimed | Private — no filing[42] |
| Cox (private) | Cresta | Revenue per chat +20–30% | Vendor-Claimed | Private — EDGAR "Cresta" hits are pre-2020 false positives[42] |
| PwC UK (private) | Harvey | Saves 1–2 hrs per liquidation | Vendor-Claimed | Private partnership — no filing[43] |
| Duolingo (NASDAQ:DUOL) | Decagon | 80% chat deflection | Vendor-Claimed | 10-K books generic "AI costs," no vendor, no metric[44] |
Tally: 1 CONFIRMED-TRIANGULATED (AIG) · 2 PARTIAL (GE, T-Mobile) · 6 VENDOR-CLAIMED-UNCORROBORATED. The one clean case is AIG → Palantir + Anthropic — the most valuable page in this report. AIG's own 2025 annual report discloses the measured outcomes (Lexington 370,000+ submissions, up ~26% YoY; submit-to-bind up ~35–40%; claims cycle times from hours to minutes; figures REPORTED-from-primary — relayed from AIG-hosted primary docs, as the AIG PDF fetched as raw bytes rather than quoted verbatim), and AIG's own investor day and earnings call name Palantir and Anthropic as the partners and describe reviewing 100% of private/non-profit submissions without adding underwriters, targeting $4B+ new premium by 2030.[40] Two independent AIG-authored channels close the triangle.
The one honest nuance on AIG
Even the category's single clean triangulation has a seam: the metrics live in the annual-report layer; the vendor attribution lives in the investor-day / earnings layer. They are both AIG-authored, but a maximally adversarial reader can argue the specific number is not audit-tied to the vendor within one document. That caveat is itself the point — this is as good as customer-side outcome evidence gets in the entire category, and it still isn't a single audited line crediting a vendor with a number. Numbers also drift across AIG's own reporting layers (e.g., +26%/+35% vs. +30%/+40% quoted/binding in different disclosures), so we cite the range, not a false-precision figure.
The two partials show the more common failure. GE Aerospace's FY2025 10-K independently reports deliveries up 26% and discloses concrete AI-productivity tools (an AI blade-inspection tool halving inspection time; an AI material assistant cutting MRO turnaround) — but a full-text EDGAR search for "Palantir" across all GE filings returns zero hits.[36][37] T-Mobile's own Q4-2025 calls corroborate the ~$3B-savings magnitude of Distyl's claim but never name Distyl (crediting a partner “with OpenAI”).[40] Each confirms an AI-driven gain; each strips or mismatches the vendor.
Why the evidence is structurally missing — and why AIG is the exception
The one clean triangulation exists only because AIG itself chose to disclose it. The category's outcome evidence is otherwise systematically non-triangulable, for three compounding reasons: (1) vendors are laundered out of audited filings — auditors and IR will not credit one software vendor for a headline operating metric, so even real gains lose the attribution (GE/Palantir). (2) The crispest vendor numbers cluster on private buyers (Associated Materials, Rippling, Cox, PwC) with no filing obligation — the loudest claims are the least checkable. (3) Public buyers disclose the benefit but not the vendor, and often a different number (Cigna quantifies $200M but never names Sierra; T-Mobile credits “OpenAI” not Distyl). The customer and the vendor each hold one edge of the triangle and almost never publish the third — AIG is the rare buyer that published all three.
Method note — the logo wall vs. the outcome
A deployco “logo wall” proves a relationship, not a result. We separate Logo-Confirmed / Outcome-Unattributable — the customer is real and confirmed (press release, named case study, marketplace listing) but no quantified P&L number is customer-attributed — from genuine outcome triangulation. Most of the category's “proof” is logo-confirmed-but-outcome-unattributable; treating a logo as an outcome is the single most common diligence error in reading these decks.
Investment implication: any deployco pitch resting on "named enterprise + quantified outcome" is, evidentially, vendor-attested marketing — even when the enterprise is real and public. Diligence should treat these as leads, then (a) check EDGAR full-text for the deployco's name in the customer's filings (usually 0 hits) and (b) look for whether the customer independently quantifies any aligned metric. The GE-style partial is the ceiling of available evidence, not the floor. The base rate confirms it: ~90% of the S&P 500 mention AI in 10-Ks[45], yet only ~12% of CEOs cite both cost and revenue benefit, and 56% report no significant financial benefit.[46] Reported
7.2 — Pilot→production: grading the "95%," not parroting it
You've seen the “95% of AI pilots fail” line — it was everywhere. It comes from one source, MIT/NANDA's The GenAI Divide: State of AI in Business 2025, and we read all 26 pages so you don't have to. Here's what's actually behind the number: a look at “over 300 publicly disclosed AI initiatives,” 52 interviews, and a survey of “153 senior leaders collected across four major industry conferences.”[31] That is a conference-hallway survey, not a census — and it gets shakier from there:
| Dimension | Finding | Grade |
| Sample | n=153 survey (conference-recruited) + 52 interviews; 300 initiatives are publicly disclosed (selection-biased) | Thin |
| Self-reported vs. measured | Self-reported; the report's own note calls figures "directionally accurate based on individual interviews rather than official company reporting" | Weak |
| Definition of "success" | Verbatim: a tool "users or executives have remarked as causing a marked and sustained… P&L impact." Success = a subjective remark, not an audited outcome | Soft |
| Window / review | 6-month window; labeled "Preliminary Findings," v0.1; not peer-reviewed | Not peer-reviewed |
Crucially, the 95% refers specifically to custom/task-specific enterprise GenAI tools; generic chatbots (ChatGPT/Copilot) are reported at ~83% adoption and explicitly excluded.[31] Autopsy verdict: a defensible qualitative signal that custom GenAI struggles to reach durable production — but it does NOT support precise quantitative claims. The number's virality outran its rigor. Consensus-but-Unproven Cite the pattern; never the precision.
The defensible benchmarks come from higher-rigor sources, all verified against primary pages:
- ISG State of Enterprise AI Adoption 2025 (n=1,200 defined use cases): 31% reached full production (double 2024); ~$1.3M average spend to date; 1 in 4 initiatives achieving expected ROI on growth; 50% hitting expected efficiency gains — all four figures confirmed verbatim.[32] Confirmed The most defensible public benchmark of pilot→production conversion.
- Ramp AI Index — corporate-card transaction data from 50,000+ businesses (measured spend, not survey): ~4× monthly AI spend growth Feb 2025 → Feb 2026. Ramp underestimates adoption (misses free tools) — the honest direction of bias.[48] Reported The highest-rigor source on the spend question.
- Menlo Ventures — enterprise GenAI spend $37B in 2025, 3.2× YoY; 76% of use cases now purchased vs. built; buyers convert at 47% to production vs. 25% for traditional SaaS.[30] Reported (VC-modeled — flag the book.)
Reconciling the numbers
MIT (single-digit% for hard custom builds) and ISG (31% of prioritized use cases) are not contradictory — different denominators. The truth is bracketed:
~single-digit% for hard custom builds, ~30% for the broader prioritized portfolio. And two independent sources agree on the one robust best-practice:
buy-over-build converts faster (Menlo 47% vs. 25%; MIT ~2× for partnerships). Everything else labeled "best practice" — learning/memory systems, mid-market speed — is consensus, not measured. Public production-case libraries (e.g. Evidently AI's ~805 cases across 185 companies
[55]) are useful for "what shipped looks like" but are survivorship-biased numerators with no denominator — they cannot ground a base rate.
7.3 — The academic corpus: the failure mechanism is real and named
The Valency academic corpus returned 11 on-thesis, real arXiv papers. The consensus is consistent and useful: (1) single-run benchmark success ≠ production reliability (ReliabilityBench introduces consistency/robustness/fault-tolerance metrics)[33]; (2) agent/RAG failures are frequently silent (MAS-FIRE: semantic failures propagate without exceptions)[34]; (3) enterprise-specific gaps recur — domain process knowledge, evolving requirements, role-based data access. IBM's own deployment paper is the most on-thesis case study: "evidence of [generalist agents'] use in production enterprise settings remains limited" — benchmark strength does not equal deployed business value.[35] Reported (arXiv preprints — verify ids before external citation.)
The honest bottom line on evidence
No rigorous, peer-reviewed, population-representative study of pilot→production failure exists publicly. The most-cited stat is the weakest source; the strongest sources measure spend (real and exploding) rather than success. The money is unambiguous — $37B, 3.2× YoY, ~4× in card data. Whether that money is working is genuinely opaque, bracketed between single-digit and ~30% production conversion depending on how you count. Anyone claiming precise certainty about "what works" in this category is selling something.
8 · The Incumbents
Threat, channel, or roadkill? Mostly channel.
Our June writeup dismissed the consulting firms in two sentences. That was a mistake. Accenture, McKinsey (QuantumBlack), Deloitte, EY, BCG X, and IBM Consulting are the biggest, best-funded sales channels into exactly the giant, cautious enterprise buyers that a startup can't reach on its own. And the important fact is what they did: every one of them chose to partner with or buy the deployment layer rather than try to out-build it. When the incumbents decide to resell you instead of copy you, that tells you where the value is.
| Firm | Delivery model | Disclosed GenAI $ (labeled) | Key startup/platform bridge | Verdict |
| Accenture | Bodyshop → asset/accelerator; absorbing FDE model via M&A | GenAI bookings $5.9B FY25; advanced-AI rev $2.7B Confirmed | Palantir Business Group (>2,000 certified); acquired RANGR, Decho | ChannelThreat-that-absorbs |
| McKinsey (QuantumBlack) | Advisory-led, high-touch; strong internal assets (Lilli, Kedro) | Not separately disclosed n/a | OpenAI Frontier Alliance; Anthropic | Channel (upstream) |
| Deloitte | Bodyshop → agentic product (Zora AI, Omnia) | Not disclosed (total $70.5B FY25) n/a | Anthropic (Claude to ~470K staff — largest deployment); Palantir EOS | ThreatChannel |
| EY | Bodyshop renting a startup's factory | Not disclosed n/a | 8090 (founding partner, EY.ai PDLC); Microsoft >$1B/5yr | Channel (dependent) |
| BCG X | Consulting-led build-and-scale; real product org | ~25% of 2025 rev ≈ ~$3.6B AI-related (not GenAI-specific) Reported [54] | OpenAI Elite/Frontier; Anthropic (since 2023) | ChannelThreat |
| IBM Consulting | Services welded to owned stack (watsonx) | GenAI book >$12.5B cumulative Q4'25; ~$1.5B GenAI consulting Confirmed | Owns watsonx; integrates OpenAI/Anthropic | Channel (captive)Roadkill in open GenAI |
The tell is in what they actually did. EY rents 8090's software factory rather than building one[52]; Deloitte put Claude in front of ~470,000 of its own people (Anthropic's largest deployment anywhere) and wraps Palantir's platform[51]; McKinsey and BCG joined OpenAI's Frontier Alliance[50]; Accenture stood up a 2,000-plus-person Palantir unit and simply bought two embedded-engineering shops (RANGR, Decho).[49] Every one of these is a firm admitting it can't out-build the startups — so it resells them instead.
The disclosed numbers frame the stakes. Accenture's $5.9B GenAI bookings and $2.7B advanced-AI revenue[49] Confirmed and IBM's >$12.5B cumulative GenAI book[53] Confirmed prove demand is real and large — but both are single-digit shares of firm revenue and heavily bookings-weighted (i.e., early). EY, McKinsey, Deloitte, and BCG are private and do not disclose GenAI-specific revenue; BCG's ~$3.6B is AI-related (a broader bucket than GenAI-specific), so it is not comparable to Accenture's number.
Overall verdict
For a startup or an investor, the consulting giants are mostly a CHANNEL — not roadkill, and only a conditional threat. What they're good at: scale, the trust to get through procurement and legal, managing the human change inside a big company, C-suite access, and — for Accenture — a balance sheet big enough to just buy anyone who threatens them. What they're structurally bad at: their whole business is billing for hours, not selling software at software margins, so they lose on speed, product depth, and increasingly on talent (the best deployment engineers keep leaving for labs and startups). Net: treat Accenture/Deloitte/BCG X as frenemies who both distribute you and compete for the build, EY/McKinsey as pure channels, IBM as locked inside its own stack. The threat only bites a startup that stays small long enough to get absorbed. The channel opportunity is available right now.
9 · Similarities vs. Differences — The Distillation
Where the category converges, where the real forks are, and what remains unknown
Having scored 22 companies across five axes and hunted hard for evidence, the payoff synthesis:
Where the whole category genuinely converges
- The model is not the moat. 13 of 14 independents orchestrate the same frontier models; routing is increasingly automated. Differentiation lives above and below the model.
- Everyone has a verifier. The verification/eval layer is the universal building block — the mechanism that makes probabilistic output billable. It is the most convergent and most load-bearing component in the category.
- The planner/executor pattern is commoditizing. Orchestrator-directs-task-model is now table stakes in the agentic players.
- "Own the outcome" is mostly a framing, not a metered invoice — except in CX, where a resolution is countable.
Where the real strategic forks are
- Depth of the deterministic write-back substrate (the surviving core of the ontology fork). Owning a governed, transactional, systems-of-record write layer is the deepest and hardest-to-copy moat — and the accumulated integration surface behind it is a genuine head-start advantage. This is where Palantir's real daylight over Distyl remains.
- Substrate ownership vs. rental. Own it (Palantir, Distyl) → platform margins and platform valuation. Rent it (Northslope on Palantir) → app-layer margins and dependency risk. This is the single clearest value-capture fork.
- FDE where the substrate is complex; product-led where it is legible. A delivery-model choice dictated by problem structure, not preference — and the frontier of "complex enough to need FDEs" retreats as models improve.
- Outcome pricing where the unit is countable. A durable edge in CX; structurally unavailable for complex transformation.
What the evidence (weakly or strongly) suggests works vs. fails
| Claim | Verdict | Strength |
| Buy-over-build converts to production faster | Works | Reported — two independent sources (Menlo, MIT) |
| A verification layer is required for billable output | Works | Reported — universal in practice + academic corpus |
| Deterministic substrate compounds; workflow-first ceilings | Plausible | Consensus-Unproven — architecture + Distyl's revealed pivot, no audited proof |
| FDE-embed drives production conversion | Conditional | Inferred — true where substrate is complex; unproven as a universal |
| Outcome pricing is a superior model | Only in CX | Reported — requires a countable unit |
| "95% of pilots fail" | Directional only | Consensus-Unproven — low-rigor source; pattern real, precision false |
What remains genuinely unknown — and worth stating plainly
The honest limits of this analysis
- Whether ontology-first actually out-compounds workflow-first in production is not knowable from public data. Production-conversion and compounding rates are undisclosed for both poles. This is the single most important diligence gap in the category.
- Exactly one clean customer-triangulated outcome exists (AIG → Palantir + Anthropic), and even it has a seam — metrics and vendor-attribution sit in different AIG documents. Every other quantified claim is vendor-attested or attaches to a private buyer. Outside AIG, the category's outcome evidence is systematically non-triangulable.
- Unit economics are entirely inferred. No rate cards; all pricing/margin structure is analytical inference, flagged throughout.
- Distyl's Context Mesh depth is vendor-asserted only — no independent audit confirms it compounds in production at Palantir depth. Treat the convergence as "Distyl now claims Palantir-like architecture," not "Distyl has proven parity."
So, the whole thing in a paragraph. The deployment layer is real, big, and growing fast, and underneath the branding everyone is building the same machine: rent a frontier model, feed it the company's data in some organized way, check its work, deliver the result, send an invoice. The parts that are actually hard to copy are the boring ones — the plumbing that lets an agent safely change systems of record, the years of integration wiring, and the verification layer. The exciting-sounding differentiators — which model you use, “we only bill for outcomes,” armies of embedded engineers — are mostly either commoditizing, specific to customer service, or dictated by how messy the problem is rather than by any real strategy. And the secret the category doesn't advertise: almost nothing about “what works” can be proven from public data. Which is the real takeaway. In a market this foggy, being willing to say “we don't actually know, and here's exactly why” isn't a gap in the analysis — it's the most useful thing an investor can be handed.
Sources
- Palantir, Foundry Ontology — Overview / Core Concepts (digital-twin framing; semantic + kinetic elements; interfaces). palantir.com/docs/foundry/ontology/overview
- Palantir, Action Types & Functions Overview (governed writes as transactions). palantir.com/docs/foundry/action-types/overview
- Distyl AI, Technology — Context Mesh (persistent write-back knowledge graph; "each use case inherits everything"). distyl.ai/technology/context-mesh
- Distyl AI, Technology — Context Views ("the semantic layer"; Directed Review edits schema). distyl.ai/technology/context-views
- Zero Future Tech (primary-docs-derived), Palantir AIP Agent–Ontology Interaction (5-layer model). zerofuturetech.substack.com; corroborated by Palantir AIP features docs
- Distyl AI, NVIDIA Enterprise Agents ("living, traversable graph"; ReAct + deterministic pipeline; open Nemotron via NIM). distylai.substack.com/p/nvidia-enterprise-agents
- Palantir, AIP Bootcamp ("zero to use case in 5 days or less"). palantir.com/platforms/aip/bootcamp
- Palantir, Ontology Best Practices / Structural Guidance (compounding mechanism + duplication risk). palantir.com/docs/foundry/ontology/ontology-best-practices
- Sacra / Galileo prior verified fact sheet (June 2026), Distyl Distillery / Routines (SOP decomposition, auditable task chain).
- AgentMarketCap, Distyl AI $175M: Enterprise Knowledge Graphs & Agent Memory Layer ("applying Palantir's ontology thesis natively to LLM agents"). agentmarketcap.ai
- Cognition, Funding & product (Devin planner/executor; ~$400M @ $10.2B Sept 2025; >$1B Series D ~$26B May 2026). cognition.ai/blog; TechCrunch
- Happy Robot, Technical Overview / Voice AI (6-model pipeline; proprietary voice models; deterministic rule-guards; SIP/K8s). happyrobot.ai/blog/technical-overview
- Factory, GA announcement / product (model-agnostic per-task routing; no seat minimums). factory.ai/news/factory-is-ga
- Factory, Series C & Droids / Missions docs ($150M @ $1.5B, Apr 2026; role-scoped Droids; Review/Test verifier). factory.ai/news/series-c
- Harvey, Funding, FDE program & open-weight model ($200M @ $11B Mar 2026; "services business with software margins"; shipped an open-weight post-trained legal model, "Tenet"). harvey.ai/blog; harvey.ai/blog (Tenet)
- Cresta, AWS Marketplace listing & Sacra profile (per-seat + usage; 125K chats/yr at $150K; $125M Series D ~$1.6B). sacra.com/c/cresta-ai
- Glean, Series F & Enterprise Graph ($150M @ $7.2B Jun 2025; multi-model; per-seat). glean.com/press
- Cohere / Reuters, North platform, open weights & funding (trains own Command/Embed/Rerank; ships open weights under Apache-2.0 for self-host; private/air-gapped “GPU-in-a-closet” deploy; ~$6.8B val). Reuters; cohere.com/blog
- Menlo Ventures, 2025: The State of Generative AI in the Enterprise ($37B spend, 3.2× YoY; Anthropic 40%; 76% buy; 47% vs 25% conversion). menlovc.com
- MIT NANDA, The GenAI Divide: State of AI in Business 2025 (Jul 2025, "Preliminary Findings" v0.1; n=153 survey + 52 interviews + 300 initiatives; success = "remarked" P&L impact). Full PDF (mirror)
- ISG, State of Enterprise AI Adoption 2025 (n=1,200 use cases; 31% to production; $1.3M avg spend; 1-in-4 ROI; 50% efficiency). isg-one.com
- Gupta, A., ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress, arXiv 2601.06112. arxiv.org/abs/2601.06112
- Jia, J. et al., MAS-FIRE: Fault Injection & Reliability Evaluation for LLM-Based Multi-Agent Systems (silent semantic failures), arXiv 2602.19843. arxiv.org/abs/2602.19843
- Shlomov, S. et al. (IBM), From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production (CUGA), arXiv 2510.23856. arxiv.org/abs/2510.23856
- Palantir, Q1 2026 Letter to Shareholders (GE "using AIP to help increase engine production by 26%"). palantir.com/q1-2026-letter
- GE Aerospace, FY2025 Form 10-K (deliveries +26%, LEAP +28%, AI blade-inspection & material-assistant tools); SEC EDGAR full-text "Palantir" in GE filings = 0 hits. SEC 10-K
- Palantir, Q1 2026 Business Update (Associated Materials OTIF 40%→90%; private customer, no filing). investors.palantir.com
- CNBC, Goldman Sachs autonomous coder pilot (Devin); SEC EDGAR "Devin" in GS 10-K/10-Q = 0 hits. CNBC
- AIG & T-Mobile customer-side triangulation: AIG 2025 Annual Report (Lexington 370K+ submissions, ~+26% YoY, submit-to-bind ~+35–40%, claims hours→minutes) aig.com/investors; AIG Investor Day / earnings naming Palantir + Anthropic; 100% of private/non-profit submissions, $4B+ premium by 2030 target (Carrier Management, Apr 2025) carriermanagement.com; T-Mobile Q4-2025 Capital Markets Day (~$3B AI savings by 2027, ~60% chatbot containment; Distyl unnamed, credited “with OpenAI”) investor.t-mobile.com. Confirmed-Triangulated (AIG) / Partial (T-Mobile)
- Reuters, Cigna says AI tools save customers $200M over three years (2026-07-23; never names Sierra). Reuters; Sierra customers page: sierra.ai/customers
- Decagon / Cresta, Case studies (Rippling 38%→>50%; Cox +20%; private customers, no filings). decagon.ai; cresta.com/customer-stories
- Harvey, PwC UK customer story (saves 1–2 hrs per liquidation; private partnership). harvey.ai/blog
- Duolingo, FY2025 Form 10-K (generic "AI costs" in opex; no vendor, no deflection metric). SEC 10-K
- Center for Audit Quality, S&P 500 and AI Reporting (~90% mention AI in 2024 10-Ks). thecaq.org
- CFO Dive, Few 10-Ks tie AI to tangible revenue gains (85% Fortune 500 mention AI; 12% CEOs cite cost+revenue; 56% no benefit). cfodive.com
- Ramp, AI Index (card-transaction data, 50k+ businesses; ~4× monthly spend growth Feb'25→Feb'26). ramp.com/data/ai-index
- Accenture, Q4/FY2025 earnings + Palantir Business Group + M&A ($5.9B GenAI bookings; $2.7B advanced-AI revenue; RANGR, Decho acquisitions). Accenture earnings; Palantir group
- OpenAI, Frontier Alliance partners (McKinsey, BCG, Accenture). openai.com; McKinsey: mckinsey.com
- Anthropic, Deloitte–Anthropic partnership (Claude to ~470K staff — largest enterprise deployment); Deloitte–Palantir EOS: deloitte.com. anthropic.com
- EY, EY & 8090 launch EY.ai PDLC (Mar 2026; powered by 8090's Software Factory; performance claims are vendor projections). ey.com
- IBM, Q4 2025 results (GenAI book >$12.5B cumulative; ~$1.5B GenAI consulting Q3'25; consulting +3% YoY). newsroom.ibm.com
- BCG / Bloomberg, BCG says AI work brought 25% of 2025 revenue (~$3.6B AI-related); OpenAI Elite + Anthropic (2023). Bloomberg; anthropic.com/news/anthropic-bcg
- Evidently AI, Generative AI applications: 805 production use cases from 185 companies (production-only aggregator; measures claimed outcomes). evidentlyai.com
- Claritype, Company & product (ex-Palantir founders Rob Giardina + Ilya Lipkind; ontology/semantic-first; ~$6.6M seed 2022). claritype.com/company
- Sierra / TechCrunch, Funding & pay-per-resolution ($350M @ $10B Sept 2025; ~$950M @ ~$15.8B May 2026). TechCrunch
- Decagon / Reuters, Funding & per-resolution option ($131M @ $1.5B Jun 2025; $250M @ $4.5B Jan 2026). Reuters
- Northslope / Yahoo Finance, Funding & Palantir-native model (~$22M; Friends & Family Capital / ex-Palantir CFO Colin Anderson; AI app layer on AIP + Gemini). Yahoo Finance
- Happy Robot / Reuters, Funding ($44M Series B Sept 2025; $150M Series C ~$1.2B Aug 2026, freight/logistics voice). Reuters
- ZenML, LLMOps Database — what ~1,200 production deployments reveal (2025) (aggregated production case studies; “simple narrow workflows beat general-purpose agents,” “context engineering over prompt engineering”). zenml.io
- AIUseCaseHub, Enterprise AI deployment tracker (~3,814 public deployments across ~3,241 companies; filterable by maturity Production/Scaled Production and evidence strength; notes evidence signals ≠ operational audit). aiusecasehub.com/cases
- OpenAI, OpenAI launches the Deployment Company (forward-deployed engineers + deployment specialists; Tomoro acquisition adds ~150 FDEs subject to close). openai.com; WSJ (org/scope)
- Anthropic, Applied AI + Claude Partner Network (embedded Applied AI engineers for enterprise Claude deployments; $100M partner-network commitment, dedicated engineers + technical architects). anthropic.com/news/claude-partner-network; enterprise AI services
- joylarkin, Awesome-AI-Market-Maps (curated index of published AI market maps; most plot by stack layer or category, not by substrate-ownership × ontology). github.com/joylarkin/Awesome-AI-Market-Maps
- a16z, Your Data (Agents) Need Context (argues enterprise value comes from data/context gravity — permissions, metadata, workflow history — more than raw model quality). a16z.com/your-data-agents-need-context
- Kai Waehner, Enterprise Agentic AI Landscape 2026: Trust, Flexibility, and Vendor Lock-In (vendor analysis organized around lock-in / model-neutrality — the closest published lens to our data-gravity overlay). kai-waehner.de
Note on arXiv ids [33–35]: several Valency-corpus ids are forward-dated (2026 preprints); ids returned live by the corpus and HTTP-verified, but confirm each resolves before external citation. No source or metric in this report was fabricated; single-source and vendor-asserted claims are labeled as such throughout.
Generated by Galileo 🔭 · August 31, 2026 · Tomales Bay Capital