← All Reports

The Deployment Layer: What Actually Works

A category-wide strategy & architecture map of the companies that take frontier AI and make it work inside large enterprises — and an honest accounting of what the public evidence can and cannot prove.

📅 August 31, 2026 🔭 Galileo Research Tomales Bay Capital

Executive Summary

Here is the basic problem this whole industry exists to solve. A big company buys access to a frontier model — GPT-5, Claude, whatever — and discovers that a brilliant model sitting next to its messy data does approximately nothing. Someone has to wire the model into the company's actual systems, teach it the company's actual mess, point it at real work, and make it safe enough to trust. The companies that do that wiring are the "deployment layer," and they range from two-year-old startups to Accenture. This report maps them by strategy rather than one-by-one, scores 22 of them on a single matrix, and pokes hard at the one architectural argument the whole category is secretly having. The house rule throughout: every real claim gets a confidence label, and when the evidence is thin, we just say so — because in this category the honest answer is often "nobody actually knows," and pretending otherwise is how you lose money.

Bottom line: the deployment layer is real, big, and growing fast — enterprises roughly tripled their generative-AI spend to about $37B in 2025 — but “what works” is far less provable than the sales decks suggest. The edges that look durable are unglamorous: the plumbing that lets an agent safely write back into systems of record, the years of integration wiring that's annoying to rebuild, and that verification layer. The sexier moats — outcome-based pricing, armies of embedded engineers — are narrower than the hype. And the one question every investor actually wants answered — does the ontology approach compound over time, or does the workflow approach just win by getting to production faster? — cannot be answered from public data today. That gap is the single most important thing to go diligence in this category.

1 · The Deployment Layer in One Screen

What these companies are, and why the category is opaque

Start with the thing everyone gets wrong: a frontier model is not a product. You cannot buy Claude, point it at your company, and get a result — any more than you can buy a brilliant new hire and get value on day one before they know where anything is. Someone has to plug the model into your systems, teach it your specific mess, aim it at real work, check that it did the work right, and figure out how to charge for it. The companies that do that — the “deployment layer” — are where an enterprise's AI budget actually turns into anything useful. That's the whole job.

Every player in this report, from a two-year-old startup to Accenture, runs the same core loop:

The core loop everyone runs plug in the model → teach it your data → point agents at the work → deliver a result → bill for it. Every company here runs this same loop. What makes them different is which step they obsess over: some build a permanent map of the whole enterprise, some just automate one task at a time, some fly engineers out to sit with the customer, some bet the business on getting paid per result — and, crucially, each one has to build the part that double-checks the AI's work before it's safe enough to charge for.

Now, this category is genuinely hard to analyze, for three reasons — and those three reasons are why this report is written the way it is:

How to read the confidence labels

So here's the one habit that runs through the entire report: every real number gets a confidence label, so you always know how much weight it can bear. There are four:

Confirmed disclosed in a primary source (filing, official docs, company statement).   Reported credible secondary source, not independently verifiable.   Estimated / Inferred our analytical inference; no disclosure exists.   Consensus-but-Unproven widely repeated, weakly evidenced.

And when the only source is the vendor's own material, we call it Confirmed-as-claimed — meaning the company definitely said it, but nobody independent checked it. “They said so” and “it's true” are different things, and that difference is the whole game here.

2 · The Master Matrix

Twenty-two companies, five strategic axes

This is the map of the whole report — everything after it is just the walkthrough. Every company is scored on the five choices that define a deployment strategy. A note on who's here: OpenAI and Anthropic are now on the map, not just next to it. They used to be pure model suppliers — you rented the brains, someone else did the deploying. But in 2026 both stood up real deployment arms (OpenAI literally launched an "OpenAI Deployment Company" and bought a 150-person forward-deployed team; Anthropic built its "Applied AI" org and a $100M partner network of embedded engineers). So the arms dealers are now also fighting the war — which is a big enough deal that we score them and unpack it in the callout below. Scale, Mercor, and Turing are still left out on purpose; they're in the data-and-evaluation business next door, not the deployment business.

The five axes — and how to read the matrix

The big table below packs a lot into each cell, so here's the key. Each company is scored on five choices; these are the terms you'll see in the cells and what they actually mean:

AxisThe strategic choice it captures
1 · Models & weightsWhose brains does it use? Closed frontier = renting GPT/Claude/Gemini; open-weight = a model you can host yourself; agnostic router = swaps between them per task; owns models = trained its own. And whether it fine-tunes (custom-trains on a domain) or just uses models as-is.
2 · Ontology ↔ WorkflowDoes it model the whole enterprise once as shared, durable structure (ontology-first), or attack one workflow at a time (workflow-first)? The load-bearing axis (Section 3).
3 · OrchestrationThe machinery that runs the AI. Ingestion/semantic layer = getting the company's data in and organized; planner/executor = one model breaks the job into steps, others do them; verifier = the part that checks the AI's work before it counts; HITL (human-in-the-loop) = a person approves before anything ships; memory/state = what it remembers between runs.
4 · DeliveryHow you buy it. FDE (forward-deployed engineers) = the vendor sends engineers to build inside your company; product-led/self-serve = you just sign up and use it. Many land as services, expand into a platform.
5 · CommercialHow you're billed. Outcome-based = pay per result (e.g. per resolved ticket); seat/subscription = per user; consumption = per usage (tokens/calls). And who's contractually on the hook for the result.
Company1 · Models & weights2 · Ontology↔Workflow3 · Orchestration4 · Delivery5 · Commercial
PalantirModel-agnostic: open-weight (Llama/Mistral) first-class + BYOM self-host (vLLM/Ollama) for on-prem/air-gapped; frontier via catalogOntology-first (pole)Ontology semantic layer + AIP orchestrator + verifier + HITLFDE → platform/partnerLicense + consumption (value-framed)
DistylMulti-model: closed frontier (OpenAI+Anthropic via Azure) for hardest reasoning + open-weight Nemotron via NVIDIA NIM for always-on/long-running agentsWorkflow-first → convergingRoutines task-graph + Context Mesh + verifierFDE-embed (purest)Outcome-framed services → platform
NorthslopeGemini Enterprise + AIP-brokered frontierOntology-first (rented from Palantir)AIP substrate + client app layer + HITLFDE-embed (purest)Services/project + substrate passthrough
ClaritypeFrontier-orchestrated (unspecified)Ontology / semantic-first (ex-Palantir DNA)Semantic model + NL→governed-query + inspectable reasoningProduct-led self-serveSeat / SaaS (not disclosed)
SierraMulti-model “constellation” (15+): frontier + open-weight + proprietary, routed per subtaskWorkflow-firstSelf-supervising agent + verifier + HITLProduct-led + SEPure pay-per-resolution[57]
Cognition (Devin)Closed frontier (Anthropic-heavy)Workflow-firstPlanner/executor + self-test verifier + PR HITLProduct-ledSubscription + usage (ACU)
Happy Robot6-model pipeline: swappable LLM + proprietary voice modelsWorkflow-first (rule-guarded)SIP/K8s voice stack + rule-verifier + human QA + HITLProduct-led + enterprise onboardingBundled platform + usage
FactoryModel-agnostic per-task router (Claude/GPT/Gemini)Workflow-firstOrchestrator + role-scoped Droids + Review/Test verifier + HITLProduct-led, no seat minimumsSubscription / consumption
DecagonMulti-model frontierWorkflow-first (AOPs)Runtime + knowledge grounding + QA verifier + HITLProduct-led + SEPer-resolution option (most pick per-conversation)
HarveyFrontier + legal fine-tunes; ships own open-weight model (“Tenet”)Workflow-firstRetrieval + task agents + citation verifier + heavy lawyer HITLFDE program (services → platform)Seat / subscription
Cresta20+ models: frontier + fine-tuned open-weight (Mistral base) + Ocean-1 CC foundation modelWorkflow-firstReal-time transcription + agent runtime + QA verifier + HITLProduct-led + SEPer-seat + usage ("outcome-LED")
GleanMulti-model (Bedrock/Vertex/Azure OpenAI)Ontology-leaning (Enterprise Graph)Connectors + graph semantic layer + RAG + Agents + citation verifierProduct-led SaaSPer-seat subscription
CohereOwns own models (Command/Embed/Rerank); open weights (Apache-2.0); private/air-gappedWorkflow-first (platform)North RAG + agent orchestration + verifier, in-VPCModel/platform vendor (no FDE)Consumption (tokens) + platform license
8090Frontier-orchestrated (agnostic)Workflow-first (PDLC)Software Factory pipeline + engineer-verifierServices-led via EY (SI channel)Project/services (outcome-framed, undisclosed)
OpenAI (lab arm)Owns own models; single-vendor, not single-modality — deploys OpenAI models only, but that now spans closed frontier + own Apache-2.0 open-weight gpt-oss (on-prem-capable)Workflow-first (use-case-led)Own models + FDE build + evals + HITLFDE-embed ("Deployment Company"; Tomoro ~150 FDEs)Consumption (tokens) + enterprise + services
Anthropic (lab arm)Owns own models (Claude), closed frontier only; ships no open weights — purest single-model deploycoWorkflow-first (use-case-led)Own models + Applied AI build + evals + HITLFDE-embed ("Applied AI" + $100M partner network)Consumption (tokens) + enterprise + services
AccentureMulti-model; Palantir + OpenAI aligned + open-weight via AI Refinery/NVIDIA (custom Llama + Llama Nemotron NIM); on-prem/sovereignOntology-via-partner (Palantir group)Accelerators + FDE model (absorbed via M&A)Bodyshop → asset/acceleratorLabor leverage + accelerators
McKinsey (QuantumBlack)OpenAI + Anthropic (ecosystem)Advisory (n/a)Lilli / Horizon / Kedro internal assetsAdvisory-led, high-touchFees (judgment + seats)
DeloitteAnthropic (Claude to ~470K); Palantir EOS; open-weight Llama Nemotron (NVIDIA stack) powers Zora AI, on-prem-capableOntology-via-partner (Palantir EOS)Zora AI agentic + Omnia + Global Agentic NetworkBodyshop → agentic productLabor leverage + emerging product
EYFrontier via 8090 + MicrosoftWorkflow (via 8090 PDLC)Rents 8090's Software FactoryBodyshop renting a factoryFees + rented delivery
BCG XOpenAI Elite + AnthropicWorkflow (bespoke build)Build-and-scale org + ARTKIT red-teamingConsulting-led build orgFees (AI ≈ 25% of rev)
IBM ConsultingOwns Granite (Apache-2.0 open weights) + Llama/Mistral via watsonx; integrates OpenAI/Anthropic; on-prem/cloudWorkflow (watsonx-anchored)watsonx.ai/.governance/Orchestrate + consultingServices welded to owned stackSoftware + consulting signings

All matrix placements are evidence-backed in Sections 3–8. Cell confidence ranges from Confirmed (Palantir/Distyl architecture, Cohere model ownership, Factory routing, disclosed revenues) to Estimated (most pricing/delivery cells for private companies). Sources §.

The labs are now players, not just suppliers Here's the wrinkle worth staring at. For two years the deal was simple: OpenAI and Anthropic sold the models, and a whole ecosystem of startups did the messy work of deploying them inside enterprises. The labs were the arms dealers; the deploycos were the soldiers. In 2026 that stopped being true. OpenAI launched an actual “Deployment Company” — a services org with forward-deployed engineers — and bought Tomoro (~150 FDEs) to staff it overnight.[63] Reported Anthropic built “Applied AI,” its own embedded-engineer function, plus a $100M partner network of Applied AI engineers and technical architects.[64] Reported So the supplier is now walking straight into its customers' business. Why? Because the deployment layer is where the margin and the stickiness live — the model is a commodity you rent, but the deployment is a relationship you own. The strategic read for an investor: every deployco in this report now competes with its own most important vendor, and that vendor has infinite model access and a direct line to the same CIOs. It's the same move Distyl made going up into ontology, run in reverse — the labs collapsing down into deployment. The one hedge: the lab arms are single-vendor by design (Anthropic's Applied AI deploys Claude exclusively and ships no open weights — the purest single-model case; OpenAI's Deployment Company deploys OpenAI models only, though since gpt-oss its portfolio now spans closed frontier and its own Apache-2.0 open weights), while the independents stay model-agnostic and — as the matrix now reflects — most already run open-weight models (Nemotron, Llama, Mistral, Granite) somewhere in the production pipeline, especially for regulated/on-prem/air-gapped work. So “we're not locked to one lab's models” becomes a live selling point it wasn't a year ago.

The deployment layer, mapped on the two questions that decide the valuation

Two axes and a color, each a distinct question.
Left–right (X): what do you build first? Model the whole enterprise (ontology, right) or attack one workflow (workflow, left).
Up–down (Y): do you own your substrate or rent it? “Substrate” = the model + semantic layer + platform your customers build on. Own it (top) and you get platform economics and a platform multiple; rent someone else's (bottom) and you get app-layer margins plus dependency risk. This is the report's single clearest value-capture fork — which is why it's the vertical axis of the flagship chart, not a footnote.
Dot color: data gravity. ● Rose = capture (to work, your data moves into the vendor's platform — lock-in). ● Teal = sovereignty (it acts where your data already lives). ● Grey = mixed.
Placements are our analytical read Estimated-Inferred from public architecture evidence, scored on those single questions — not disclosed numbers.

Three things to notice. Palantir sits top-right and rose: it owns its substrate (platform economics) but it's a capture business — Foundry pulls your data into Foundry. That capture/lock-in is the whole rub, and the color now says so out loud while the height still credits the platform ownership. Northslope sits bottom-right: same ontology ambition as Palantir, but it rents Palantir's substrate — app-layer economics and dependency risk in one dot. The labs (OpenAI, Anthropic) sit top-left and rose: they own the deepest substrate there is (the model itself) but are workflow-led and capture-heavy — you send your data to them. The map's punchline: almost everything is rose. Genuine data sovereignty (teal — Cohere, Claritype) is rare, and that scarcity is itself a differentiator.

Prior art: market maps of this category usually plot vendors by stack layer (infra / orchestration / apps) or by category logos — see the [65] market-map index. The closest published framings to ours are a16z's argument that data/context gravity outweighs raw model quality[66] and vendor analyses organized around trust, flexibility, and lock-in.[67] We didn't find an existing map that crosses substrate ownership against ontology-vs-workflow and overlays data gravity for this specific deployco cohort — so treat this as our own synthesis, informed by those, not a restatement of a standard chart.

3 · The Load-Bearing Section

Ontology-first vs. workflow-first — and why the fork is quietly collapsing

This is the central architectural fork of the category and the deepest section of the report. Two opposite bets, same go-to-market motion: model the whole enterprise first (Palantir), or attack one workflow first (Distyl). The investor's prior was that these two might not actually be that different. We were mandated to disprove the fork before asserting it — and the disprove-first hunt landed the report's most important finding.

The money finding The clean “whole-enterprise vs. one-workflow” split is now mostly a difference in what you build first, not what you build. Distyl, supposedly the poster child for the one-workflow approach, has quietly gone and built a permanent, governed, company-wide data layer it calls Context Mesh — and describes it in language almost identical to Palantir's famous moat.[3] Meanwhile Palantir now sells 1–5-day workshops that stand up a single working workflow fast[7] — which was supposed to be Distyl's whole edge. They started at opposite ends and walked toward each other. What's left between them is real but narrow: how deep and how safely each one can write changes back into the systems that actually run the business, plus the years of unglamorous integration wiring that's a pain to rebuild.

Palantir (the ontology pole), mapped at architecture depth

Palantir's differentiation is architectural, not merely go-to-market. Foundry's Ontology is a largely deterministic, governed model of an enterprise — Palantir frames it as representing the enterprise's decisions, not merely its data.[1] It has two first-class halves:

Here's how the AI actually touches all this, in plain terms. It never gets handed the company's raw database. Instead it works through five guardrails: (1) only the relevant facts get pulled into its view on each message, never the whole schema; (2) it can walk the connections between objects — Order → Shipment → Carrier → hub — which ordinary keyword search simply can't do; (3) it can run the company's actual business logic; (4) when it wants to change something, that change goes through a controlled, optionally human-approved “Action” rather than a raw database write; and (5) the whole thing is governed, so the agent only ever acts through the sanctioned boundary.[5] Palantir's own phrase for the philosophy: “determinism where possible, autonomy where necessary, constraint everywhere.” The key point for an investor: the valuable part is not the model — Palantir just rents whichever model is best — it's this governed layer that turns the reckless version (“hand the AI your database password”) into the safe version (“let the AI act only through a controlled interface”).[5] Estimated-Inferred

Distyl (the workflow pole), and its 2026 convergence

In June our verified map of Distyl was the pure workflow-first archetype: a Standard Operating Procedure auto-decomposed into "Routines" — discrete, auditable tasks with logged inputs/outputs/tool-calls — a thin, local semantic layer per workflow.[16] That is still true at the execution layer. But Distyl's own 2026 technology pages describe a four-layer platform whose foundation is now a persistent semantic layer Confirmed-as-claimed:

Distyl's own NVIDIA-integration post describes moving "past the concept of the flat, static retrieval index" to build "a living, traversable graph… encompass[ing] episodic memory, tacit and explicit knowledge, policies, workflows, and domain logic," where agents "write back state that persists long after the session ends."[6] A third-party analysis states it plainly: "Distyl is applying [Palantir's ontology] thesis natively to LLM agents."[17] Reported

Head-to-head: same destination, different construction method

DimensionPalantir Foundry/AIP (ontology-first)Distyl Distillery 2026 (workflow-first → converging)
Persistent semantic layer?YES — Ontology (objects/props/links)YES — Context Mesh + Context Views
Typed entities + relationshipsObject types / link types (hard schema)Nodes/edges: "typed relationships" Rep
Governed write-back to the layerAction types (transactional, audited)"writes back what every agent decides" as-claimed
Multi-hop traversal by agentsObject Query Tool link traversal"living, traversable graph," multi-hop as-claimed
Data stays in placeVirtual tables enable in-place"No migration required" as-claimed
How the model is BUILTEngineers design schema up front (deterministic)SMEs + agents accrete it; "Directed Review" edits schema
Compounding mechanismInterfaces + canonical objects → workflows inherit"Each use case inherits everything Context Mesh knows"
Maturity of the layer~decade of production hardeningNew (2026 re-architecture); unproven at Palantir depth

Disprove-first: the null hypothesis is strongly supported

Evidence Distyl is moving toward Palantir: Context Mesh is a write-back knowledge graph; Context Views is an explicit "semantic layer"; the ex-Palantir founders (Arjun Prakash + Derek Ho) make the ontology DNA intentional.[3][4] Evidence Palantir is moving toward Distyl: the AIP Bootcamp promises "zero to use case in 5 days or less," with "at least one production-ready workflow by day five" — Palantir attacking the fast per-workflow motion.[7] Confirmed-as-claimed

Verdict on the fork The fork survives only as (1) a sequencing difference (build-model-first vs. ship-first-then-accrete) and (2) a schema-determinism difference (a designed, strongly-typed object model where agents cannot invent objects, vs. a mesh partly accreted from agent decisions). The genuine, surviving daylight is the depth of the deterministic action/write-back layer into systems of record — Palantir's atomic, staged, governed writes into ERP/WMS/edge over a decade of connectors are far deeper than anything Distyl publicly documents (Distyl writes back into its own mesh, not documented ERP-grade write-back). That is a head-start moat, not a categorically different architecture. As an architectural claim, "ontology-first vs. workflow-first" no longer describes two different architectures — it describes two on-ramps to the same highway. That itself is the headline.

The compounding-vs-ceiling stress-test

The thesis under test: workflow-first converts to production faster but ceilings into 40 disconnected point-solutions; ontology-first is slower-to-value but compounds because every new workflow inherits the object model.

The core diligence gap — stated plainly The "ceilings out into disconnected point-solutions" outcome is not observable in public data for Distyl or any workflow-first peer. No public deployment census shows point-solution sprawl vs. compounding; no audited before/after exists. The stress-test is architecturally sound and corroborated by Distyl's revealed behavior, but empirically unproven. Production-conversion rates are unknown for both poles. Any investment thesis that rests on "ontology-first compounds and workflow-first ceilings" is, today, resting on architecture and inference — not measured outcomes. Consensus-but-Unproven
4 · Models & Weights

Open vs. closed, by use case — and the one company that owns its models

The pattern here is refreshingly clear for once: 13 of the 14 independents just use the big frontier models as-is — they don't train their own. The deployment layer sits on top of the model, not inside it. So the only real choices are which models, how easily you can swap between them, and whether you bother fine-tuning one for a specific domain.

The single most-cited enterprise-spend datapoint of 2026: Anthropic leads enterprise LLM spend at 40%, having "unseated OpenAI" (Menlo Ventures).[30] Reported — a market-model estimate from a VC with portfolio exposure, directionally strong but not audited. Distyl, Cognition, and Harvey all skew Anthropic-heavy in their orchestration, consistent with that leadership.

So what Since nearly everyone is renting the same few models, the model itself is not the moat — picking one is now basically an automated routing decision, like choosing the cheapest flight. Whatever defensibility exists lives in the stuff around the model: how you feed it the company's data, how you orchestrate it, how you check its work, and how you deliver. Which is exactly why the Section 3 argument about ontology matters more than “which model do you use.”
5 · Orchestration Architecture

Where designs converge — and the one component everyone builds

Strip the marketing away and the orchestration stacks rhyme: ingestion/connectors → a semantic or context layer → a planner/executor split (an orchestrator model directing task models) → tool and action calls → a verification layer → human-in-the-loop checkpoints → memory/state. The interesting question is where they converge vs. genuinely differ.

The universal building block: the verification layer

Category-wide convergence Everyone, separately, built a “verifier” — the part that checks the AI's work — and it's what makes the output sellable in the first place. The flavors differ by industry: answer-quality checks for customer service (Sierra, Decagon), code review and test bots for engineering (Factory), citation-grounding for law (Harvey), hard rule-guards for freight quotes (Happy Robot), show-your-work reasoning for analytics (Claritype), and the governed “Action” system for Palantir. But it's the same idea every time, and it's the load-bearing commercial part of the whole category: without it, a customer has no reason to trust — or pay for — what the agent produced.

The academic literature corroborates why this layer is non-negotiable: single-run benchmark success does not predict production reliability, and agent failures are frequently silent — hallucination, misinterpreted instructions, and reasoning drift that propagate without throwing any runtime exception (MAS-FIRE; ReliabilityBench; IBM's CUGA deployment paper).[34][33][35] Reported A verifier is the only thing standing between "the demo worked" and "the deployment silently corrupted a workflow."

Where designs genuinely differ

The planner/executor pattern (an orchestrator model decomposing work and directing task models) is now near-universal in the agentic players — Factory's coordinator assigning role-scoped Droids[25], Cognition's long-horizon task decomposition[18], Palantir's AIP Logic. It is converging into a commodity pattern; the differentiation has moved to the context layer beneath it and the verification layer above it.

6 · Delivery & Commercial Models

FDEs are dying at the edges; outcome pricing is rare and CX-shaped

Delivery: forward-deployed engineering survives only where the substrate is complex

Forward-deployed engineering — the Palantir-invented playbook that started this category — is expensive, and across the 14 independents only four still really run it: Distyl, Northslope, Harvey (literally described as “a services business with software margins”), and 8090 (delivered through EY).[26] Reported Everywhere the job is cleaner and more self-contained — writing code (Factory, Cognition), search (Glean), support tickets (Decagon, Sierra), voice (Happy Robot) — companies skipped the embedded-engineer model and just shipped a product you can buy yourself. Factory sells off the shelf with no seat minimums[24]; Glean lands on seats and grows into agents.[28]

The pattern Whether you need embedded engineers isn't about philosophy — it's about how messy the job is. Rewiring a whole enterprise across siloed, filthy data needs humans on-site. Resolving a support ticket or shipping a code change doesn't; you can package that as software. And as the models get better, more jobs move from the first bucket to the second, so the embedded-engineer model keeps retreating to the genuinely hard, cross-system builds. That, by the way, is also why the consulting giants (Section 8) are a channel and not roadkill — they already own the relationships for exactly the big, complex, high-touch deployments where embedding still wins.

Commercial: outcome pricing is rarer than the narrative claims

“We only get paid when you get a result” is the category's favorite line. In practice, genuine pay-for-outcome billing is rare, and almost entirely a customer-service thing. Only Sierra truly bills per-resolution. Decagon [58] offers it as an option but says most customers still pick per-conversation because it's predictable. Cresta talks “outcome-led” but actually invoices per-seat plus usage (its own rate card: 125K chats/yr for $150K).[27] Confirmed (rate card) Coding, legal, and search bill the old-fashioned way — seats and usage.

The reason is simple once you see it: to bill per outcome, you need an outcome you can count — a “resolved ticket.” That clean unit exists in customer service and almost nowhere else. You can't invoice “per transformed supply chain.” So wherever you hear “own the outcome” outside of CX (Distyl, 8090), it's marketing language wrapped around what is really a services contract, not a meter that ticks per result. Estimated-Inferred — no rate cards exist; the pricing structure is our inference.

Where "own the outcome" becomes a liability Contractually owning the outcome transfers risk to the vendor. In CX with a countable resolution and a verifier, that risk is priceable. Outside CX — where the "outcome" is a multi-quarter transformation with confounded attribution — owning it means owning blame for results the vendor only partly controls. The companies loudest about outcomes (Sierra, Decagon) are precisely those where the outcome is measurable; the ones doing complex builds quietly price on services. That asymmetry is a tell about where outcome pricing actually works.
7 · What Actually Works — The Evidence

The hard data: mostly, it does not exist — and that is the finding

Everyone wants the same thing here: hard proof of what actually converts a pilot into production and moves the P&L — not another testimonial. We looked three ways: at what the big deployment aggregators show across thousands of projects, at the customer-side filings behind specific wins, and at the academic literature. The honest headline: the gold-standard proof barely exists in public, and the statistic everyone quotes is the flimsiest one in the pile. But the aggregate data is real and it does say something — so start there.

7.0 — What the aggregators actually show, at scale

Rather than cherry-pick a few flattering case studies, start with the databases that have already collected thousands of real deployments and asked “what shipped, and what held up?” Three are worth knowing, and they agree with each other:

The catch with all of them These databases are enormous and genuinely useful for one question — what does a shipped system look like? — and useless for the question everyone actually asks — what's the success rate? They're built from published deployments, which means they're a list of survivors with no graveyard next to it: a big numerator and no denominator. And crucially, they're indexed by model (Copilot, Bedrock, Claude), almost never by deployco — so they tell you AI is being used, not whose deployment approach won. To get to “who actually moved the number,” you still have to do the filing-by-filing dig below. Which is where it gets bleak.

Where AI actually ships — deployments cluster in a handful of jobs

Pool the aggregator catalogs and one thing jumps out: production AI isn't spread evenly across the enterprise — it piles up in a few well-defined jobs (customer support, writing code, search/knowledge) and thins fast after that. This is the “narrow workflows win” thesis as a picture: the cleaner and more countable the task, the more real deployments exist. Bars show the rough share of catalogued production deployments by function, pooled and normalized across the ZenML and Evidently corpora.[61][55] Estimated-Illustrative — aggregators use different taxonomies; treat as directional shape, not precise percentages.

The enterprise-AI funnel: the money is huge, the throughput is not

The single hardest fact about this category, in one chart. Enterprises are spending enormous money — ~$37B on GenAI in 2025, roughly triple the prior year.[48] But that money narrows sharply as it moves toward results: the widely-quoted (and weak) MIT/NANDA line puts pilot success in the single digits[31], ISG's survey work puts production conversion around ~31% with roughly 1-in-4 reaching clear ROI.[32] Different studies, different denominators — so read the shape, not the exact widths: lots of spend in, not much proven value out. Estimated-Illustrative — stages come from different sources and are not a single measured cohort.

7.1 — Named contract outcomes: exactly one clean triangulation (AIG), and it took the customer choosing to disclose it

So the aggregators can't tell you who won. Can anyone? The gold standard would be a public company saying, in its own audited filing, “this specific vendor produced this specific result.” That beats any case study, because the customer has to answer to auditors. We took the deployco case studies that named public customers and went looking for the number in the customer's own annual reports, 10-Ks, investor days, and earnings calls. Nine worked examples. Here's how many survived:

Grading vocabulary:  Confirmed-Triangulated customer's own disclosure names the deployco AND a quantified outcome.  Partial customer corroborates the outcome magnitude but not the vendor name (or vice-versa).  Vendor-Claimed deployco-published only, no customer-side corroboration.  A separate class, Logo-Confirmed / Outcome-Unattributable, is where a customer confirms the relationship (logo wall, press release) but no number attaches to it.
EnterpriseDeploycoClaimed resultGradeCustomer-side check
AIG (NYSE:AIG)Palantir + AnthropicLexington 370K+ submissions (~+26% YoY); submit-to-bind ~+35–40%; underwriting cycle weeks→<1 day; 100% of private/non-profit submissions reviewed without adding underwriters; target $4B+ new premium by 2030Confirmed-TriangulatedAIG's own 2025 annual report / investor day / earnings disclose the outcomes and name Palantir + Anthropic as the partners[40]
GE Aerospace (NYSE:GE)Palantir (AIP)"AIP to help increase engine production by 26%"Partial10-K corroborates deliveries +26%, LEAP +28%, names AI tools — but 0 EDGAR hits for "Palantir"[37]
T-Mobile (NASDAQ:TMUS)Distyl~$3B AI savings by 2027; ~60% chatbot containmentPartialT-Mobile's own Q4-2025 CMD calls corroborate the magnitude but never name Distyl — credits a partner “with OpenAI”[40]
Assoc. Materials (private)PalantirOTIF 40%→90% in 9 monthsVendor-ClaimedPrivate — no filing to triangulate[38]
Goldman Sachs (NYSE:GS)Cognition (Devin)"Hundreds of Devins, scaling to thousands"Vendor-ClaimedExec press quote, no P&L number; 0 EDGAR hits for "Devin"[39]
Cigna (NYSE:CI)SierraAuth time −80%, live in 8 weeksVendor-ClaimedCigna discloses $200M AI savings but never names Sierra; different metric[41]
Rippling (private)DecagonDeflection 38%→>50%Vendor-ClaimedPrivate — no filing[42]
Cox (private)CrestaRevenue per chat +20–30%Vendor-ClaimedPrivate — EDGAR "Cresta" hits are pre-2020 false positives[42]
PwC UK (private)HarveySaves 1–2 hrs per liquidationVendor-ClaimedPrivate partnership — no filing[43]
Duolingo (NASDAQ:DUOL)Decagon80% chat deflectionVendor-Claimed10-K books generic "AI costs," no vendor, no metric[44]

Tally: 1 CONFIRMED-TRIANGULATED (AIG) · 2 PARTIAL (GE, T-Mobile) · 6 VENDOR-CLAIMED-UNCORROBORATED. The one clean case is AIG → Palantir + Anthropic — the most valuable page in this report. AIG's own 2025 annual report discloses the measured outcomes (Lexington 370,000+ submissions, up ~26% YoY; submit-to-bind up ~35–40%; claims cycle times from hours to minutes; figures REPORTED-from-primary — relayed from AIG-hosted primary docs, as the AIG PDF fetched as raw bytes rather than quoted verbatim), and AIG's own investor day and earnings call name Palantir and Anthropic as the partners and describe reviewing 100% of private/non-profit submissions without adding underwriters, targeting $4B+ new premium by 2030.[40] Two independent AIG-authored channels close the triangle.

The one honest nuance on AIG Even the category's single clean triangulation has a seam: the metrics live in the annual-report layer; the vendor attribution lives in the investor-day / earnings layer. They are both AIG-authored, but a maximally adversarial reader can argue the specific number is not audit-tied to the vendor within one document. That caveat is itself the point — this is as good as customer-side outcome evidence gets in the entire category, and it still isn't a single audited line crediting a vendor with a number. Numbers also drift across AIG's own reporting layers (e.g., +26%/+35% vs. +30%/+40% quoted/binding in different disclosures), so we cite the range, not a false-precision figure.

The two partials show the more common failure. GE Aerospace's FY2025 10-K independently reports deliveries up 26% and discloses concrete AI-productivity tools (an AI blade-inspection tool halving inspection time; an AI material assistant cutting MRO turnaround) — but a full-text EDGAR search for "Palantir" across all GE filings returns zero hits.[36][37] T-Mobile's own Q4-2025 calls corroborate the ~$3B-savings magnitude of Distyl's claim but never name Distyl (crediting a partner “with OpenAI”).[40] Each confirms an AI-driven gain; each strips or mismatches the vendor.

Why the evidence is structurally missing — and why AIG is the exception The one clean triangulation exists only because AIG itself chose to disclose it. The category's outcome evidence is otherwise systematically non-triangulable, for three compounding reasons: (1) vendors are laundered out of audited filings — auditors and IR will not credit one software vendor for a headline operating metric, so even real gains lose the attribution (GE/Palantir). (2) The crispest vendor numbers cluster on private buyers (Associated Materials, Rippling, Cox, PwC) with no filing obligation — the loudest claims are the least checkable. (3) Public buyers disclose the benefit but not the vendor, and often a different number (Cigna quantifies $200M but never names Sierra; T-Mobile credits “OpenAI” not Distyl). The customer and the vendor each hold one edge of the triangle and almost never publish the third — AIG is the rare buyer that published all three.
Method note — the logo wall vs. the outcome A deployco “logo wall” proves a relationship, not a result. We separate Logo-Confirmed / Outcome-Unattributable — the customer is real and confirmed (press release, named case study, marketplace listing) but no quantified P&L number is customer-attributed — from genuine outcome triangulation. Most of the category's “proof” is logo-confirmed-but-outcome-unattributable; treating a logo as an outcome is the single most common diligence error in reading these decks.

Investment implication: any deployco pitch resting on "named enterprise + quantified outcome" is, evidentially, vendor-attested marketing — even when the enterprise is real and public. Diligence should treat these as leads, then (a) check EDGAR full-text for the deployco's name in the customer's filings (usually 0 hits) and (b) look for whether the customer independently quantifies any aligned metric. The GE-style partial is the ceiling of available evidence, not the floor. The base rate confirms it: ~90% of the S&P 500 mention AI in 10-Ks[45], yet only ~12% of CEOs cite both cost and revenue benefit, and 56% report no significant financial benefit.[46] Reported

7.2 — Pilot→production: grading the "95%," not parroting it

You've seen the “95% of AI pilots fail” line — it was everywhere. It comes from one source, MIT/NANDA's The GenAI Divide: State of AI in Business 2025, and we read all 26 pages so you don't have to. Here's what's actually behind the number: a look at “over 300 publicly disclosed AI initiatives,” 52 interviews, and a survey of “153 senior leaders collected across four major industry conferences.”[31] That is a conference-hallway survey, not a census — and it gets shakier from there:

DimensionFindingGrade
Samplen=153 survey (conference-recruited) + 52 interviews; 300 initiatives are publicly disclosed (selection-biased)Thin
Self-reported vs. measuredSelf-reported; the report's own note calls figures "directionally accurate based on individual interviews rather than official company reporting"Weak
Definition of "success"Verbatim: a tool "users or executives have remarked as causing a marked and sustained… P&L impact." Success = a subjective remark, not an audited outcomeSoft
Window / review6-month window; labeled "Preliminary Findings," v0.1; not peer-reviewedNot peer-reviewed

Crucially, the 95% refers specifically to custom/task-specific enterprise GenAI tools; generic chatbots (ChatGPT/Copilot) are reported at ~83% adoption and explicitly excluded.[31] Autopsy verdict: a defensible qualitative signal that custom GenAI struggles to reach durable production — but it does NOT support precise quantitative claims. The number's virality outran its rigor. Consensus-but-Unproven Cite the pattern; never the precision.

The defensible benchmarks come from higher-rigor sources, all verified against primary pages:

Reconciling the numbers MIT (single-digit% for hard custom builds) and ISG (31% of prioritized use cases) are not contradictory — different denominators. The truth is bracketed: ~single-digit% for hard custom builds, ~30% for the broader prioritized portfolio. And two independent sources agree on the one robust best-practice: buy-over-build converts faster (Menlo 47% vs. 25%; MIT ~2× for partnerships). Everything else labeled "best practice" — learning/memory systems, mid-market speed — is consensus, not measured. Public production-case libraries (e.g. Evidently AI's ~805 cases across 185 companies [55]) are useful for "what shipped looks like" but are survivorship-biased numerators with no denominator — they cannot ground a base rate.

7.3 — The academic corpus: the failure mechanism is real and named

The Valency academic corpus returned 11 on-thesis, real arXiv papers. The consensus is consistent and useful: (1) single-run benchmark success ≠ production reliability (ReliabilityBench introduces consistency/robustness/fault-tolerance metrics)[33]; (2) agent/RAG failures are frequently silent (MAS-FIRE: semantic failures propagate without exceptions)[34]; (3) enterprise-specific gaps recur — domain process knowledge, evolving requirements, role-based data access. IBM's own deployment paper is the most on-thesis case study: "evidence of [generalist agents'] use in production enterprise settings remains limited" — benchmark strength does not equal deployed business value.[35] Reported (arXiv preprints — verify ids before external citation.)

The honest bottom line on evidence No rigorous, peer-reviewed, population-representative study of pilot→production failure exists publicly. The most-cited stat is the weakest source; the strongest sources measure spend (real and exploding) rather than success. The money is unambiguous — $37B, 3.2× YoY, ~4× in card data. Whether that money is working is genuinely opaque, bracketed between single-digit and ~30% production conversion depending on how you count. Anyone claiming precise certainty about "what works" in this category is selling something.
8 · The Incumbents

Threat, channel, or roadkill? Mostly channel.

Our June writeup dismissed the consulting firms in two sentences. That was a mistake. Accenture, McKinsey (QuantumBlack), Deloitte, EY, BCG X, and IBM Consulting are the biggest, best-funded sales channels into exactly the giant, cautious enterprise buyers that a startup can't reach on its own. And the important fact is what they did: every one of them chose to partner with or buy the deployment layer rather than try to out-build it. When the incumbents decide to resell you instead of copy you, that tells you where the value is.

FirmDelivery modelDisclosed GenAI $ (labeled)Key startup/platform bridgeVerdict
AccentureBodyshop → asset/accelerator; absorbing FDE model via M&AGenAI bookings $5.9B FY25; advanced-AI rev $2.7B ConfirmedPalantir Business Group (>2,000 certified); acquired RANGR, DechoChannelThreat-that-absorbs
McKinsey (QuantumBlack)Advisory-led, high-touch; strong internal assets (Lilli, Kedro)Not separately disclosed n/aOpenAI Frontier Alliance; AnthropicChannel (upstream)
DeloitteBodyshop → agentic product (Zora AI, Omnia)Not disclosed (total $70.5B FY25) n/aAnthropic (Claude to ~470K staff — largest deployment); Palantir EOSThreatChannel
EYBodyshop renting a startup's factoryNot disclosed n/a8090 (founding partner, EY.ai PDLC); Microsoft >$1B/5yrChannel (dependent)
BCG XConsulting-led build-and-scale; real product org~25% of 2025 rev ≈ ~$3.6B AI-related (not GenAI-specific) Reported [54]OpenAI Elite/Frontier; Anthropic (since 2023)ChannelThreat
IBM ConsultingServices welded to owned stack (watsonx)GenAI book >$12.5B cumulative Q4'25; ~$1.5B GenAI consulting ConfirmedOwns watsonx; integrates OpenAI/AnthropicChannel (captive)Roadkill in open GenAI

The tell is in what they actually did. EY rents 8090's software factory rather than building one[52]; Deloitte put Claude in front of ~470,000 of its own people (Anthropic's largest deployment anywhere) and wraps Palantir's platform[51]; McKinsey and BCG joined OpenAI's Frontier Alliance[50]; Accenture stood up a 2,000-plus-person Palantir unit and simply bought two embedded-engineering shops (RANGR, Decho).[49] Every one of these is a firm admitting it can't out-build the startups — so it resells them instead.

The disclosed numbers frame the stakes. Accenture's $5.9B GenAI bookings and $2.7B advanced-AI revenue[49] Confirmed and IBM's >$12.5B cumulative GenAI book[53] Confirmed prove demand is real and large — but both are single-digit shares of firm revenue and heavily bookings-weighted (i.e., early). EY, McKinsey, Deloitte, and BCG are private and do not disclose GenAI-specific revenue; BCG's ~$3.6B is AI-related (a broader bucket than GenAI-specific), so it is not comparable to Accenture's number.

Overall verdict For a startup or an investor, the consulting giants are mostly a CHANNEL — not roadkill, and only a conditional threat. What they're good at: scale, the trust to get through procurement and legal, managing the human change inside a big company, C-suite access, and — for Accenture — a balance sheet big enough to just buy anyone who threatens them. What they're structurally bad at: their whole business is billing for hours, not selling software at software margins, so they lose on speed, product depth, and increasingly on talent (the best deployment engineers keep leaving for labs and startups). Net: treat Accenture/Deloitte/BCG X as frenemies who both distribute you and compete for the build, EY/McKinsey as pure channels, IBM as locked inside its own stack. The threat only bites a startup that stays small long enough to get absorbed. The channel opportunity is available right now.
9 · Similarities vs. Differences — The Distillation

Where the category converges, where the real forks are, and what remains unknown

Having scored 22 companies across five axes and hunted hard for evidence, the payoff synthesis:

Where the whole category genuinely converges

Where the real strategic forks are

What the evidence (weakly or strongly) suggests works vs. fails

ClaimVerdictStrength
Buy-over-build converts to production fasterWorksReported — two independent sources (Menlo, MIT)
A verification layer is required for billable outputWorksReported — universal in practice + academic corpus
Deterministic substrate compounds; workflow-first ceilingsPlausibleConsensus-Unproven — architecture + Distyl's revealed pivot, no audited proof
FDE-embed drives production conversionConditionalInferred — true where substrate is complex; unproven as a universal
Outcome pricing is a superior modelOnly in CXReported — requires a countable unit
"95% of pilots fail"Directional onlyConsensus-Unproven — low-rigor source; pattern real, precision false

What remains genuinely unknown — and worth stating plainly

The honest limits of this analysis
  • Whether ontology-first actually out-compounds workflow-first in production is not knowable from public data. Production-conversion and compounding rates are undisclosed for both poles. This is the single most important diligence gap in the category.
  • Exactly one clean customer-triangulated outcome exists (AIG → Palantir + Anthropic), and even it has a seam — metrics and vendor-attribution sit in different AIG documents. Every other quantified claim is vendor-attested or attaches to a private buyer. Outside AIG, the category's outcome evidence is systematically non-triangulable.
  • Unit economics are entirely inferred. No rate cards; all pricing/margin structure is analytical inference, flagged throughout.
  • Distyl's Context Mesh depth is vendor-asserted only — no independent audit confirms it compounds in production at Palantir depth. Treat the convergence as "Distyl now claims Palantir-like architecture," not "Distyl has proven parity."

So, the whole thing in a paragraph. The deployment layer is real, big, and growing fast, and underneath the branding everyone is building the same machine: rent a frontier model, feed it the company's data in some organized way, check its work, deliver the result, send an invoice. The parts that are actually hard to copy are the boring ones — the plumbing that lets an agent safely change systems of record, the years of integration wiring, and the verification layer. The exciting-sounding differentiators — which model you use, “we only bill for outcomes,” armies of embedded engineers — are mostly either commoditizing, specific to customer service, or dictated by how messy the problem is rather than by any real strategy. And the secret the category doesn't advertise: almost nothing about “what works” can be proven from public data. Which is the real takeaway. In a market this foggy, being willing to say “we don't actually know, and here's exactly why” isn't a gap in the analysis — it's the most useful thing an investor can be handed.

Sources

  1. Palantir, Foundry Ontology — Overview / Core Concepts (digital-twin framing; semantic + kinetic elements; interfaces). palantir.com/docs/foundry/ontology/overview
  2. Palantir, Action Types & Functions Overview (governed writes as transactions). palantir.com/docs/foundry/action-types/overview
  3. Distyl AI, Technology — Context Mesh (persistent write-back knowledge graph; "each use case inherits everything"). distyl.ai/technology/context-mesh
  4. Distyl AI, Technology — Context Views ("the semantic layer"; Directed Review edits schema). distyl.ai/technology/context-views
  5. Zero Future Tech (primary-docs-derived), Palantir AIP Agent–Ontology Interaction (5-layer model). zerofuturetech.substack.com; corroborated by Palantir AIP features docs
  6. Distyl AI, NVIDIA Enterprise Agents ("living, traversable graph"; ReAct + deterministic pipeline; open Nemotron via NIM). distylai.substack.com/p/nvidia-enterprise-agents
  7. Palantir, AIP Bootcamp ("zero to use case in 5 days or less"). palantir.com/platforms/aip/bootcamp
  8. Palantir, Ontology Best Practices / Structural Guidance (compounding mechanism + duplication risk). palantir.com/docs/foundry/ontology/ontology-best-practices
  9. Sacra / Galileo prior verified fact sheet (June 2026), Distyl Distillery / Routines (SOP decomposition, auditable task chain).
  10. AgentMarketCap, Distyl AI $175M: Enterprise Knowledge Graphs & Agent Memory Layer ("applying Palantir's ontology thesis natively to LLM agents"). agentmarketcap.ai
  11. Cognition, Funding & product (Devin planner/executor; ~$400M @ $10.2B Sept 2025; >$1B Series D ~$26B May 2026). cognition.ai/blog; TechCrunch
  12. Happy Robot, Technical Overview / Voice AI (6-model pipeline; proprietary voice models; deterministic rule-guards; SIP/K8s). happyrobot.ai/blog/technical-overview
  13. Factory, GA announcement / product (model-agnostic per-task routing; no seat minimums). factory.ai/news/factory-is-ga
  14. Factory, Series C & Droids / Missions docs ($150M @ $1.5B, Apr 2026; role-scoped Droids; Review/Test verifier). factory.ai/news/series-c
  15. Harvey, Funding, FDE program & open-weight model ($200M @ $11B Mar 2026; "services business with software margins"; shipped an open-weight post-trained legal model, "Tenet"). harvey.ai/blog; harvey.ai/blog (Tenet)
  16. Cresta, AWS Marketplace listing & Sacra profile (per-seat + usage; 125K chats/yr at $150K; $125M Series D ~$1.6B). sacra.com/c/cresta-ai
  17. Glean, Series F & Enterprise Graph ($150M @ $7.2B Jun 2025; multi-model; per-seat). glean.com/press
  18. Cohere / Reuters, North platform, open weights & funding (trains own Command/Embed/Rerank; ships open weights under Apache-2.0 for self-host; private/air-gapped “GPU-in-a-closet” deploy; ~$6.8B val). Reuters; cohere.com/blog
  19. Menlo Ventures, 2025: The State of Generative AI in the Enterprise ($37B spend, 3.2× YoY; Anthropic 40%; 76% buy; 47% vs 25% conversion). menlovc.com
  20. MIT NANDA, The GenAI Divide: State of AI in Business 2025 (Jul 2025, "Preliminary Findings" v0.1; n=153 survey + 52 interviews + 300 initiatives; success = "remarked" P&L impact). Full PDF (mirror)
  21. ISG, State of Enterprise AI Adoption 2025 (n=1,200 use cases; 31% to production; $1.3M avg spend; 1-in-4 ROI; 50% efficiency). isg-one.com
  22. Gupta, A., ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress, arXiv 2601.06112. arxiv.org/abs/2601.06112
  23. Jia, J. et al., MAS-FIRE: Fault Injection & Reliability Evaluation for LLM-Based Multi-Agent Systems (silent semantic failures), arXiv 2602.19843. arxiv.org/abs/2602.19843
  24. Shlomov, S. et al. (IBM), From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production (CUGA), arXiv 2510.23856. arxiv.org/abs/2510.23856
  25. Palantir, Q1 2026 Letter to Shareholders (GE "using AIP to help increase engine production by 26%"). palantir.com/q1-2026-letter
  26. GE Aerospace, FY2025 Form 10-K (deliveries +26%, LEAP +28%, AI blade-inspection & material-assistant tools); SEC EDGAR full-text "Palantir" in GE filings = 0 hits. SEC 10-K
  27. Palantir, Q1 2026 Business Update (Associated Materials OTIF 40%→90%; private customer, no filing). investors.palantir.com
  28. CNBC, Goldman Sachs autonomous coder pilot (Devin); SEC EDGAR "Devin" in GS 10-K/10-Q = 0 hits. CNBC
  29. AIG & T-Mobile customer-side triangulation: AIG 2025 Annual Report (Lexington 370K+ submissions, ~+26% YoY, submit-to-bind ~+35–40%, claims hours→minutes) aig.com/investors; AIG Investor Day / earnings naming Palantir + Anthropic; 100% of private/non-profit submissions, $4B+ premium by 2030 target (Carrier Management, Apr 2025) carriermanagement.com; T-Mobile Q4-2025 Capital Markets Day (~$3B AI savings by 2027, ~60% chatbot containment; Distyl unnamed, credited “with OpenAI”) investor.t-mobile.com. Confirmed-Triangulated (AIG) / Partial (T-Mobile)
  30. Reuters, Cigna says AI tools save customers $200M over three years (2026-07-23; never names Sierra). Reuters; Sierra customers page: sierra.ai/customers
  31. Decagon / Cresta, Case studies (Rippling 38%→>50%; Cox +20%; private customers, no filings). decagon.ai; cresta.com/customer-stories
  32. Harvey, PwC UK customer story (saves 1–2 hrs per liquidation; private partnership). harvey.ai/blog
  33. Duolingo, FY2025 Form 10-K (generic "AI costs" in opex; no vendor, no deflection metric). SEC 10-K
  34. Center for Audit Quality, S&P 500 and AI Reporting (~90% mention AI in 2024 10-Ks). thecaq.org
  35. CFO Dive, Few 10-Ks tie AI to tangible revenue gains (85% Fortune 500 mention AI; 12% CEOs cite cost+revenue; 56% no benefit). cfodive.com
  36. Ramp, AI Index (card-transaction data, 50k+ businesses; ~4× monthly spend growth Feb'25→Feb'26). ramp.com/data/ai-index
  37. Accenture, Q4/FY2025 earnings + Palantir Business Group + M&A ($5.9B GenAI bookings; $2.7B advanced-AI revenue; RANGR, Decho acquisitions). Accenture earnings; Palantir group
  38. OpenAI, Frontier Alliance partners (McKinsey, BCG, Accenture). openai.com; McKinsey: mckinsey.com
  39. Anthropic, Deloitte–Anthropic partnership (Claude to ~470K staff — largest enterprise deployment); Deloitte–Palantir EOS: deloitte.com. anthropic.com
  40. EY, EY & 8090 launch EY.ai PDLC (Mar 2026; powered by 8090's Software Factory; performance claims are vendor projections). ey.com
  41. IBM, Q4 2025 results (GenAI book >$12.5B cumulative; ~$1.5B GenAI consulting Q3'25; consulting +3% YoY). newsroom.ibm.com
  42. BCG / Bloomberg, BCG says AI work brought 25% of 2025 revenue (~$3.6B AI-related); OpenAI Elite + Anthropic (2023). Bloomberg; anthropic.com/news/anthropic-bcg
  43. Evidently AI, Generative AI applications: 805 production use cases from 185 companies (production-only aggregator; measures claimed outcomes). evidentlyai.com
  44. Claritype, Company & product (ex-Palantir founders Rob Giardina + Ilya Lipkind; ontology/semantic-first; ~$6.6M seed 2022). claritype.com/company
  45. Sierra / TechCrunch, Funding & pay-per-resolution ($350M @ $10B Sept 2025; ~$950M @ ~$15.8B May 2026). TechCrunch
  46. Decagon / Reuters, Funding & per-resolution option ($131M @ $1.5B Jun 2025; $250M @ $4.5B Jan 2026). Reuters
  47. Northslope / Yahoo Finance, Funding & Palantir-native model (~$22M; Friends & Family Capital / ex-Palantir CFO Colin Anderson; AI app layer on AIP + Gemini). Yahoo Finance
  48. Happy Robot / Reuters, Funding ($44M Series B Sept 2025; $150M Series C ~$1.2B Aug 2026, freight/logistics voice). Reuters
  49. ZenML, LLMOps Database — what ~1,200 production deployments reveal (2025) (aggregated production case studies; “simple narrow workflows beat general-purpose agents,” “context engineering over prompt engineering”). zenml.io
  50. AIUseCaseHub, Enterprise AI deployment tracker (~3,814 public deployments across ~3,241 companies; filterable by maturity Production/Scaled Production and evidence strength; notes evidence signals ≠ operational audit). aiusecasehub.com/cases
  51. OpenAI, OpenAI launches the Deployment Company (forward-deployed engineers + deployment specialists; Tomoro acquisition adds ~150 FDEs subject to close). openai.com; WSJ (org/scope)
  52. Anthropic, Applied AI + Claude Partner Network (embedded Applied AI engineers for enterprise Claude deployments; $100M partner-network commitment, dedicated engineers + technical architects). anthropic.com/news/claude-partner-network; enterprise AI services
  53. joylarkin, Awesome-AI-Market-Maps (curated index of published AI market maps; most plot by stack layer or category, not by substrate-ownership × ontology). github.com/joylarkin/Awesome-AI-Market-Maps
  54. a16z, Your Data (Agents) Need Context (argues enterprise value comes from data/context gravity — permissions, metadata, workflow history — more than raw model quality). a16z.com/your-data-agents-need-context
  55. Kai Waehner, Enterprise Agentic AI Landscape 2026: Trust, Flexibility, and Vendor Lock-In (vendor analysis organized around lock-in / model-neutrality — the closest published lens to our data-gravity overlay). kai-waehner.de

Note on arXiv ids [33–35]: several Valency-corpus ids are forward-dated (2026 preprints); ids returned live by the corpus and HTTP-verified, but confirm each resolves before external citation. No source or metric in this report was fabricated; single-source and vendor-asserted claims are labeled as such throughout.

Generated by Galileo 🔭 · August 31, 2026 · Tomales Bay Capital