
If your AI bill keeps climbing and your agents' answers still need checking, McKnight Consulting Group's August 2026 benchmark, "Stop the Token Bleed," offers an uncomfortable explanation: both problems share a cause. It compared three ways an agent can reach enterprise data and found that the route determined both what a question cost and whether the answer was right.
Governed metadata for AI agents is the study's answer: a catalog that has already classified which of millions of objects matter, queried by the agent instead of rebuilt on every question. In McKnight's runs, that choice cut token use by up to 29.8x and returned a perfect answer at every scale tested.
Quick Answer
Governed metadata for AI agents means giving an agent a pre-classified catalog of enterprise data that it queries as a tool, instead of scanning raw sources or reading a prompt stuffed with schemas. McKnight Consulting Group's August 2026 benchmark, which features Informatica Cloud Data Governance and Catalog (CDGC) from Informatica, a Salesforce company, found the governed route used 2,112 to 4,019 tokens at a perfect F1 of 1.000 across 300,000 to 15 million objects, while ungoverned routes used up to 29.8x more tokens and never scored above 0.66. It matters for CIOs, CFOs, data leaders, and Agentforce owners deciding whether a governed layer belongs in place before agents scale. Vantage Point designs and implements the Salesforce data foundation those agents depend on.
Key Takeaways (TL;DR)
- Cost and correctness share a root cause. McKnight's benchmark found that an agent with no governed path to data re-derives which objects matter on every question: expensive, inconsistent, often wrong.
- The numbers: governed access, 2,112 to 4,019 tokens at F1 1.000 at every tier; ungoverned routes, 16,563 to 67,089 tokens at F1 0.29 to 0.66 (input plus output tokens).
- Bigger context windows did not help. Loading source metadata into the prompt produced the study's worst result, F1 0.294.
- Read the caveats. The benchmark is vendor-featured and covers one task and one model; its authors name two conditions where the governed layer loses.
- How Vantage Point helps: we design and implement the Salesforce data foundation that Agentforce agents depend on.
What is "token bleed" in enterprise AI?
"Token bleed" is the study's term for the soaring cost, latency, and context bloat caused by feeding raw, uncurated enterprise data into large language models. An agent that must scan millions of objects to find the right few is, the authors note, expensive at any token price.
The authors describe the pressure behind it: boards asking why AI costs keep climbing, CFOs fielding questions about a usage line growing faster than any other. Much of it, they argue, is "token maxing": treating usage as a proxy for adoption. "How do we spend fewer tokens?" is only half the question, they write; the real one is how to get correct answers with fewer tokens.
What did the benchmark actually test?
The study held everything constant except the route to the data: the same model, the same synthetic estate, and the same request, "Show me all the assets needed to build a report that involves Government ID." Only the access instruction changed:
- Direct database access. Live connection details; the agent scans and queries raw sources itself.
- Context stuffing. Extracted schema, table, and column definitions loaded into the context window, unorganized.
- Governed metadata layer. A query interface to Informatica CDGC, classified ahead of time, via a Model Context Protocol (MCP) server.
The estate was synthetic metadata modeled on a financial services enterprise; no customer data was used, and the catalog was populated by automated discovery with no human tuning.
| Element | What McKnight's benchmark used |
|---|---|
| Model | claude-opus-4-8 in the Claude Code harness; 200K context; high effort |
| Scale | 300K, 1.5M, 3M, and 15M objects across Snowflake, Oracle, and Amazon S3 |
| Answer key | Columns matching SSN, SUBJECT_SSN, ID_DOC_NUMBER; frozen before the catalog existed |
| Decoys | TAX_ID, LICENSE_NO/DL_NUMBER, DOC_TYPE |
| Scoring | Precision, recall, and F1 against the key; best-F1 run of 3 reported; tokens counted as input plus output, cache excluded |
What did the benchmark find?
The table uses the study's input-plus-output token basis, which excludes cache reads and writes.
| Tier | Approach | Total tokens | vs. governed | F1 |
|---|---|---|---|---|
| 300K | Governed layer | 2,711 | – | 1.000 |
| 300K | Context stuffing | 40,243 | 14.8x | 0.656 |
| 300K | Direct database access | 34,500 | 12.7x | 0.415 |
| 1.5M | Governed layer | 2,112 | – | 1.000 |
| 1.5M | Context stuffing | 62,957 | 29.8x | 0.294 |
| 1.5M | Direct database access | 44,700 | 21.2x | 0.592 |
| 3M | Governed layer | 2,805 | – | 1.000 |
| 3M | Context stuffing | 16,563 | 5.9x | 0.522 |
| 3M | Direct database access | 29,199 | 10.4x | 0.508 |
| 15M | Governed layer | 4,019 | – | 1.000 |
| 15M | Context stuffing | 24,569 | 6.1x | 0.450 |
| 15M | Direct database access | 67,089 | 16.7x | 0.549 |
The governed route stayed inside a 2,112 to 4,019 token band at F1 1.000 at all four tiers. The ungoverned routes swung between 16,563 and 67,089 tokens with "no stable relationship to data size" and never exceeded F1 0.66. At 15 million objects, direct access used 16.7x the governed route's tokens and returned fewer than half the correct columns. As the source grew 50x, governed cost grew 1.48x: the estate was understood once, at ingest.
The gap is almost entirely output tokens: 4,000 versus 66,321 at 15 million objects. Output is priced higher than input, so the dollar gap is wider than the token gap, and output volume measures how much work the agent did to convince itself it had an answer. In the study's words, the governed route did not think harder. It had less to think about.
Two different ways to be wrong
The two ungoverned routes failed in opposite directions.
| Route | At 15M objects | Why it matters |
|---|---|---|
| Context stuffing over-returns | 53,183 columns returned, 35,708 of them wrong | An inventory that is two-thirds decoys cannot be acted on |
| Direct access under-returns | 12,719 of 24,540 correct columns missed | A short, clean list reads as authoritative, which the study calls the more dangerous compliance error |
Both failures share one cause. Telling SSN from TAX_ID is not a reasoning problem; the model can do that. The hard part is knowing which of millions of objects to look at, and when nobody decides that ahead of time, the agent decides at run time, once per question.
These were not hallucinations. Every wrong pick was a real column in one of the three sources; the study found none fabricated. The fix is not a better model or a stricter prompt; it is a selection problem, and selection is what governance does.
Does a bigger context window fix it?
No. Context stuffing was included because it is the strongest objection to the governed approach: windows are large and getting cheaper, so why not put the schema in the prompt? Measured, it produced the benchmark's worst single result, F1 0.294 at 1.5 million objects.
A production-shaped scenario reinforced the point: a simulated 25-million-record airline estate across five source systems, queried for the assets behind a passenger contact report. On the input-plus-output basis, the governed route answered for 3,078 tokens at F1 1.000; direct access spent 82,533 tokens for F1 0.667. A context-stuffing run pointed at the schema and allowed to iterate nearly doubled its spend to 19,247 tokens and found fewer attributes. Under the study's full accounting, which adds cache reads and writes, the governed run cost 137,778 tokens against 12.2 million for direct access, widening the gap from 26.8x to 88.8x.
What does this mean for Salesforce customers?
Informatica is now part of Salesforce. The catalog this benchmark features sits in the Salesforce portfolio alongside Salesforce Data 360 (formerly Data Cloud) and MuleSoft. The study makes no claim about how Agentforce or Data 360 use CDGC internally, and neither do we. It does make governed metadata a data-foundation question, one we examined in our analysis of the Informatica acquisition's implications for AI-driven enterprise solutions.
The pattern generalizes. What won was an architecture, not a brand: a governed retrieval layer, classified ahead of time, that the agent queries as a tool. MCP was the mechanism, and Agentforce can use the same protocol. Our guide to how Agentforce uses MCP to connect external systems covers the wiring; this study covers why what sits at the other end should be governed.
Agents inherit the foundation. An Agentforce agent answering from Data 360 and an external estate is doing discovery whether or not anyone measures it. That puts the data foundation ahead of agent count on the roadmap.
Where governed metadata does not help (read this before you buy)
The authors write that "a benchmark that only reports where the sponsor's architecture wins is marketing," and publish two conditions under which the governed approach loses.
- Freshness is load-bearing. The catalog measured had just been ingested. A governed layer six months stale will route an agent confidently toward objects that no longer exist, and cheaply. "Cheap and wrong is not an improvement over expensive and wrong."
- Classifier coverage is the ceiling. The catalog scored 1.000 because its out-of-the-box classifiers recognized the identifier patterns present. Against a proprietary identifier scheme with no matching classifier, the governed route inherits the same selection problem.
Vantage Point adds its own caveats: the benchmark features Informatica's product, and Informatica is a Salesforce company; it tests one task type (data discovery) with one model (claude-opus-4-8 in the Claude Code harness) on synthetic metadata; and it reports the best-F1 run of three per cell. Treat the ratios as directional and the method (a frozen answer key, planted decoys, precision and recall per route) as the reusable part.
What should your team do next?
- Inventory where agents touch enterprise data. List every agent and copilot that reads outside its own system, and how each finds the objects it uses.
- Run the afternoon benchmark. Write the answer key by hand, plant three or four look-alike columns, run your current setup, and score precision and recall. Our companion post, How to Benchmark AI Agent Token Cost and Accuracy in an Afternoon, walks through it.
- Decide the governed layer deliberately. Whatever the catalog, set a freshness SLA and review classifier coverage against your own identifier schemes first.
- Budget on predictability, not just price. The study's most quotable line is a finance line: "A cost that cannot be forecasted cannot be budgeted or capped."
- Put discovery accuracy on the agent scorecard. Most teams, the study observes, have measured only that their agents return something.
How Vantage Point Helps
Vantage Point designs and implements the Salesforce data foundation that agents depend on: a data-foundation assessment and design across Salesforce Data 360, MuleSoft, and Informatica-fed metadata, so the governed layer exists before the first agent queries it; Agentforce implementation with governed data access, where discovery is measured rather than assumed; and managed services that keep the foundation current, because a stale catalog fails cheaply. Senior consultants only — no junior handoffs; the experts you meet are the experts who deliver. To begin, start a conversation with our team.
Ready to Find Out What Your Agents Actually Return?
An afternoon benchmark and a data-foundation review will show whether your agents find the right objects or re-derive them on every question. Talk to Vantage Point about governed data access for Agentforce, or explore our Salesforce services.
Frequently Asked Questions
What is token bleed in enterprise AI?
"Token bleed" is McKnight Consulting Group's term for the soaring cost, latency, and context bloat caused by feeding raw, uncurated enterprise data into large language models.
What is a governed metadata layer for AI agents?
A governed metadata layer is a catalog of enterprise data assets, scanned and classified ahead of time, that an AI agent queries as a tool. In McKnight's benchmark it was Informatica Cloud Data Governance and Catalog (CDGC), reached through a Model Context Protocol (MCP) server.
Does this apply to Agentforce and Salesforce Data 360?
The study did not test Agentforce or Data 360 and makes no claim about how either uses CDGC internally. Its finding is architectural: an agent querying a pre-classified, governed layer as a tool beat one scanning raw sources or reading a stuffed prompt on both cost and accuracy, and that applies to any agent reaching enterprise data.
Were the wrong answers hallucinations?
No. Every wrong pick was a real column in one of the sources; the study found no hallucinated identifiers. The failures were selection errors, returning look-alikes such as TAX_ID or missing correct columns, which is why the fix is governance rather than a different model.
How reliable is a vendor-featured benchmark?
Treat it as directional. The study features Informatica CDGC, and Informatica is a Salesforce company; its authors acknowledge this and publish two conditions where the governed approach loses. It also covers one task, one model, and synthetic metadata, and reports the best of three runs. The method transfers; the exact ratios may not.
What is the fastest way to test this in my own environment?
Follow the study's recipe: write the answer key by hand before building anything, plant three or four look-alike columns, then run your current agent setup while capturing tokens and scoring precision and recall. No production data is required.
Sources
- Stop the Token Bleed: Benchmarking the Benefits of Governed Metadata for Enterprise AI (McKnight Consulting Group, August 2026)
- Informatica Cloud Data Governance and Catalog (CDGC)
- Salesforce's Informatica Acquisition and AI-Driven Enterprise Solutions (Vantage Point)
- How Agentforce Uses MCP to Connect External Systems (Vantage Point)
Vantage Point is a boutique CRM consulting firm helping businesses transform with Salesforce, HubSpot, and AI — 150+ clients, 400+ engagements, and a 4.71/5 average engagement rating. Learn more at vantagepoint.io.
