Skip to content

Governed Metadata for AI Agents: What the Token Bleed Study Found

McKnight's 2026 benchmark found governed metadata for AI agents cut tokens up to 29.8x at perfect accuracy. See what it means for Agentforce and Data 360.

Governed Metadata for AI Agents: What the Token Bleed Study Found
Governed Metadata for AI Agents: What the Token Bleed Study Found

If your AI bill keeps climbing and your agents' answers still need checking, McKnight Consulting Group's August 2026 benchmark, "Stop the Token Bleed," offers an uncomfortable explanation: both problems share a cause. It compared three ways an agent can reach enterprise data and found that the route determined both what a question cost and whether the answer was right.

Governed metadata for AI agents is the study's answer: a catalog that has already classified which of millions of objects matter, queried by the agent instead of rebuilt on every question. In McKnight's runs, that choice cut token use by up to 29.8x and returned a perfect answer at every scale tested.

Quick Answer

 

Governed metadata for AI agents means giving an agent a pre-classified catalog of enterprise data that it queries as a tool, instead of scanning raw sources or reading a prompt stuffed with schemas. McKnight Consulting Group's August 2026 benchmark, which features Informatica Cloud Data Governance and Catalog (CDGC) from Informatica, a Salesforce company, found the governed route used 2,112 to 4,019 tokens at a perfect F1 of 1.000 across 300,000 to 15 million objects, while ungoverned routes used up to 29.8x more tokens and never scored above 0.66. It matters for CIOs, CFOs, data leaders, and Agentforce owners deciding whether a governed layer belongs in place before agents scale. Vantage Point designs and implements the Salesforce data foundation those agents depend on.

Key Takeaways (TL;DR)

  • Cost and correctness share a root cause. McKnight's benchmark found that an agent with no governed path to data re-derives which objects matter on every question: expensive, inconsistent, often wrong.
  • The numbers: governed access, 2,112 to 4,019 tokens at F1 1.000 at every tier; ungoverned routes, 16,563 to 67,089 tokens at F1 0.29 to 0.66 (input plus output tokens).
  • Bigger context windows did not help. Loading source metadata into the prompt produced the study's worst result, F1 0.294.
  • Read the caveats. The benchmark is vendor-featured and covers one task and one model; its authors name two conditions where the governed layer loses.
  • How Vantage Point helps: we design and implement the Salesforce data foundation that Agentforce agents depend on.

What is "token bleed" in enterprise AI?

"Token bleed" is the study's term for the soaring cost, latency, and context bloat caused by feeding raw, uncurated enterprise data into large language models. An agent that must scan millions of objects to find the right few is, the authors note, expensive at any token price.

The authors describe the pressure behind it: boards asking why AI costs keep climbing, CFOs fielding questions about a usage line growing faster than any other. Much of it, they argue, is "token maxing": treating usage as a proxy for adoption. "How do we spend fewer tokens?" is only half the question, they write; the real one is how to get correct answers with fewer tokens.

What did the benchmark actually test?

The study held everything constant except the route to the data: the same model, the same synthetic estate, and the same request, "Show me all the assets needed to build a report that involves Government ID." Only the access instruction changed:

  1. Direct database access. Live connection details; the agent scans and queries raw sources itself.
  2. Context stuffing. Extracted schema, table, and column definitions loaded into the context window, unorganized.
  3. Governed metadata layer. A query interface to Informatica CDGC, classified ahead of time, via a Model Context Protocol (MCP) server.

The estate was synthetic metadata modeled on a financial services enterprise; no customer data was used, and the catalog was populated by automated discovery with no human tuning.

Element What McKnight's benchmark used
Model claude-opus-4-8 in the Claude Code harness; 200K context; high effort
Scale 300K, 1.5M, 3M, and 15M objects across Snowflake, Oracle, and Amazon S3
Answer key Columns matching SSN, SUBJECT_SSN, ID_DOC_NUMBER; frozen before the catalog existed
Decoys TAX_ID, LICENSE_NO/DL_NUMBER, DOC_TYPE
Scoring Precision, recall, and F1 against the key; best-F1 run of 3 reported; tokens counted as input plus output, cache excluded

What did the benchmark find?

The table uses the study's input-plus-output token basis, which excludes cache reads and writes.

Tier Approach Total tokens vs. governed F1
300K Governed layer 2,711 – 1.000
300K Context stuffing 40,243 14.8x 0.656
300K Direct database access 34,500 12.7x 0.415
1.5M Governed layer 2,112 – 1.000
1.5M Context stuffing 62,957 29.8x 0.294
1.5M Direct database access 44,700 21.2x 0.592
3M Governed layer 2,805 – 1.000
3M Context stuffing 16,563 5.9x 0.522
3M Direct database access 29,199 10.4x 0.508
15M Governed layer 4,019 – 1.000
15M Context stuffing 24,569 6.1x 0.450
15M Direct database access 67,089 16.7x 0.549

The governed route stayed inside a 2,112 to 4,019 token band at F1 1.000 at all four tiers. The ungoverned routes swung between 16,563 and 67,089 tokens with "no stable relationship to data size" and never exceeded F1 0.66. At 15 million objects, direct access used 16.7x the governed route's tokens and returned fewer than half the correct columns. As the source grew 50x, governed cost grew 1.48x: the estate was understood once, at ingest.

The gap is almost entirely output tokens: 4,000 versus 66,321 at 15 million objects. Output is priced higher than input, so the dollar gap is wider than the token gap, and output volume measures how much work the agent did to convince itself it had an answer. In the study's words, the governed route did not think harder. It had less to think about.

Two different ways to be wrong

The two ungoverned routes failed in opposite directions.

Route At 15M objects Why it matters
Context stuffing over-returns 53,183 columns returned, 35,708 of them wrong An inventory that is two-thirds decoys cannot be acted on
Direct access under-returns 12,719 of 24,540 correct columns missed A short, clean list reads as authoritative, which the study calls the more dangerous compliance error

Both failures share one cause. Telling SSN from TAX_ID is not a reasoning problem; the model can do that. The hard part is knowing which of millions of objects to look at, and when nobody decides that ahead of time, the agent decides at run time, once per question.

These were not hallucinations. Every wrong pick was a real column in one of the three sources; the study found none fabricated. The fix is not a better model or a stricter prompt; it is a selection problem, and selection is what governance does.

Does a bigger context window fix it?

No. Context stuffing was included because it is the strongest objection to the governed approach: windows are large and getting cheaper, so why not put the schema in the prompt? Measured, it produced the benchmark's worst single result, F1 0.294 at 1.5 million objects.

A production-shaped scenario reinforced the point: a simulated 25-million-record airline estate across five source systems, queried for the assets behind a passenger contact report. On the input-plus-output basis, the governed route answered for 3,078 tokens at F1 1.000; direct access spent 82,533 tokens for F1 0.667. A context-stuffing run pointed at the schema and allowed to iterate nearly doubled its spend to 19,247 tokens and found fewer attributes. Under the study's full accounting, which adds cache reads and writes, the governed run cost 137,778 tokens against 12.2 million for direct access, widening the gap from 26.8x to 88.8x.

What does this mean for Salesforce customers?

Informatica is now part of Salesforce. The catalog this benchmark features sits in the Salesforce portfolio alongside Salesforce Data 360 (formerly Data Cloud) and MuleSoft. The study makes no claim about how Agentforce or Data 360 use CDGC internally, and neither do we. It does make governed metadata a data-foundation question, one we examined in our analysis of the Informatica acquisition's implications for AI-driven enterprise solutions.

The pattern generalizes. What won was an architecture, not a brand: a governed retrieval layer, classified ahead of time, that the agent queries as a tool. MCP was the mechanism, and Agentforce can use the same protocol. Our guide to how Agentforce uses MCP to connect external systems covers the wiring; this study covers why what sits at the other end should be governed.

Agents inherit the foundation. An Agentforce agent answering from Data 360 and an external estate is doing discovery whether or not anyone measures it. That puts the data foundation ahead of agent count on the roadmap.

Where governed metadata does not help (read this before you buy)

The authors write that "a benchmark that only reports where the sponsor's architecture wins is marketing," and publish two conditions under which the governed approach loses.

  • Freshness is load-bearing. The catalog measured had just been ingested. A governed layer six months stale will route an agent confidently toward objects that no longer exist, and cheaply. "Cheap and wrong is not an improvement over expensive and wrong."
  • Classifier coverage is the ceiling. The catalog scored 1.000 because its out-of-the-box classifiers recognized the identifier patterns present. Against a proprietary identifier scheme with no matching classifier, the governed route inherits the same selection problem.

Vantage Point adds its own caveats: the benchmark features Informatica's product, and Informatica is a Salesforce company; it tests one task type (data discovery) with one model (claude-opus-4-8 in the Claude Code harness) on synthetic metadata; and it reports the best-F1 run of three per cell. Treat the ratios as directional and the method (a frozen answer key, planted decoys, precision and recall per route) as the reusable part.

What should your team do next?

  1. Inventory where agents touch enterprise data. List every agent and copilot that reads outside its own system, and how each finds the objects it uses.
  2. Run the afternoon benchmark. Write the answer key by hand, plant three or four look-alike columns, run your current setup, and score precision and recall. Our companion post, How to Benchmark AI Agent Token Cost and Accuracy in an Afternoon, walks through it.
  3. Decide the governed layer deliberately. Whatever the catalog, set a freshness SLA and review classifier coverage against your own identifier schemes first.
  4. Budget on predictability, not just price. The study's most quotable line is a finance line: "A cost that cannot be forecasted cannot be budgeted or capped."
  5. Put discovery accuracy on the agent scorecard. Most teams, the study observes, have measured only that their agents return something.

How Vantage Point Helps

Vantage Point designs and implements the Salesforce data foundation that agents depend on: a data-foundation assessment and design across Salesforce Data 360, MuleSoft, and Informatica-fed metadata, so the governed layer exists before the first agent queries it; Agentforce implementation with governed data access, where discovery is measured rather than assumed; and managed services that keep the foundation current, because a stale catalog fails cheaply. Senior consultants only — no junior handoffs; the experts you meet are the experts who deliver. To begin, start a conversation with our team.

Ready to Find Out What Your Agents Actually Return?

 

An afternoon benchmark and a data-foundation review will show whether your agents find the right objects or re-derive them on every question. Talk to Vantage Point about governed data access for Agentforce, or explore our Salesforce services.

Frequently Asked Questions

What is token bleed in enterprise AI?

"Token bleed" is McKnight Consulting Group's term for the soaring cost, latency, and context bloat caused by feeding raw, uncurated enterprise data into large language models.

What is a governed metadata layer for AI agents?

A governed metadata layer is a catalog of enterprise data assets, scanned and classified ahead of time, that an AI agent queries as a tool. In McKnight's benchmark it was Informatica Cloud Data Governance and Catalog (CDGC), reached through a Model Context Protocol (MCP) server.

Does this apply to Agentforce and Salesforce Data 360?

The study did not test Agentforce or Data 360 and makes no claim about how either uses CDGC internally. Its finding is architectural: an agent querying a pre-classified, governed layer as a tool beat one scanning raw sources or reading a stuffed prompt on both cost and accuracy, and that applies to any agent reaching enterprise data.

Were the wrong answers hallucinations?

No. Every wrong pick was a real column in one of the sources; the study found no hallucinated identifiers. The failures were selection errors, returning look-alikes such as TAX_ID or missing correct columns, which is why the fix is governance rather than a different model.

How reliable is a vendor-featured benchmark?

Treat it as directional. The study features Informatica CDGC, and Informatica is a Salesforce company; its authors acknowledge this and publish two conditions where the governed approach loses. It also covers one task, one model, and synthetic metadata, and reports the best of three runs. The method transfers; the exact ratios may not.

What is the fastest way to test this in my own environment?

Follow the study's recipe: write the answer key by hand before building anything, plant three or four look-alike columns, then run your current agent setup while capturing tokens and scoring precision and recall. No production data is required.

Sources


Vantage Point is a boutique CRM consulting firm helping businesses transform with Salesforce, HubSpot, and AI — 150+ clients, 400+ engagements, and a 4.71/5 average engagement rating. Learn more at vantagepoint.io.

David Cockrum

David Cockrum

David Cockrum is the founder and CEO of Vantage Point, a specialized Salesforce consultancy exclusively serving financial services organizations. As a former Chief Operating Officer in the financial services industry with over 13 years as a Salesforce user, David recognized the unique technology challenges facing banks, wealth management firms, insurers, and fintech companies—and created Vantage Point to bridge the gap between powerful CRM platforms and industry-specific needs. Under David’s leadership, Vantage Point has achieved over 150 clients, 400+ completed engagements, a 4.71/5 client satisfaction rating, and 95% client retention. His commitment to Ownership Mentality, Collaborative Partnership, Tenacious Execution, and Humble Confidence drives the company’s high-touch, results-oriented approach, delivering measurable improvements in operational efficiency, compliance, and client relationships. David’s previous experience includes founder and CEO of Cockrum Consulting, LLC, and consulting roles at Hitachi Consulting. He holds a B.B.A. from Southern Methodist University’s Cox School of Business.

Elements Image

Subscribe to our Blog

Get the latest articles and exclusive content delivered straight to your inbox. Join our community today—simply enter your email below!

Need help applying this to your CRM roadmap?

Talk to Vantage Point

Vantage Point helps regulated and growth-focused teams implement Salesforce, HubSpot, integrations, data migration, and managed services with practical, senior-led guidance.

Latest Articles

Salesforce Marketing Cloud Next Consultant Exam: What's on It and Who Should Take It

Salesforce Marketing Cloud Next Consultant Exam: What's on It and Who Should Take It

Salesforce launched the Marketing Cloud Next Consultant certification. See the official exam outline, format, fee, prerequisites, and who s...

Claudeforce Requires Premium Edition: The Licensing Decision

Claudeforce Requires Premium Edition: The Licensing Decision

Claudeforce requires Salesforce's premium edition, and Milano says only 5% of Sales and Service Cloud seats have it. Here's how to budget f...

Salesforce Outage September 2026: What Admins Should Do Now

Salesforce Outage September 2026: What Admins Should Do Now

A global Salesforce outage on September 16, 2026 disrupted all regions. Get the status timeline, admin actions, and post-incident checks.