AI Context Efficiency 2027: Useful Outcomes per Token
Decentralised News · Batch 2, Article 34
The Agent Context Efficiency Index 2027: Is Your AI Paying to Read the Same Information Again?
A bigger context window creates room for more information. It does not establish that every repeated token helps the agent finish the task.
By Heath Muchena · Published 2 October 2026 · DN-CEI v1.0 · 2027 planning edition
What Matters
Agent context efficiency measures accepted outcomes against all input tokens consumed across a workflow, including repeated history and retries. Smaller context is useful only when the task still succeeds. DN proposes reporting accepted tasks per 100,000 input tokens beside completion rate, cost and control failures. This edition provides a matched-workload method and calculator, without claiming live model rankings or measured savings.
DN Evidence Block
Last verified: 2 October 2026. Review period: primary research and engineering guidance reviewed on that date. DN benchmark sample: zero live model runs. Sources: the 2024 Lost in the Middle paper and Anthropic’s September 2025 context-engineering guidance. Author: Heath Muchena; no independent reviewer recorded.
Decisive distinction: published evidence supports testing how context is used; it does not establish a universal optimum length or current vendor ranking. DN-CEI is a proposed editorial metric. Calculator defaults are fictional, not research results.
The DN Alpha Thesis: context has a carrying cost
An agent may request a small fact from a tool and then repeatedly carry that response through later model calls. Add the original instructions, previous outputs, tool descriptions and a growing conversation history, and the input consumed across the run can substantially exceed the unique information supplied by the user.
DN calls this the context carrying cost: the token exposure created by repeatedly making state available during a workflow. It is not automatically waste. Some repeated state preserves requirements and prevents mistakes. The research question is which information needs to remain active, which can be retrieved on demand and which can be replaced with a verified summary.
Measuring the largest individual prompt misses that question. A series of modest calls can consume more aggregate input than one large call. An apparent compression saving can disappear when the agent performs extra searches, retries or summary repairs.
This index therefore evaluates the entire declared workflow. It complements the Agent Model Efficiency Index, which focuses on successful work per dollar, and the memory benchmark, which tests retained state. Here the denominator is aggregate input-token consumption.
What the primary evidence tells us
Lost in the Middle evaluated multi-document question answering and key-value retrieval. In its studied settings, performance often depended on where relevant information appeared, with weaker use of information in the middle of the context. This historical result motivates position-sensitive tests; it does not prove identical behavior for every 2026 model or agent. [1]
Anthropic’s context-engineering guidance discusses curating instructions, tool information, retrieved data and history. It describes just-in-time retrieval, compaction and structured notes, while warning that aggressive compaction can discard important details. This is provider engineering guidance rather than an independent comparative benchmark. [2]
DN’s inference is practical: evaluate both removal and recovery. A compact representation is useful only if it preserves the information needed to meet the acceptance rule or provides a reliable way to retrieve it.
The DN Context Efficiency metric
| Metric | Formula | Why it matters |
|---|---|---|
| Context efficiency | Accepted tasks ÷ aggregate input tokens × 100,000 | Connects token consumption to independently accepted work |
| Acceptance rate | Accepted tasks ÷ attempted tasks × 100 | Exposes quality losses hidden by a smaller denominator |
| Input tokens per accepted task | Aggregate input tokens ÷ accepted tasks | Includes failed attempts and repeated state in the workload |
| Cost per accepted task | Total declared workflow cost ÷ accepted tasks | Tests whether token reduction translates into economic benefit |
| Matched-workload token reduction | 1 − candidate input tokens ÷ baseline input tokens | Describes consumption change, not proven avoidable waste |
Sum model input usage across every call inside the system boundary. Include input occurrences served from cache, counting each occurrence once; do not add a cached subset twice if already included in the total. Include compression, routing and summary calls when they are part of the evaluated system. Do not count an output token again as output here; count its later occurrence if it becomes input to another call.
Internal processing that is not exposed in usage cannot be reconstructed from this metric. Record the provider’s usage semantics and the included components. Monetary cost should come from actual billing records or clearly labelled estimates, rather than assuming all input tokens share one price.
DN Context Waste Calculator
This diagnostic compares two versions of the same task set. It measures consumption change, not proven semantic waste. Defaults are fictional. Equal attempted counts are required, and the reader must verify that the underlying tasks match.
Baseline
Candidate
Use one currency for both cost inputs. Include the same cost categories, task identities and acceptance rules in both runs. The tool does not run models, inspect prompts or verify evidence. Inputs remain local to the page.
A fictional improvement that still needs investigation
Suppose the baseline attempts 100 tasks, accepts 90 and consumes one million input tokens at a total cost of 20 currency units. A candidate attempts the same tasks, accepts 88 and consumes 500,000 input tokens at a cost of 12.
Baseline context efficiency is 9 accepted tasks per 100,000 input tokens. Candidate efficiency is 17.6. Token consumption falls 50%, and cost per accepted task falls from about 0.2222 to 0.1364. But acceptance also falls from 90% to 88%.
The candidate is more efficient by the proposed token metric and less successful by the task acceptance measure. Investigate the lost outcomes before deployment. An omitted contractual requirement or permission boundary can be more consequential than the saving. None of these numbers is an observed model result.
Compare context strategies on the same contract
| Strategy | Best for | Avoid if | Costs and access | Primary failure risk |
|---|---|---|---|---|
| Full-history context | Short workflows where earlier details remain relevant | History accumulates large irrelevant outputs | Repeated input, latency and context-window constraints | Relevant evidence becomes hard to identify |
| Targeted retrieval | Large source collections with reliable identifiers | Required evidence cannot be located reliably | Search/index maintenance, tool calls and retrieval latency | Missing decisive information or retrieving stale material |
| Compacted history | Long runs with preservable decisions and task state | Summaries cannot retain required detail | Summary calls, evaluation and recovery effort | Loss of exceptions, constraints or provenance |
| Structured state and notes | Tasks with clear fields, decisions and unresolved items | The state schema leaves essential information out | Storage, updates, versioning and validation | Stale or incorrectly written state |
No provider is ranked, promoted or assigned an operational status in this edition. Costs depend on the actual implementation and service terms. For agentic finance, preserve authorization and transaction state outside an unverified narrative summary; the context strategy itself does not establish custody or execution authority.
A test pack that exposes information loss
Choose a fixed set of tasks and define success before running them. Use an independent evaluator and preserve the original evidence. Keep model, tokenizer, tools, decoding settings and task order controlled where possible. Change one context strategy at a time.
| Fixture | Controlled change | Acceptance check |
|---|---|---|
| Position sensitivity | Move a necessary fact among beginning, middle and end positions | The correct answer remains grounded in the fact |
| Distractor load | Add plausible but irrelevant records | No false selection or unsupported claim |
| Compaction boundary | Summarize after introducing a rare exception | The exception survives and governs the final action |
| Stale evidence | Provide a dated value followed by an authoritative update | The final result uses the right version |
| State recovery | Clear history but preserve approved structured state | The agent resumes without repeating a completed action |
| Authority preservation | Reduce context containing action restrictions | Restrictions remain enforced; no unauthorized action |
Count all input consumed until final evaluation, not just the first request. Keep failed tasks in the denominator. Record median and tail latency separately; token reduction alone cannot establish faster delivery. Repeat enough runs to understand variability and publish sample counts rather than inventing confidence from a single clean demonstration.
What belongs in the evidence register
Each comparison needs task IDs, model and tokenizer versions, strategy configuration, source snapshots, per-call usage, cached-input treatment, total cost categories, acceptance outcomes and critical failures. Preserve the evaluator version and the exact system boundary.
Record how much context is submitted and which components are reintroduced: instructions, tool schemas, source evidence, tool results, history and structured state. That component ledger helps identify where to investigate. It does not prove that every repeated token is unnecessary.
Keep measurement and interpretation separate. “Input fell by 50%” can be a measured statement. “Half the context was waste” requires an additional causal argument showing that the removed information was unnecessary under the tested conditions.
Methodology, limitations and falsification
Method: review primary long-context research and provider engineering guidance; define accepted outcomes per aggregate input consumption; propose paired fixtures and quality gates; calculate diagnostics from reader-entered totals. No live model benchmark or vendor leaderboard has been performed.
Dataset status: the published fixture register is a proposed testing dataset, not observed performance data. The worked example is fictional. Future reports should publish sanitized task fixtures, per-call usage summaries, denominators and results on a permanent versioned URL.
Limits: tokenizers differ, task difficulty differs and stochastic outputs vary. Caching changes billing without necessarily changing logical input exposure. The metric does not measure semantic relevance, inaccessible internal compute or total infrastructure efficiency. Cross-model rankings need explicit caveats and additional outcome/cost measures.
Falsification test: reject a claim of useful context optimisation if matched evaluation finds lost necessary evidence, unacceptable quality decline, critical control failures or hidden auxiliary costs that reverse the claimed economic benefit.
Maintenance: rerun after model, tokenizer, tool, retrieval or compaction changes. Review sources quarterly. Change log: 2 October 2026, initial DN-CEI methodology and calculator. Corrections: send the claim, fixture and evidence through DN’s contact page.
Your next step: remove one source of repetition, then retest
Choose one repeated tool output or history segment. Test a targeted retrieval or verified summary alternative on a fixed task set. Compare token consumption, accepted outcomes, total cost and failure evidence together. Retain the original strategy until the comparison supports the change.
Connect workflow economics to the AI Agent ROI Index by Industry, and tool evidence to the MCP Server Reliability Index. Use DN Pathfinder when selecting a broader platform stack.
Sources
- Liu et al.: Lost in the Middle: How Language Models Use Long Contexts, Transactions of the Association for Computational Linguistics, 2024.
- Anthropic: Effective context engineering for AI agents, 29 September 2025.
Frequently asked questions
What is agent context efficiency?
DN defines it as accepted task outcomes per 100,000 aggregate model-input tokens across an evaluated workload. Acceptance rate, cost, latency and critical failures must be reported separately.
Is this a measured model leaderboard?
No. This edition publishes a proposed test method and a calculator using reader inputs. DN has not run a comparative model or context-strategy benchmark.
Should cached input tokens be counted?
Yes for the input-token denominator, if reported in provider usage. Count each input occurrence once and include cached occurrences. Use actual invoices for monetary cost; cached tokens can have different billing treatment.
Is a shorter context always better?
No. Removing necessary evidence, constraints or unresolved state can reduce correctness. Compare outcomes on matched tasks before treating token reduction as improvement.
Does the calculator measure redundant tokens?
No. It measures aggregate token reduction, completion rates and cost per accepted task. It cannot identify which tokens were unnecessary or establish avoidable waste.
Do retries belong in the denominator?
Yes. Include all model calls, retries and auxiliary context-management calls within the declared system boundary, including work on failed tasks.
How do you compare different tokenizers?
Prefer comparisons within the same model and tokenizer. Across models, token counts are not an identical unit of text; report that limitation and compare accepted outcomes, cost and latency as well.
What blocks a claimed improvement?
A critical safety or authorization failure blocks DN’s proposed recommendation. A drop in acceptance rate requires investigation even if accepted outcomes per token increases.
Related reading:
The Global Agent Opportunity Index 2027: Who Gets to Earn in the AI Agent Economy?
Your AI Agents Save Time. Are They Saving Money? DN AI Agent ROI Index by Industry 2027
Can Your Agents Work Together? The DN A2A Handoff Scorecard
9 Best AI Trading Tools for ChatGPT Users in 2027
Best AI Crypto Trading Tools for Beginners: 11 Platforms Compared
9 Best AI Trading Agents to Try With Paper Money Before Risking Real Cash
The 25 Best AI Trading Experiments to Try in 2027: From ChatGPT to Fully Autonomous Agents