How to Test MCP Servers Before Giving Agents Real Authority
MCP Server Reliability Index 2027: Which Tools Can Agents Actually Depend On?
A server can be online while the agent's workflow fails. The DN framework measures usable outcomes, schema compatibility, recovery and permission enforcement, with failure gates that a high average cannot erase.
Decentralised News Research · Last verified 30 September 2026 · Methodology v1.0
What Matters
Judge an MCP server by accepted tool results within your workflow's deadline, not endpoint uptime alone. Test discovery, valid inputs, returned data, permission boundaries and recovery after disruption. DN's proposed index combines five measured dimensions with separate safety gates. This first edition provides a reproducible methodology and calculator; it does not claim live server rankings or independently measured performance.
DN Evidence Block
Verification: 30 September 2026. Source period: the published MCP 2026-07-28 specification, reviewed today. Scope: tools, transport, version compatibility and authorization. Decisive facts: tools have named interfaces and schemas; clients must handle protocol compatibility and validate access boundaries. Protocol conformance alone does not establish workflow success. Original DN contribution: five-dimensional index, task ledger, failure gates and local calculator. Test sample: no live endpoints tested in this edition. Reviewer: DN editorial desk; no independent technical reviewer. Methodology and evidence.
DN Alpha Thesis: the reliability unit is the accepted outcome
An agent's plan is a chain of dependencies. Discovery must expose the intended tool. Credentials must authorize the requested action. The call must use a compatible schema. The upstream service must respond. The result must contain the right information and arrive soon enough to matter. Reporting only an HTTP status hides most of that chain.
DN calls the difference between endpoint availability and useful completion the Tool Reliability Gap. A market-data server that returns a stale price is available but fails a time-sensitive trading task. A CRM tool that creates a record twice after a timeout has completed requests while violating the intended outcome. A successful read performed under revoked credentials crosses an authorization boundary.
The right benchmark holds the client, credential policy, task set, region and deadline constant. Otherwise the comparison measures different workloads and different failure opportunities.
The DN MCP Server Reliability Index
| Dimension | Weight | Measurement | Required evidence |
|---|---|---|---|
| Accepted tool outcomes | 40% | Original tasks completed correctly within deadline ÷ all eligible original tasks. | Request IDs, task acceptance rules, timestamps and validated outcomes. |
| Availability | 15% | Successful scheduled health/discovery probes ÷ scheduled probes. | Fixed probe schedule and observation window, including failures. |
| Schema compatibility | 15% | Passing declared schema regression cases ÷ executed cases. | Versioned inputs, tool definitions and acceptance fixtures. |
| Recovery | 15% | Interrupted tasks recovered to the correct final state within recovery deadline ÷ injected interruption cases. | Fault scenario, retries, final state and duplicate-action checks. |
| Authorization enforcement | 15% | Expected permission decisions correctly enforced ÷ tested decisions. | Expired, wrong-audience, revoked and insufficient-scope cases where applicable. |
Formula: DN-MRI = 0.40T + 0.15A + 0.15S + 0.15R + 0.15P, with each input expressed from 0 to 100. The weights are DN editorial choices, not requirements of the MCP specification.
Publish p50, p95 and p99 latency, first-attempt success, eventual success, retries, data freshness and cost alongside the headline. Score transport reliability and task correctness separately when possible. If a dimension was not tested, mark the composite incomplete; do not replace missing data with 100.
DN MCP Reliability Calculator
Enter measured percentages from one declared test cohort. The defaults are a fictional example, not observations of any server. Select “Untested” to see why a full score cannot be published.
The calculator cannot verify your percentages or detect missing incidents. Export includes the evidence label, entered values and gate outcome. It stores no data remotely.
The publication gates
The average score describes the sample. It should never authorize deployment by itself. DN proposes separate publication gates: any observed unauthorized action or duplicated consequential write blocks a “recommended for unattended actions” verdict pending remediation and retest. Unknown results remain unverified. Missing dimensions block the composite. A p95 above the declared latency budget blocks a “meets latency budget” label, even if eventual correctness is high.
A zero-failure sample is evidence about that sample, not proof that failure is impossible. Report sample size, task mix and observation period. One hundred reads and five writes should not be marketed as a reliable benchmark of payment execution.
A practical test plan
- Freeze the cohort: record endpoint or package version, MCP version, client, transport, region, authorization configuration and upstream dependencies.
- Declare acceptance: specify the correct result and deadline before each task runs. Separate reads, writes and money-moving actions.
- Record discovery: test pagination, tool names, declared schemas and changing tool lists under the intended credentials.
- Run baseline tasks: capture first-attempt outcome and eventual outcome with bounded retries.
- Inject controlled failures: interrupt the network, expire a token or return an upstream error in a staging environment.
- Inspect final state: establish whether a write executed before a retry. Check for duplicate records, orders or payments.
- Publish the ledger: make task IDs, timing, error category and acceptance decisions reproducible without exposing secrets or private customer data.
Do not perform destructive fault injection on a third party's production service. Use authorized staging systems and deterministic mock dependencies, then distinguish those tests from production observations.
Choose by workload, not one universal leaderboard
| Workflow | Most important evidence | Avoid if | Cost, access and custody |
|---|---|---|---|
| Read-only research | Correct sources, freshness, stable schema and accepted results. | Responses lack provenance or return stale material. | API/hosting charges and source access; usually no financial custody. |
| CRM or operations writes | Correct final state, duplicate suppression and attributable changes. | A timeout leads to blind re-execution. | Vendor access, write scopes and rework cost; holds business data rather than funds. |
| Trading or wallet actions | Permission enforcement, quote freshness, approval and execution reconciliation. | Authority is unbounded or executed actions cannot be reconciled. | Trading/network fees; custody depends on the connected exchange or wallet. |
| Long-running enterprise workflow | Recovery to correct state, credential lifecycle and version compatibility. | The workflow cannot resume safely after disruption. | Runtime, observability and support costs; account and geography restrictions vary. |
No vendors are ranked here because equivalent workload evidence is not yet available. A server listed in a registry has passed a discovery step, not this reliability test. Operational status and commercial eligibility should be checked before any future vendor recommendation.
Schema change and retries can create hidden costs
A tool can preserve its name while changing arguments or output meaning. Pin test fixtures to a version and rerun them when the schema changes. Contract tests should catch breaking semantics, not just whether JSON parses.
Retries make an eventual-success percentage look better while adding delay and cost. Group all attempts under the original task. For reads, a bounded retry may be reasonable. For consequential writes, inspect final state before repeating the action or use application-supported idempotency. A client-generated request ID alone does not guarantee the upstream system prevents duplicates.
What would prove DN's framework wrong?
The weighting should change if matched tests show another metric better predicts accepted completion or incident loss. A server with lower synthetic uptime might outperform one with higher uptime on the actual task cohort. If schema and recovery tests add no predictive value to a simple workload, simplify the test without claiming broader coverage. Publish counterexamples and retain the previous version so the methodology can improve.
Methodology, evidence and change log
Edition v1.0 is a proposed benchmark specification, not measured market data. DN reviewed the MCP 2026-07-28 tools, transport, versioning and authorization documentation on 30 September 2026. The dimensions, weights, gates and test sequence are DN's original editorial design. No live server, SDK or provider uptime was measured; the calculator's defaults are fictional. A production leaderboard would require declared cohorts, comparable credentials, repetition, independent validation and a disclosed test window.
- MCP tools, schemas and annotation trust.
- MCP transports.
- MCP versioning and compatibility.
- MCP authorization.
Update policy: review after published protocol changes and quarterly; measured cohorts must carry their own dates. Change log: 30 September 2026, initial framework and calculator. Submit corrections with evidence via DN Contact. Find eligible crypto platforms through DN Pathfinder.
Frequently asked questions
What is MCP server reliability?
The ability to expose usable tools and return correct, authorized results within a declared deadline, including during interruptions and recovery.
Does uptime prove that an MCP tool works?
No. A reachable endpoint can still return incorrect, stale or unusable results. Measure accepted tool outcomes separately.
Has DN ranked live MCP servers in this edition?
No. This edition publishes a methodology and an illustrative calculator. No independently tested vendor leaderboard is claimed.
What counts as a successful tool call?
A result that passes the task-specific acceptance rule and arrives within the declared deadline. A network or protocol success alone is insufficient.
Should retries count as new tasks?
No. Group attempts under the original request and record retry count and added cost. Report first-attempt and eventual success separately.
Are tool annotations a safety guarantee?
No. The MCP tools specification says annotations should be treated as untrusted unless they come from trusted servers. Verify actual behavior.
How much testing is enough?
There is no universal sample size. Start with a declared task set and record the observation window and number of attempts; expand testing for rare failures and consequential operations.
Can this calculator monitor a server?
No. It calculates from the metrics you enter. It does not contact endpoints or verify the measurements.
What stops a high score from hiding a serious fault?
DN proposes publication gates for unauthorized actions, duplicated consequential writes, missing evidence and latency-budget failures. Those gates are evaluated separately from the weighted score.