We Tested Five AI Finance Systems. Here Is What the Evidence Actually Proves.
DN’s Agent Trust dataset makes authorization, recovery and execution evidence inspectable. Its most important result is the boundary between a passing test and justified financial authority.
By Heath Muchena · Research snapshot: 11 October 2026 · Agent Runtime Evidence Dataset v0.1 · Risk Graph v1.6
- DN now tracks five financial-agent systems through nine recurring surveillance feeds, with exact upstream commit pins and retained observation history.
- All nine feeds passed in the launch snapshot. These are bounded test results, not production safety certifications.
- Almanak holds a provisional E3 exposure class. Coinbase AgentKit, Definitive Flash, Liquid Co-Invest and Virtuals ACP remain unclassified.
- Passing surveillance does not upgrade a class. Missing, stale or failed observations cannot support a Verified badge.
The decisive question for an AI finance system is not whether it can describe a trade. It is whether its authority is bounded when a trade, payment or tool call goes wrong.
A model can follow an instruction while the surrounding software mishandles a retry. A trading client can generate an idempotency key while the server fails to enforce it. An agent can stop issuing new orders while an existing position remains open. Each distinction changes what a passing demonstration actually establishes.
DN’s new Agent Trust release starts with those boundaries. The first cohort comprises Almanak, Coinbase AgentKit, Definitive Flash, Liquid Co-Invest and Virtuals EconomyOS / Agent Commerce Protocol. The public comparison connects versioned profiles, recurring tests, source observations, verification responses and embeddable badges.
This is a financial-agent evidence cohort. It is separate from DN’s managed-runtime comparison of cloud infrastructure and self-hosted frameworks. It does not rank model intelligence, general runtime portability or trading returns.
Evidence has a scope, a version and an expiry.
A useful trust signal must identify what ran, what it tested, which software version it exercised and what it did not establish. A green status without those boundaries is easy to overinterpret.
DN’s working principle is therefore narrow: publish the evidence that exists, preserve the uncertainty that remains, and stop presenting old or incomplete observations as current verification.
Observation period: the recurring observations in this launch snapshot were recorded on 11 October 2026 between 08:13 and 09:42 UTC. The dataset was retrieved at 2026-10-11T10:04:24.408Z. Individual feeds retain their precise timestamps.
Evidence surface: selected upstream unit and boundary suites, mocked or fake-client interactions, and Almanak managed-fork recovery and position-containment scenarios. No funded production trading or production security certification is claimed.
Version discipline: each recurring observation records an upstream commit, DN workflow run and stated limitations. A pinned test result does not establish that the newest upstream release, a hosted deployment or a reader’s configuration behaves identically.
Release distinction: Dataset v0.1 is the recurring-evidence comparison interface; Risk Graph v1.6 is the underlying profile dataset. Five systems and nine feeds describe this launch cohort, not the entire market.
The first five systems: what DN can and cannot conclude
The comparison is alphabetical. Different suites exercise different surfaces, so test counts, control counts and confidence scores do not support a safest-to-least-safe ranking. DN’s evidence-confidence figures are editorial assessment scores out of 100, not statistical probabilities of safety or measured percentages of system coverage.
| System | Exposure class at launch | Executed evidence | Material boundary |
|---|---|---|---|
| Almanak | E3, provisional | Policy and gateway checks plus managed-fork recovery and containment | Production-mainnet operation, complete live exposure inputs, scoped gateway authorization and broader owner recovery remain unproven. |
| Coinbase AgentKit | Unclassified | Selected capability, schema and failure-boundary unit tests | Live wallet execution, principal revocation, spending limits and production transaction containment remain unproven. |
| Definitive Flash | Unclassified | Selected signed-order validation and simulated submission reconciliation | Production API authorization, live custody and signer revocation, venue behavior and live limit enforcement remain unproven. |
| Liquid Co-Invest / Co-Invest Computer | Unclassified | Local MCP write suppression, auth propagation and client idempotency-key behavior | Private backend enforcement, duplicate-order prevention, live limits and token revocation remain unproven. |
| Virtuals EconomyOS / ACP | Unclassified | Selected commerce-phase, payment-structure and event-identity unit tests | Live escrow settlement, dispute enforcement, production wallet authority and recovery remain unproven. |
Almanak: stronger execution evidence, still provisional
Almanak has the broadest tested financial-control surface in the initial cohort. Its recurring core program exercises allocation limits, aggregate exposure and gateway authentication. Its monthly deep program tests signer recovery and open-position containment on a managed Arbitrum fork.
That distinction matters. Revoking a signer and closing an existing position are different control problems. The former restricts future authority; the latter addresses exposure already created. Evidence for one cannot stand in for the other.
The public classification remains E3, provisional. Current passing feeds support the label “E3 Provisional · Current,” not an unconditional Verified badge. Mainnet operation, complete live exposure inputs, more granular gateway authorization and broader owner recovery remain outside the demonstrated boundary.
Coinbase AgentKit: capability checks are not spending controls
DN’s selected AgentKit tests exercise protocol-family gating, unsupported-network and unknown-token failures, invalid-address schema rejection, transaction failure propagation and a wallet-provider signer surface. These are deterministic upstream unit tests with mocks.
They provide useful evidence about specific rejection and failure paths. They do not establish live delegated-wallet revocation, enforced spending limits, compromised-key recovery or production transaction containment. AgentKit remains unclassified.
Definitive Flash: signed execution still needs state certainty
The selected Flash evidence covers signed-order field validation, separation of EVM and SVM field families, pairing of Permit2 typed data and signatures, and simulated handling of ambiguous submission and retry behavior.
A signature boundary and a reconciliation boundary answer different questions. One concerns the submitted authorization material; the other concerns whether the client can determine what happened after submission. Fake-client tests can exercise these paths without proving production API authorization, venue behavior or live cancellation and limits. Flash remains unclassified.
Liquid Co-Invest: the client cannot prove the server
The Liquid benchmark exercises local MCP write suppression, WRITE capability classification, token propagation, HTTP failure surfacing, catalog filtering and per-write idempotency-key behavior.
The key caveat is precise: generating a client idempotency key does not prove that a private backend prevents duplicate orders. That requires evidence from the enforcement boundary itself. DN has not established live order reconciliation, server-side authorization scope, position limits, cancel/close behavior or token revocation latency. Liquid remains unclassified.
Virtuals ACP: commerce state is not settlement assurance
The ACP evidence exercises commerce-phase gates, x402 payment structure and nonce requirements, failure handling, entity-bound signature packing and job-event identity reconciliation through selected deterministic unit suites.
Correct phase transitions in tests do not establish live escrow funding and release, effective dispute enforcement, production wallet authority, replay prevention in the deployed payment path or smart-wallet recovery. Those remain separate research questions. Virtuals remains unclassified.
DN Agent Evidence Boundary Explorer
Select a system to inspect its recurring feeds, source pins and stated limitations. The scenario control illustrates how evidence can lose current status; it does not change DN’s public dataset or simulate a financial loss.
How to read an exposure class without turning it into a guarantee
DN’s classes describe the authority its available evidence can justify under stated conditions. They run from E0 Observe Only to E4 High-Assurance Autonomous Counterparty. They are not credit ratings, return forecasts or universal deployment approvals.
An unclassified profile is an explicit evidence gap, not a finding that the product is unsafe. A provisional class is also a real constraint: repeated passes do not silently remove the qualification. Stronger evidence and a classification review would be needed.
The recurring-state calculation evaluates each manifest-declared feed. A passing weekly observation becomes stale after eight days; a monthly observation after 35 days. Missing observations are pending. Non-passing observations require review, and invalid or future timestamps are flagged. The system summary preserves the most restrictive feed state.
A failed run may reflect an environment or public-RPC problem rather than a product defect. That distinction requires diagnosis. Until then, the evidence cannot honestly support an unrestricted current-verification claim.
The economic implication: assurance must survive software change
DN’s original thesis is that financial-agent adoption will increasingly depend on the continuity of control evidence. A buyer considering persistent financial authority needs more than a convincing demo or a one-time evaluation.
We call the resulting uncertainty the Evidence Expiry Discount: the reduction in decision value when an observation no longer matches the deployed version, configuration or relevant time window. This is a qualitative research concept, not a calculated valuation factor in the dataset.
There are two clocks. One measures how recently a benchmark ran. The other measures whether the benchmark’s pinned software still corresponds to the deployment under consideration. A fresh run on an old commit resets the first clock, not the second. DN’s pinned surveillance deliberately makes that limitation visible.
Three developments would support this thesis: procurement teams asking for version-linked control records; integration teams requiring verified recovery paths before expanding permissions; and service providers making recurring evidence part of deployment documentation. These are predictions to test, not outcomes demonstrated by this five-system release.
The counterargument is substantial. Unit tests can be inexpensive to pass, public repositories may omit the real enforcement layer, and a monitoring product can accumulate impressive-looking data without reducing operational risk. The answer is to expand the tested boundary, not inflate the score.
The valuable trust product is a maintained evidence trail.
The opportunity is to make changes in authority, failure handling and recovery visible to people and machines. The defensible asset is the observation history, its version discipline and its explicit unknowns.
This release is a starting cohort. Its value will grow if DN can demonstrate where controls regress, which gaps close and whether stronger production evidence changes a classification. Adding names without adding proof would weaken it.
Use the release as an evidence interface
- Open the profile. Check the actual test surface and class qualification before interpreting a badge.
- Inspect each observation. Match its upstream commit, timestamp, workflow run and limitations to the claim you care about.
- Compare the deployed configuration. A public SDK result may not transfer to a private backend, hosted service or modified integration.
- Keep unknowns explicit. Missing spending limits, recovery or settlement evidence should remain missing in your own assessment.
- Recheck after change. Version advancement and configuration changes require renewed validation of the relevant controls.
Developers can consume the JSON dataset, download the CSV comparison or query a system’s verification response. SVG badges and trust cards are available from individual profiles. Preserve the qualification and source link when displaying them; “unclassified” or “provisional” must not become an invented green check.
Methodology and maintenance
DN selected the five-system cohort for its financial-agent research program and tested the public executable surfaces available to that program. This is not a representative market sample or a complete security audit. The compared systems have different architectures and different tested boundaries.
The core and deep workflows preserve latest records and timestamped history. Writers are serialized, PR validation does not commit production observations, and main-branch persistence uses fetch/rebase/retry protection. These controls improve the evidence pipeline; they do not certify the tested products.
Future work includes controlled upstream-pin updates, broader enforcement and recovery tests, and carefully bounded production observations. Neither a larger cohort nor more passing unit tests will automatically justify a higher exposure class.
Frequently asked questions
Does a passing DN benchmark mean an AI agent is safe?
No. A pass applies to the stated tests, upstream commit and environment. It does not establish production safety, custody security or every financial control.
Why are four systems unclassified?
Their selected executable tests provide useful bounded evidence but do not establish enough authorization, limits, recovery and containment evidence for a DN exposure class. Unclassified does not mean unsafe.
Is Almanak fully verified?
No. Almanak has a provisional E3 classification and current passing core and managed-fork surveillance. Its badge remains provisional, and production-mainnet assurance is outside the evidence boundary.
What is the difference between Dataset v0.1 and Risk Graph v1.6?
Dataset v0.1 is the comparison and recurring-evidence interface. Risk Graph v1.6 is the underlying versioned profile dataset. The two version numbers describe different artifacts.
How often is surveillance refreshed?
Core Almanak checks and the other four systems are scheduled weekly. Almanak signer recovery and position containment are scheduled monthly. Passing weekly feeds become stale after eight days and monthly feeds after 35 days.
Can developers use the dataset?
Yes. DN provides public JSON, CSV, per-system verification responses and SVG badges. Consumers should preserve source attribution and inspect each observation’s timestamp, upstream commit and limitations.
Is this a managed-runtime performance ranking?
No. This five-system financial-agent cohort does not compare cloud runtime portability, model intelligence, throughput or general hosting performance.
Sources and reproducibility
The table is a dated launch snapshot. The linked dataset and profiles may change as new observations are recorded.
- DN Agent Runtime Evidence Dataset v0.1 · CSV comparison
- DN Agent Risk Graph v1.6
- DN verification methodology
- Almanak profile · surveillance schedule and limitations
- Coinbase AgentKit profile · surveillance schedule and limitations
- Definitive Flash profile · surveillance schedule and limitations
- Liquid Co-Invest / Co-Invest Computer profile · surveillance schedule and limitations
- Virtuals EconomyOS / ACP profile · surveillance schedule and limitations
- almanak-co/sdk at 39cfabb780d2
- coinbase/agentkit at 2e6dbaf725b9
- DefinitiveCo/flash-mcp at 25064e2ce10d
- liquid-public/coinvest-mcp at 98d7f7744152
- Virtual-Protocol/acp-python at 398241f6f132
DN workflow links may require repository access. Public observation JSON exposes the run identifiers, commit pins and stated test boundaries.
Observation JSON files link the test evidence to exact upstream repositories and workflow runs. The explorer exposes those links for the selected system.
Disclosure
This is DN’s own research release. DN designed the comparison and classifications; readers should inspect the source observations rather than treat DN’s assessment as an independent third-party certification. Systems were not ordered by commercial relationships. This article contains no direct affiliate registration links.
Tests described here do not establish legal compliance, financial suitability, custody protection, insurance coverage or future trading performance. No exposure class is permission to move a reader’s funds. Product availability and rules may differ by jurisdiction and configuration.