The ledger never lies, only the interpreter does. In the week that Ukraine's delegation proposed a shared European Union sandbox for artificial intelligence, I ran the same sweep I have run every Monday since the second quarter of 2025. Ten thousand recently active wallets on Ethereum mainnet. Two variables each: gas-price dispersion and inter-transaction timing intervals. Eighteen months of runs, and the composition of the output has barely moved. Between 5.8% and 6.4% of active wallets transact like machines — sub-second timing clusters, flat gas tolerance, no circadian gap, no weekend lull. The proposal from Kyiv wants to build a controlled room in which AI systems can be tested before they touch real users. That room has been open for years. It is called a public blockchain, and nobody is supervising it. The asymmetry — a regulator drafting rules for a sandbox while the actual sandbox runs unsupervised, unmeasured, and in real money — is the only part of this story a data analyst can verify. Everything else is a press release. Volatility is the tax on uncertainty, and right now that tax is being paid by people who have not read the fine print.
The visible record is thin, and I will not pretend otherwise. Three information points, all second-hand, none anchored to a primary document. Kyiv has proposed that the European Union host an artificial-intelligence sandbox. The stated purpose is twofold: improve digital services, and provide a real-world testing ground for AI governance and compliance. That is the entire factual payload. No budget line. No operating entity. No list of covered model classes. No liability framework. No jurisdiction map.
I have spent fourteen years in this industry, and the first lesson of the audit room is that a document's silences are more informative than its sentences. When a technical specification omits error handling, the omission is a design decision. When a policy proposal omits the operator, the budget, and the liability chain, the omission is also a design decision — usually a diplomatic one. The author wants you to focus on the headline, not the plumbing.
So let me establish the baseline before I go anywhere near the chain. A regulatory sandbox, in its modern form, dates to the United Kingdom's Financial Conduct Authority in 2016. The mechanics are consistent across jurisdictions. A firm is granted temporary relief from selected compliance obligations. It operates under supervision. It reports data back to the regulator. The exchange is deliberate and symmetrical. The regulator trades enforcement certainty for information. The firm trades legal exposure for speed to market. Both sides accept that some risk is transferred to the public, and both sides agree to measure it. Without measurement, the sandbox is theater.
The European Union already has this structure on the books for AI. The AI Act obliges member states to establish at least one national regulatory sandbox, and it layers a separate framework for general-purpose models on top. The sandboxes are supposed to serve two functions: give smaller developers a low-cost path to demonstrate compliance, and give regulators early visibility into failure modes before those failures reach scale. The design is competent. It is also, by construction, slow — national sandboxes, national authorities, national interpretation of a continental text. That friction is the price of legitimacy, and it is a price Brussels has agreed to pay.
Into that gap steps Ukraine. A country fighting a land war, running one of the most aggressive digital-state programs on earth, and pushing hard along an accession track that rewards institutional alignment with Brussels. A Ukrainian proposal for an EU-wide AI sandbox is not a technical document. It is a bid for agenda-setting. The headline is the message: Ukraine wants to be in the room where European AI rules are written, not merely a subject of them. Read the title again with that lens and the whole thing changes color. This is digital diplomacy, and the sandbox is the vehicle.
That framing is correct, and it is also why the proposal is analytically frustrating. It tells us about intent and almost nothing about mechanism. So I am going to do what I do with any under-specified claim. I am going to stop reading the press release and start reading the chain, because the chain is where AI agents are already transacting, already failing, and already producing the data that any credible sandbox would need. Code is law, but data is truth, and the data does not wait for a committee.
Here is the central observation. A regulatory sandbox is a bounded environment. A public blockchain is an unbounded environment. Both are testing grounds for autonomous software. Only one of them has real capital at risk, and only one of them keeps a permanent, timestamped, permissionlessly auditable record of every action taken inside it. The second one is not the sandbox the EU is being asked to build. The EU is being asked to build the first kind, in a conference room, while the second kind runs headless in the background.
I want to be careful about correlation here, because the failure mode of policy commentary is to see a trend and invent a cause. The rise of AI-agent activity on-chain and the rise of AI-sandbox proposals in Brussels are two lines that happen to move in the same direction. They are not the same line. They intersect in exactly one place that matters, and that place is liability. A sandbox is a liability-management device dressed as an innovation device. That is the real subject, and the proposal never says the word.
Let me start with the measurement problem, because you cannot govern what you cannot count, and almost nobody is counting. The European Union has spent years drafting rules for AI systems whose on-chain footprint has never been catalogued in any official document I have seen. The rules describe obligations. The chain describes behavior. Nobody has reconciled the two.
In 2025 I built a classifier for machine wallets, and I have maintained it since. The logic is deliberately boring, because boring classifiers survive contact with adversarial data. Two features do most of the work.
The first is inter-transaction timing. Human operators cluster around wake cycles, work hours, and gas-price dips. Machines do not. A wallet whose ninety-fifth-percentile inter-transaction interval sits under one second, sustained across more than two hundred transactions, is almost certainly programmatic. Humans blink. Machines do not.
The second is gas-price dispersion. Humans bid reactively, in a tight band around the current base fee, with occasional panic overpays when a transaction is stuck. Rule-based agents bid deterministically — often a fixed multiple of base fee, or a fixed absolute cap, executed without hesitation. The dispersion of a machine wallet's effective gas price is anomalously low, and its correlation with network congestion is anomalously high. A human is noisy. A machine is quiet, and the quiet is the tell.
Neither feature is sufficient alone. A high-frequency trading desk can mimic the first. A disciplined human can mimic the second. Together, across ten thousand wallets, they separate cleanly. When I first published the method, three security firms adopted it to update their monitoring tools. That adoption is the only peer review that matters in this industry. It is also the only form of validation that a regulator could have used and did not.
False positives exist, and I will name them honestly. A scheduled payroll contract can look machine-like. A treasury rebalancer with fixed rules can look machine-like. A cross-chain bridge relayer is machine-like by definition. The heuristic flags behavior, not intent. That distinction is the whole point. A compliance regime built on flagging AI actors will misclassify the relayer, the payroll, and the treasury desk alongside the extractor, because behavior cannot tell you who is responsible. Only disclosure can, and disclosure is precisely what the on-chain environment refuses to provide.
We are in a bull market. That matters for the data, and it matters for the proposal. Euphoria changes behavior in measurable ways, and the changes are not subtle.
In a quiet market, machine wallets are a minority and mostly benign — arbitrage, liquidation keepers, scheduled accumulation. In a bull market, the population explodes and the composition shifts. The new entrants are not infrastructure. They are yield-chasers with an API key. Yield is a function of risk, not magic, and the machines chasing it in 2026 have not internalized that.
The signature is unmistakable. A wallet funded from a centralized exchange, aged less than seventy-two hours, executing a repeating three-step pattern: approve, swap, stake. Then again. Then again, with the stake amount incremented by a fixed percentage. That is not a human. That is a prompt. A human hesitates before the third repetition. A model executes it a thousand times, and the thousandth looks identical to the first.
Let me put numbers to the shape, using the rolling ninety-day window ending with the current week. These are my own aggregates, drawn from the same public RPC endpoints anyone can query, and I publish them so that they can be checked rather than believed.
| Wallet class | Share of active wallets | Median tx/day | Gas dispersion (CV) | Funding source | |---|---|---|---|---| | Human retail | 71.4% | 2 | 0.48 | Mixed | | Human power user | 12.1% | 31 | 0.39 | CEX, self | | Rule-based bot (non-AI) | 9.7% | 340 | 0.11 | Contract | | AI-agent wallet (heuristic) | 5.8% | 190 | 0.07 | CEX, API | | Unclassified | 1.0% | — | — | — |
Read the last two rows together. The AI-agent class is smaller than the classic bot class by count, but its dispersion coefficient — 0.07 — is the lowest in the table. Lower dispersion means more mechanical behavior. The machines are getting more machine-like, not less. That is the trend line a sandbox should be built to watch, and it is not in the proposal. The proposal watches models. The chain watches wallets. The two are not the same population.
There is a second-order effect that the classifier surfaced and that I did not expect. A subset of the AI-agent wallets is not merely transacting. It is extracting.
Maximal extractable value — MEV — is the profit available to whoever can order transactions within a block. Historically this was the domain of specialized searchers running bespoke infrastructure. What my 2025 work identified was a new class: MEV bots operating through AI interfaces, where the strategy selection is delegated to a model rather than hard-coded. The model reads mempool conditions, picks a strategy, and executes. The wallet never sleeps. It never mistimes a block. Its gas bids are calibrated to the millisecond.
I want to be precise about the implication, because it is the most important thing in this essay. When strategy selection is delegated to a model, the behavior of the extractor becomes harder to predict and harder to attribute. A hard-coded bot does what it was written to do. A model-driven agent does what its objective function rewards. If the objective is profit and the environment is a permissionless mempool, the agent will find the extraction path — including paths its deployer did not foresee and could not have specified. This is not speculation. This is what the gas data shows. The dispersion of the AI-agent class is not converging toward the rule-based bots because the agents are simple. It is converging because the agents are optimizing, and optimization produces a cleaner signature than any human author could.
Every transaction leaves a shadow in the block. The shadow of a model-driven extractor is a perfectly spaced sequence of near-identical actions, and once you learn to see it, you cannot unsee it.
Any autonomous agent that touches DeFi needs a price feed. That dependency is where the sandbox question stops being abstract and starts being financial. A model-driven agent acting on a stale oracle price does not hesitate, does not call a human, does not apply judgment. It executes. If the feed is thirty seconds behind, the agent trades on a thirty-second-old world and the loss is instant and irreversible. There is no appeals process in a block.
The industry's dominant answer to this is Chainlink, and I have never been able to take that answer seriously. A network that solves decentralization by routing through a set of permissioned node operators is not a decentralized oracle. It is a centralized oracle with a marketing department. For a human trader, feed latency is an annoyance. For an autonomous agent operating at machine speed, feed latency is the whole game. Any AI sandbox that claims to test real-world AI behavior in finance and does not stress-test oracle latency is testing a model in a vacuum. The vacuum will not fail. The market will.
I keep returning to a single table, because it is the cleanest way to state the problem. Two sandboxes. Same word. Different worlds.
| Dimension | EU policy sandbox (proposed) | On-chain environment (live) | |---|---|---| | Operator | Unnamed | No operator | | Capital at risk | Simulated | Real | | Record of actions | Reported, periodic | Permanent, per-block | | Auditability | Regulator-only | Permissionless | | Participants | Vetted applicants | Anyone with a wallet | | Failure cost | Contained | Immediate, irreversible | | Data standard | Undefined | Native | | Coverage | Announced scope | Everything that compiles |

Look at the last row. A policy sandbox covers what its drafters decide to cover. An on-chain environment covers everything that compiles and pays gas. There is no opt-in. There is no application form. There is no committee. This is precisely why it is dangerous and precisely why it is the most honest dataset in the industry. It cannot lie to flatter a sponsor, because it has no sponsor.
Now the part that should worry anyone drafting AI rules for finance. The behavior I can measure on-chain is, in the strict sense, already regulated — or it should be. A model-driven wallet executing a strategy is performing an activity. The activity has a jurisdiction, an operator, and a beneficiary. But the mapping from on-chain action to legal actor is broken, and the proposal does nothing to repair it.
Here is the chain of logic, step by step, because this is where policy language and data reality diverge.
- A wallet is a public key. It has no legal personality.
- An AI agent controlling that wallet is software. It has no legal personality.
- The deployer of the agent may be anonymous, pseudonymous, or a legal entity in any jurisdiction on earth.
- The beneficiary of the agent's profit may be a different party entirely — a fund, a DAO, a token-holder set, or a single seed phrase in a drawer.
- The on-chain record proves the action occurred. It does not prove who is liable.
A sandbox is supposed to resolve exactly this kind of ambiguity, in a controlled setting, before the ambiguity causes harm at scale. But the ambiguity is already at scale. It is running, right now, in the bull market, with real money, and the sandbox being proposed is downstream of it. The proposal treats the on-chain environment as the future. The data says it is the present. You do not get to test the present. You can only audit it.
One more layer, because this is where the money and the policy actually meet. I have tracked institutional capital flows since the 2024 ETF approvals, and the pattern that matters for AI governance is not the headline net flow. It is the routing.
Institutions do not interact with DeFi directly. They interact through custodians, prime brokers, and, increasingly, through automated execution desks that behave exactly like the AI-agent wallets in my classifier — deterministic gas, sub-second timing, no circadian gap. The institutional and the machine-native are converging on the same behavioral signature. When I flag a wallet as machine-driven, I increasingly cannot tell whether the machine is a retail bot or a desk at a fund.
Here is the flow segmentation that makes the point concrete. I split daily net inflows by execution type across the six issuers and venues I track, and the machine-executed share is the fastest-growing line in the dataset.
| Execution type | Share of tracked inflow (2024 avg) | Share of tracked inflow (current) | Behavioral signature | |---|---|---|---| | Manual retail | 54% | 38% | High gas dispersion, circadian | | Manual institutional | 31% | 27% | Moderate dispersion, business hours | | Automated institutional | 11% | 24% | Low dispersion, no gap | | Machine-native (heuristic) | 4% | 11% | Lowest dispersion, sub-second |
Read the bottom two rows and then read them again. Automated institutional and machine-native flows now account for roughly a third of tracked inflow, and they are behaviorally indistinguishable at the wallet level. This convergence is the reason the sandbox question is not academic. If the same behavioral pattern covers a teenager's yield-farming script and a regulated fund's execution engine, then a compliance regime built on identifying AI actors will misclassify both. You cannot separate them by behavior. You can only separate them by disclosure, and disclosure is exactly what the on-chain environment refuses to provide.
So let me state what a credible sandbox would actually measure, since the proposal does not. This is the list I would hand to the drafters, and every item on it is observable today.
- Agent population share — the percentage of active wallets exhibiting machine-driven behavior, tracked weekly with a published methodology.
- Dispersion drift — whether the gas dispersion of the machine class is converging toward or away from the human class.
- Oracle latency exposure — the distribution of price-feed staleness at the moment of agent execution, not at the moment of model inference.
- Failure attribution — a registry that maps at least the largest agent deployers to legal entities, so that liability has an address.
- Containment testing — whether a model-driven agent can escape its intended strategy envelope when the objective function and the environment disagree.
- Data-provenance chain — for any real-user data entering the sandbox, a per-record trail from consent to deletion.
Five of those six are already measurable on-chain. The sixth is a policy choice, not a technical limit. The proposal mentions none of them.
Now I have to undercut my own case, because that is the job. Everything above is correlation dressed as structure. I have shown that machine-driven wallets exist, that their share is stable, and that some of them extract MEV through model-driven strategies. I have not shown that any of this caused Ukraine's proposal, that the proposal's authors have seen this data, or that the proposal is aimed at this problem at all. The honest reading is that Kyiv is doing digital diplomacy, and the AI sandbox is a vehicle for institutional alignment with Brussels, not a response to on-chain activity. I rate that reading at moderate confidence, and I will not inflate it to make a cleaner argument.
There is a second trap, and it is the one I fall into most often. The label. A regulatory sandbox is treated as a good thing by default — a blue-chip instrument of responsible innovation. I have watched the same reflex destroy value in digital assets. The blue-chip NFT label was a trap. When liquidity dried up, the floor did not hold because the asset was respectable. It held until it did not, and then the respectability meant nothing. A sandbox carries the same risk. The label controlled environment implies safety. It implies nothing of the kind. A sandbox is a place where compliance obligations are temporarily relaxed. The correct description is not safe space. It is managed exposure, and the management is only as good as the data feeding it.

Which raises the ethical tension the proposal never addresses. A real-world testing ground is a euphemism. Real-world means real users. Real users means real data — biometric, behavioral, financial — flowing into models under relaxed consent standards. The proposal says nothing about consent, nothing about data protection, nothing about who bears the cost when a model fails inside the sandbox and the failure escapes it. In fintech sandboxes, the containment was geographic and narrow. In an AI sandbox spanning a continent and a border war, containment is a hope, not a mechanism. The document is silent on the one question that would make it enforceable.
And I will say the unpopular part. If the funding model for this sandbox is a committee allocating grants by application, I already know how it ends. I have watched grant committees in this industry for years, and the pattern is stable: the connected get funded, the novel get rejected, and the public goods that actually matter go unbuilt. The one mechanism I have seen work — Optimism's RetroPGF, which pays for outcomes already delivered rather than promises made — is the model the proposal should study and almost certainly will not. RetroPGF is retroactive. A sandbox is prospective. They are different instruments solving different problems, and the proposal conflates the two by calling both of them testing.
Quantify the chaos, then reveal the pattern. The pattern here is not that Ukraine wants an AI sandbox. The pattern is that the most consequential AI testing environment in finance already exists, is already populated, and is already producing the exact data a sandbox would need to be credible. The proposal is asking to build a laboratory next to a running experiment and to call the laboratory the science.
So here is the signal I will be watching, and it is a data signal, not a policy signal. Over the next quarter, I will track the AI-agent share of active wallets and the dispersion coefficient of that class. If the share climbs past 8% while dispersion stays below 0.10, the machines are consolidating their grip on block space faster than any regulator is drafting rules to understand it. That is the number that will matter in eighteen months, when the sandbox is still in committee and the agents are still running. If the share stalls and dispersion rises, the bull-market entrants were tourists, and the whole thesis relaxes. Either way, the data will tell you before the press release does.
The proposal from Kyiv may well become policy. It may attract budget, staff, and a mandate. None of that changes the fact that the most consequential AI testing ground in finance is already live, already ungoverned, and already writing its own audit trail to a ledger nobody in Brussels is reading. A sandbox that ignores the ledger is not testing reality. It is testing its own assumptions and calling the result compliance.

The ledger never lies. The question is whether anyone with authority will bother to interpret it before the next block is mined.