Last week a research pipeline I was asked to review produced a nine-dimension assessment of a tokenized infrastructure project. The document was immaculate. Technical evaluation: N/A. Supply structure: N/A. Governance health: N/A. Risk matrix: six categories, zero enumerated risks. It closed with a confidence disclaimer and a recommendation for further diligence. It took four minutes to notice: there was no token. Retrieval had returned nothing — no whitepaper, no contract address, no repository — and the model, rather than halting, completed the schema anyway. The output was not a lie. It was a form. And the form had already been pasted into a deal memo whose only function was to make a third party believe diligence had occurred.

Structured due diligence is now the default in crypto. Since spot Bitcoin ETFs reopened the institutional door, every fund has bolted some version of an LLM research agent onto deal flow, because headcount cannot scale with the token count. The scaffold is consistent across implementations: nine dimensions — technical, tokenomics, market, ecosystem position, compliance, team and governance, risk, narrative, and supply-chain transmission. Each dimension has sub-fields, each sub-field has a required value, and the required value must be machine-readable. On paper, this is exactly the checklist a careful human analyst would run. The architecture, however, has three stages that fail differently. Retrieval gathers sources. Extraction maps text to fields. Scoring ranks what was extracted. Only the third stage is visible in the deliverable. The first two are invisible — and when retrieval returns an empty set, extraction does not error out. It writes N/A into every cell, and scoring produces a verdict anyway, because a schema with all its slots filled is a valid schema.
I have seen this pattern before, in a different costume. In 2017, as a high school junior, I pulled apart the $1.4 billion ParagonCoin raise and found no whitepaper worth reading and no contract worth auditing — a template with every section present and every section empty. The ICO era's deliverable was the whitepaper; the template was the product. 2017's dream is today's regulation, and its research artifact has merely been upgraded: the whitepaper became a dashboard.
The mechanism deserves precision, because it is not a bug in the model. It is a property of every system whose success metric is format validity. When a pipeline is judged on whether it emitted a conforming document, it will emit a conforming document. Three failure classes show up in production, and they are not equally dangerous. The silent null — empty retrieval, blank fields, formatted output — is the loudest and the least harmful, because a human eventually notices the holes. The fabricated fill is quieter: the model infers a plausible team allocation, a plausible TVL, a plausible audit date. Nothing in the interface distinguishes an inferred value from a retrieved one. The third class, correct halting, is rare, because halting produces no artifact.
The dangerous artifact is not the one that says nothing. It is the one that says something plausible in a format that discourages verification.
I spent a year inside a payments research lab co-developing a prototype for a privacy-preserving digital dollar, and the lesson that transferred most cleanly was not about zero-knowledge proofs. It was about staleness. A consumer contract reading a price feed checks one number. It does not check whether that number was updated in the last thirty seconds. An oracle reporting the same price for six hours and an oracle reporting nothing at all are indistinguishable at the interface — unless the consumer explicitly reads the timestamp. That is the whole architecture of the problem. The interface is designed for the answer, and nobody designed it for the absence of an answer.
The same asymmetry runs through the scaling debate. There are dozens of Layer 2s and roughly one user base, so liquidity is sliced thinner with every launch — and every launch still produces a routed, formatted, apparently functional network. Fragmentation reads as growth at the dashboard level. The nine-dimension template has the identical pathology: thin retrieval sliced into nine outputs produces the texture of breadth. A single honest sentence — insufficient verifiable information to assess — has been distributed across nine headings until it reads like a report.
The bull market makes this expensive rather than merely embarrassing. When risk appetite is high, a research artifact's function is justification, not discovery. A fabricated diligence memorandum has asymmetric returns: it is consumed to support a position, and the position's eventual outcome is attributed to the market, not to the memo. Nobody audits the auditor. The assets that most need real diligence are the ones with the least retrievable data. The more uncertain the asset, the more confident the report. That is not a research failure; it is an incentive structure.
There is a version of this that works. Bitcoin's fee market was structurally thin until Ordinals pushed inscription demand into block space. Without that demand, the security-budget conversation would already be a funding conversation. The point is not that inscriptions are good. The point is that fee revenue is a number a model cannot generate — it is paid, or it is not. Liquidity behaves the same way. It is the only input in the stack that cannot be synthesized.
The consensus reaction to a null result is that the pipeline hallucinated and needs guardrails. The contrarian read is nearly the opposite: the empty assessment was the most honest document in the folder, and the real failure is institutional intolerance for N/A. For the overwhelming majority of tokens, there is genuinely not enough verifiable information to score nine dimensions. That is a finding, not a defect. But a finding of no data produces nothing allocatable; it terminates the process. So the system runs anti-inductively: confidence is rewarded in proportion to obscurity, and the one class of output that is always true gets discarded. In my own audit work, a report that ended in a halt has never once been invalidated by a later event. The reports that filled every cell have been wrong in both directions.
The next two years of agent payments will be judged on one capability: whether machines can settle with each other autonomously, without a human in the loop. None of that works if the agent cannot distinguish an empty feed from a stale one. The missing infrastructure is not a better model. It is provenance at the retrieval layer — signed nulls, retrieval receipts, machine-checkable proof that something was not found. When the oracle is silent, does your contract know?