Public contract databases are version histories, not ledgers. Sum a 111,420,683-record global corpus of official award records as published and you get $194.23 trillion — of which $39.90 trillion sits on superseded versions of amended contracts, counted again. The first corpus-scale measurement of procurement's most common counting error: what creates it, where it concentrates, and how to count correctly.
Ask what a record in a procurement database is, and the intuitive answer — "a contract" — is wrong often enough to break arithmetic at the trillion-dollar scale.
A public contract that runs for years accumulates events: option years exercised, scope modified, values re-priced, periods extended. Transparency regimes in most jurisdictions require each of these events to be published. The publication mechanism is almost always the same: a new record, carrying the contract's identifiers and its full current value, is added to the public dataset. The earlier records are not deleted — deleting them would destroy the public audit trail, which is the entire point of publishing.
The result is that a procurement database is a version history. Each long-running contract appears as a chain of records, every link of which states a full contract value as of its date. For documentation, this is exactly right. For arithmetic, it is a trap: the chain documents one procurement, and only its latest link states what the contract is now worth.
A worked example from the corpus, chosen for being ordinary rather than extreme — a single US federal contract:
| Record | Published | Stated value |
|---|---|---|
| Version 1 — original award | 2024-09-24 | $918,276.48 |
| Version 2 — amendment | 2025-08-15 | $1,710,049.32 |
| Version 3 — amendment | 2026-08-07 | $1,809,912.09 |
| Naive sum of records | $4,438,237.89 | |
| Actual contract value | $1,809,912.09 |
One procurement, three records, and a naive reading overstates it by 145.2%. Nothing about this contract is unusual — no error, no correction, no anomaly. Amendments happened, and the source documented them the way sources document amendments. Multiply this ordinary case by 14,292,691 superseded versions and the mirage reaches $39.90 trillion.
We measure across the Valan procurement corpus — 111,420,683 award records from 125 official publication systems, with amendment chains resolved deterministically on source-native award identifiers (no similarity matching; a record joins a chain only when the source's own identifiers say so).
| Measurement | Records | Stated value (USD) |
|---|---|---|
| Award records, as published | 111,420,683 | $194.23 trillion |
| — of which superseded versions of another record | 14,292,691 (12.8%) | $39.90 trillion |
| Deduplicated countable awards (one record per award) | 96,281,500 | $152.26 trillion |
The remaining 846,492 records, carrying $2.07 trillion, are not superseded versions but are excluded from the countable set by the corpus's canonical classification for non-versioning reasons; they are stated for completeness and form no part of the versioning measurement.
Three properties of the $39.90 trillion matter more than its size.
It is not an error rate. Among the 10,802,777 award identifiers that appear more than once in the corpus, 97.6% are version chains of a single procurement — the records are correct, individually and collectively. The overstatement appears only when a consumer adds them.
Version-tracked awards whose current value is $10 million or more make up just 1.07% of amended awards — but their superseded versions carry $28.69 trillion, 71.9% of all phantom value. The mirage sits precisely on top of the rows that dominate any "biggest awards," "top suppliers," or "total spending" analysis.
It is concentrated exactly where analysis is most tempting. Large contracts do not have dramatically longer amendment chains; they simply restate a large value with every amendment. The consequence is that the mirage does not spread thinly across a database: it stacks on the biggest rows first.
It compounds quietly. The 10,599,664 amended awards average only 2.35 versions each — most contracts are counted twice, not ten times. A doubling hides easily inside a large aggregate, which is why totals inflated this way rarely look implausible on their face.
Versioning is not the only way published value fields diverge from money spent. In US federal data, the headline value field on an award record is base_and_all_options: the contract's ceiling if every option is exercised. It is supposed to exceed what the government has actually obligated; a contract can carry a large ceiling and a fraction of it in obligations, and both figures are correct answers to different questions.
Across the 34,209,443 deduplicated US federal records in the corpus that state both fields, ceilings total $6.80 trillion against $4.66 trillion of recorded obligations. The headline field runs 46% above the money actually obligated — a $2.14 trillion gap. A consumer who quotes the headline field as "spending" inherits that gap on top of the versioning mirage; the two layers stack.
It is tempting to read this as a criticism of procurement transparency. It is the opposite. A regime that published only current net positions would be less useful: the version chain is what lets an auditor, a journalist, or a supplier reconstruct what was known and decided at every point in a contract's life. Point-in-time history is a feature bought at the price of making the dataset unsafe for naive aggregation — and that price is worth paying, because aggregation can be fixed at the consumer side and a destroyed audit trail cannot be fixed at all.
The obligation this creates sits with anyone who quotes a total.
Four checks, applicable to any procurement dataset:
1. Establish what a record is. Does the source publish amendments as new records? (Almost all do.) If so, the dataset is a version history and must be deduplicated to latest-version-per-award before any sum.
2. Resolve chains on the source's own identifiers. Award and modification identifiers are published for exactly this purpose. Never resolve chains by name or value similarity; deterministic identifiers or nothing.
3. Establish what the value field means. Ceiling, obligation, estimate, and outlay are different numbers that answer different questions. US federal data publishes several of them side by side; most systems publish one, and which one varies by system.
4. Treat any cross-source total with suspicion until (1)–(3) are answered per source. A pooled total inherits the worst-understood semantics of any source inside it.
The $39.90 trillion figure is a property of this corpus — official records from 125 publication systems, weighted by what those systems publish — not an estimate of the overstatement in any single official headline statistic. Official statistical publications typically apply their own deduplication; ad-hoc analyses of the raw data typically do not, and it is the latter this measurement speaks to.
Version-chain resolution is deterministic on source identifiers, so it under-counts chains wherever a source re-keys a contract on amendment. The true number of superseded versions — and therefore the true mirage — is a floor, not a ceiling.
Value fields are converted to USD at point-in-time rates; currency conversion does not create the versioning effect, but the two have not been separately decomposed.
The ceiling-versus-obligation measurement is US federal only; other systems publish different value semantics that we have not audited to the same depth.
Dollar aggregates are displayed to two decimal places of a trillion; record counts are exact.
The world's procurement transparency systems have quietly produced one of the largest public financial datasets in existence — and it is routinely misread, because it looks like a ledger and is actually an archive. The distinction is worth $39.90 trillion in this corpus alone. Any total quoted from public contract data carries an implicit claim: versions were resolved, and value semantics were checked. Where that claim is not stated, the safest assumption about the total is that it is a statement about the dataset's publishing mechanics, not about public money.
New research from Valan's 111M-record procurement dataset, delivered as it publishes. No noise.