Quantitative funds buy public-procurement data for one dominant purpose: as a revenue proxy. That reading consumes the least reliable layer of the record — value fields that stack version duplication, contract ceilings, and shared vehicle ceilings — while discarding the most reliable layer: who supplies which government, under which prime contractor, in which jurisdiction. What quantitative buyers of procurement data measure, and what they leave on the table.
Somewhere in the US federal procurement record sits a contract stating a ceiling of $1,233,340,762,115 — $1.23 trillion, on a single vendor — with $0.00 obligated against it. Query the live USAspending API today and it returns exactly those figures. It is not a data-entry error, and nothing about it violates the publishing rules of the source. It is what the headline value field on a procurement record is allowed to mean.
That record is a useful lens on an entire industry practice. Public-procurement data is sold to and bought by quantitative investors overwhelmingly for one purpose: as a revenue proxy. Match award records to listed suppliers, sum the stated values, trade the change ahead of quarterly reporting. Every link in that chain runs through the value field — and the value field is where procurement data is weakest.
Three terms matter, and they are routinely conflated:
Obligation — money the government has actually committed against the contract. The closest thing in the record to cash. Ceiling — the maximum the contract could reach if every option is exercised (in US federal data, base_and_all_options). It is supposed to exceed the obligation; both are correct answers to different questions. Vehicle (IDV) — an indefinite delivery vehicle: not a purchase but a licence to sell, under which actual orders are placed later. On multiple-award vehicles, many companies hold the same licence — and the shared program ceiling appears on each holder's record.
These layers stack into the headline number a signal pipeline consumes.
Across 16,568,857 deduplicated definitive US federal award records stating both fields, ceilings total $8.74 trillion against $3.81 trillion obligated — the headline field runs at 2.3× the committed money. A revenue proxy built on stated values is, in aggregate, measuring budget authority, not spending.
103,738 US federal vehicle (IDV) records state $976.76 trillion of combined ceilings against $0.05 trillion obligated — because a shared program ceiling is restated on every awardee of a multiple-award vehicle. Contract N0017806D4872, a US Navy services vehicle, states $1,233,340,762,115 with $0.00 obligated; other holders of vehicles in the same family each carry an identical $1,160,791,305,520. All figures verified against the live USAspending API on 23 August 2026. A pipeline that ingests vehicle records as awards does not add noise to its signal; it destroys it.
The versioning layer sits beneath both: official sources publish contract amendments as new full-value records while the old records remain published — which placed $39.90 trillion of phantom value on 14,292,691 superseded record versions in this corpus as of 16 August 2026 (The $39.9 Trillion Mirage).
And beneath the layers, a noise floor. Even a pipeline disciplined enough to filter strictly to obligation actions does not get a clean series. Fiscal-year-end de-obligations, administrative accounting true-ups, and delayed subaward reporting create artificial volatility spikes that reflect government budget cycles rather than corporate operating performance. Individual records also continue to revise in both directions after first publication — the definitive award records we sampled against the live source while preparing this article had moved since capture, up as well as down. The "cleanest" field in the record still carries the government's accounting rhythm, not the supplier's.
None of this is hidden. All of it is in the documented semantics of the source systems. And all of it lands directly on the field that value-signal pipelines treat as their input.
A global revenue proxy also assumes global coverage. The record disagrees. In a 93,805,814-record corpus of current award records spanning 102 official sources, awards are attributable to 222 buyer jurisdictions — but the distribution is a cliff.
| Jurisdiction | Award records available as structured data |
|---|---|
| United States | 62,280,482 |
| India | 29,561 |
| Japan | 9,484 |
| South Korea | 5,094 |
| China | 2,448 |
| Russia | 282 |
26 buyer jurisdictions have at least 100,000 award records available as structured, machine-readable data. 117 of 223 jurisdiction codes have fewer than 1,000 — in total, ever. Any "global" procurement signal is in practice a signal about a handful of high-disclosure jurisdictions, with the rest of the world contributing statistical decoration.
To be precise about what this measures: it is the availability of procurement records as structured, machine-readable data — not the volume of government purchasing, and not a claim about any government's internal record-keeping. But for a data consumer the distinction is academic: what cannot be obtained as data cannot enter a model.
Everything relational — which is the part of the record the source systems get right.
A procurement record is not primarily a number. It is an edge: buyer → supplier, dated, jurisdictioned, and classified. Aggregated, those edges answer questions no value signal touches.
Dependency. Which suppliers derive their business from which government customers — and how concentrated that dependency is.
The subcontracting graph. Within the same US federal corpus sit 2,588,594 subcontract records — 2,366,654 with a resolvable link to their prime award — spanning 228,701 distinct prime contracts and 371,817 distinct subcontractors: who actually performs the work, beneath the company that signed. A supplier invisible at the award level can be load-bearing one tier down, and a sampled traversal of these links resolves end-to-end (100 of 100 sampled edges).
Jurisdiction exposure. Which counterparty, in which country, under which legal regime — the input to sanctions screening, supply-chain risk, and counterparty due diligence, where a wrong answer is a compliance event rather than a bad trade.
These are the questions risk and procurement-intelligence desks ask, and the record answers them from its most trustworthy fields: identities, dates, and relationships — not values.
Two reasons.
First, the signal is built on the record's worst layer, and this article's figures put bounds on how bad that layer is: 2.3× at the ceiling level, effectively unbounded at the vehicle level. A pipeline can engineer around each distortion — deduplicate versions, prefer obligations, drop vehicles, smooth the budget-cycle spikes — but every correction is a choice made against the grain of the published format, and each choice is another way for two consumers of the same data to compute different answers.
Second, alpha is adversarial and structure is not. A revenue proxy decays as more funds consume the same feed. A dependency graph does not decay when a second desk reads it: the fact that a supplier's business concentrates in one government customer, or that a prime rests on a specific subcontractor, is a fact about the world. It stays true, it compounds as coverage deepens, and it is verifiable by anyone against the public record.
The trade in procurement data has been to sell its least reliable layer to its most sophisticated buyers. The durable product is the layer underneath.
The ceiling and vehicle analyses are US federal only; other systems publish different value semantics that we have not audited to the same depth.
Corpus figures are point-in-time as of 23 August 2026 and measure the availability of structured, machine-readable records — not total government purchasing activity. Individual award records evolve at source as contracts are amended; figures cited from our prior article are as of its publication date, 16 August 2026.
The subcontract layer is US federal and reflects records that disclose a prime linkage; subcontracting not reported in source systems is not counted.
Jurisdiction counts use ISO country codes of the buyer; a small number of records carry no attributable buyer jurisdiction. Record counts are exact; dollar aggregates are displayed to two decimal places of a trillion.
Procurement data entered the institutional market as an alpha product, priced and consumed on the layer of the record least able to carry the weight — stated values that mean ceiling, licence, or superseded history at least as often as they mean money. The layer that holds — identities, dates, relationships, jurisdictions — answers a different desk's questions, and answers them from the fields the source systems publish most reliably. The durable information in public procurement is structural, not numerical. Anyone consuming the data the other way around is trading its noise and discarding its signal.
New research from Valan's 114M-record procurement dataset, delivered as it publishes. No noise.