Revisions Ledger — Methodology

First prints are drafts

Economic data is published before it is finished. The first print of a number is an estimate built from partial source data; the agencies then revise it — for months, sometimes for decades. The worst quarter of the financial crisis is the clearest case in our own archive: 2008Q4 real GDP first printed at an annualized −3.80%. A year later the record said −5.37%. Three years later it said −8.89% — the advance estimate had missed more than five points of the collapse. The vintage standing today says −8.47%. Anyone who made a judgment in February 2009 made it against a draft.

Every number on the ledger is first-print vs the value three years later — the +3y yardstick — from the archived record. Nothing is restated from memory: every value is read from the vintage that actually stood on a given day.

The news-number basis, series by series

The ledger compares like with like: the "news number" is the figure the release actually led with, computed within a single vintage snapshot — the value standing for the month and for the prior month, from the same day's archive. Quarterly SAAR figures are computed from the same-vintage level pair, ((cur/prev)⁴ − 1) × 100.

seriesFRED idnews basisledger beginsrevision anatomy
Nonfarm payrollsPAYEMSm/m change in jobs, thousands1955-06monthly rounds, then annual benchmarks against unemployment-insurance records
Personal incomePInominal m/m %1966-02monthly rounds, then annual and comprehensive NIPA revisions
Retail salesRSAFSnominal m/m %2001-07monthly rounds, then annual benchmarks against the retail census
Industrial productionINDPROindex m/m %1928-01monthly rounds, then annual revisions
Core PCE pricesPCEPILFEprice index m/m %2000-09monthly rounds, then annual NIPA revisions
Real GDP growthGDPC1real SAAR from the same-vintage level pair1992Q2advance → second → third estimates, then annual and comprehensive revisions
Corporate profitsCPATAX% level change q/q1981Q1no advance print — profits enter later in the estimate cycle
GDIGDInominal SAAR from the same-vintage level pair2013Q2arrives later still; its early estimates settle hard

The +3y yardstick

"Settled" has to mean something fixed, or the ledger would be rewritten every time an agency runs a benchmark. We freeze each observation at first vintage + 3 years. The 2008Q4 case shows why: at +1 year the record still said −5.37 — barely half the eventual collapse on the books. At +3 years it said −8.89, within half a point of the −8.47 standing today, seventeen years and many benchmark rounds later. Three years captures the bulk of settling; waiting longer buys little and costs the ledger its stability. A settled value is written once and never touched again; observations younger than three years carry no settled row — "not yet settled" is an honest state, not a gap.

Honest floors — why our numbers differ from naive recomputations

The ALFRED archives begin mid-history: payrolls vintages start in 1955, personal income in 1966, retail sales in 2001. For observations BEFORE a series' archive floor, the oldest value on file is not a first print at all — it is a modern restatement, and its "revision" is near zero by construction. A naive recomputation across the whole archive counts those fake firsts and drags the typical revision toward zero:

payrolls +3y medianbasis
+3knaive — every archived observation, fake firsts included
+19kfloored — first prints only, the ledger's display basis

The ledger structurally excludes pre-floor observations from every table. If you recompute our numbers without the floors, you will get the first row — and it will be wrong about what first prints do.

The bases — nominal, real, and the site's own constructions

Each dossier carries a basis chip because the eight series do not share one. Personal income, retail sales, corporate profits, and GDI are nominal — the ledger revises what the agencies published, so it stays on the published basis. Core PCE is a price index; payrolls is a jobs count; industrial production is an output index. The one exception is GDP: its headline concept is real growth, so the ledger's GDP is GDPC1, the real series — the real-SAAR exception, stated on the chip. The site's GDP page reads the published growth print (A191RL) while the ledger derives the same quantity from same-vintage GDPC1 level pairs — A191RL's vintage archive is too shallow to measure revisions from (the H-2 sourcing choice); id-different, quantity-identical, which is why the GDP page carries the ledger's chip (the IND-2 basis-match ruling). None of these are the site's derived real_* constructions (like the real 10y yield), which deflate one published series by another; the ledger never mixes a derived construction into a revision comparison.

Months like this one — the conditioned counts

Each dossier's amber box splits the record by the regime that the DATA month later proved to be in — the site's settled monthly timeline, joined on the data month (quarterlies join on their mid-quarter month). The live row is chosen by the CURRENT regime read; the counts themselves never move with it. Families follow the site's taxonomy exactly: stress = Contraction; stress-adjacent = Soft Patch, Transitional, Snapback, Early Recovery; benign = all others. No new grouping was invented for this page.

Any row with fewer than 12 observations renders as "insufficient history" — a count that thin is an anecdote, and the page says so rather than showing it. The page shows family rows; the full per-label record for payrolls (the C1-B measurement of record, settled through 2023-06):

labelnmedshare up
Below-Target Drift19+17k53% up
Contraction44-60k25% up
Cooling35+37k74% up
Disinflationn=8insufficient history
Early Recovery40-10k42% up
Healthy Expansion213+25k63% up
Overheating22+6k50% up
Snapbackn=8insufficient history
Soft Patch84+16k58% up
Transitional18-32k17% up
Unclassified12+19k58% up

Span vs step — how benchmarks are counted

Benchmark revisions arrive as level wedges, and the size you quote depends on how you measure. The March 2024 payroll benchmark is the worked example: measured as a single vintage step, the level fell −589k on the day the benchmark was incorporated — February 7, 2025, the day U.S. Bureau of Labor Statistics published The Employment Situation — January 2025, the release that carried the annual benchmark revision; measured as the span from the first print's vintage (April 5, 2024) to that same day, −616k — the interim monthly rounds absorb the difference. Both figures are read off the archived PAYEMS vintages in this ledger's own store. The ledger counts benchmark-class wedges as SPAN, dated at incorporation, and carries the single-step figure as a note where the two differ materially. (The famous "−818k" from August 2024 was the PRELIMINARY March-over-March announcement — an announcement, not a vintage event; it never appears in the vintage record and so never appears here.)

The great-revisions register's admission bar

A revision enters the register when |first → settled| clears the series' own 95th percentile. The bar is per-series because the units are not comparable: a 0.14pp core PCE revision and a 400k payrolls revision are both once-in-twenty events in their own distributions, and the register ranks by that own-percentile — never by raw size across series. The native-units column rides beside the rank so you can still see the magnitudes.

Recognition — when the data admitted it

For each modern recession we record the first vintage in which the data conceded it: payrolls' first negative print near the start, and real GDP's first vintage showing the opening quarter negative. The asymmetry is the finding. Payrolls said "recession" nearly in real time every cycle — April 2001's −86k printed on 2001-04-06; April 2008's −20k printed on 2008-05-02. Real GDP conceded 2001 only at the July 2002 annual revision (16 months late) and 2008 only in July 2009 (18 months). And 2001 carries its own asymmetry: NOMINAL GDP never admitted the 2001 recession at any vintage — only the real series ever went negative. 2020 is the exception that proves the scale: a collapse large enough that the advance estimate caught it in real time.

The contrast row

Core PCE is on the ledger to show what a quiet series looks like: 160 of its 274 settled observations moved less than 0.05pp — at display precision, they did not move at all. Revisions are a property of how a series is built, not of economic statistics in general; the ledger shows the loud and the quiet on the same yardstick.

Limitations

The floors are real constraints: nothing before a series' archive floor can appear anywhere on the page, and the floors differ by decades across the eight series. Observations younger than +3 years are not yet settled and carry no verdict. Conditioned counts are counts — small denominators are flagged, and none of them are probabilities. The ledger is a descriptive record of how first prints settled. It is never a forecast of how the next print will settle.

Provenance

Vintages come from ALFRED (the St. Louis Fed's archival FRED) plus the site's own vintage store, which snapshots every indicator nightly. The derived tables are incremental and write-once: new first prints append, and each observation's settled value is written exactly once when its +3y birthday passes. The pipeline's anchor rows — 2008-11 payrolls −533k → −802k, 2020-04 personal income +10.51 → +11.74, 2008Q4 GDP −3.80 → −8.89, and the 2008-05-02 payrolls recognition date — are reproduced digit-exact from the archives in the test suite on every run.