Most countries' household microdata is restricted. Rules engines are
no longer the bottleneck to modeling their tax and benefit systems —
data access is. This page is one application built on populace's six
strategies: recalibrating populace's openly redistributable US
support to another geography's published totals, producing
microsimulation inputs there — validated where ground truth exists,
deployed where it doesn't.
Reweighting a fixed sample is adaptation with zero new parameters.
Treat the support — the record set — as a frozen representation, and
the weights as a task head fit on top of it. Recalibrating that
frozen support to a new geography's published totals is head-only
adaptation: every household stays exactly as observed; only its
influence on the aggregate changes. Importance weighting of a fixed
sample is the zero-parameter-update case of adapting a generative
model — the same operation populace already runs to hit a
congressional district's targets, run again at a larger domain
distance.
A capacity ladder sits above that zero-parameter case: frozen support
plus new weights, then frozen support plus imputed extensions for
variables the target geography needs and the US support doesn't
carry, then a fully adapted generative model that replaces the
support itself. Higher rungs spend more capacity and can close gaps
reweighting cannot. The mechanism at each rung is swappable — the
question this page and its paper ask is whether the adaptation
framing survives the climb, not which rung is correct.
02 — one operation, three distances
Reweighting a congressional district, the UK, and Belgium is the same operator.
populace's local-area method already reweights one national file to a
US congressional district's published totals — same country, same
institutions, the smallest domain shift the operator runs on. Moving
the same operator to the UK and to Belgium does not change what it
does; it changes how far the source domain sits from the target. The
diagram below places three actual runs on that axis.
Three actual runs of the same reweighting operator, placed by domain distance. The congressional-district leg (populace's local-area method), the UK leg (v1.2 measured against the held-out FRS, below), and the Belgium leg (measured below) all sit on solid ground. The UK leg's held-out verdict — two of four survey views clearly distinct from native microdata, down from three of four on the as-found file and unchanged since the repaired feed — is the paper's thesis made quantitative; the same repair nearly doubled a different view's continuous energy-distance gap to its floor, and a further declared mapping fix then closed nearly all of that gap while the verdict itself never moved — results reported alongside each other, not one in place of the last. Same operator, increasing distance — small and same-institutions at the congressional-district end, large and different-institutions at the Belgium end.
stop 1
Congressional district
Described on the sparsity strategy page and
used throughout populace's US releases: one national file reweighted
to sub-national published totals. Same country, same tax-benefit
institutions, the smallest domain shift this page discusses —
referenced here descriptively, not re-validated on this page.
stop 2
United Kingdom
The held-out ground-truth leg. Its first evaluation scored the
recalibrated US support against the Family Resources Survey —
native UK household microdata withheld from the adaptation step —
with three of four survey views clearly distinct from native, and
caught a concrete feed defect. A repaired v1.1 rerun moved the
headline to two of four views clearly distinct while nearly
doubling a different view's energy-distance gap to its floor; a
further v1.2 rerun then applied a single declared mapping fix and
closed nearly all of that same gap, on the unchanged scoring
configuration throughout. Verdict and all three runs below.
stop 3
Belgium
The zero-shot deployment leg, where no comparable native microdata
exists to score against. Validated instead by an
identical-population cross-engine check: the Axiom Belgian engine
against EUROMOD, isolating engine agreement from data quality.
Numbers below.
03 — the validation protocol
Validate where truth exists. Deploy where it doesn't.
The protocol borrows directly from held-out benchmark practice:
adapt to a target domain without touching its held-out ground truth,
then score against that ground truth once. The UK supplies the
ground-truth leg because native survey microdata — the Family
Resources Survey — exists to hold out and score against. Belgium
supplies the deployment leg precisely because that check is
unavailable: no comparably granular Belgian household microdata sits
behind the recalibration, so deployment is checked the only way
available, cross-engine agreement on an identical population, rather
than against native microdata.
Calibration forces target-geography totals to match by
construction. Reweighting is fit to hit published aggregates,
so once it runs, the aggregates match almost by definition. Every
residual gap that remains against the FRS microdata — not the totals,
the joint distribution underneath them — measures something
reweighting cannot touch: the joint structure the recalibrated file
inherited from the US support it started from. That gap is the
quantitative test of the populace thesis stated on the
support strategy page: everything downstream
can reweight the support; nothing downstream can repair it. The UK
benchmark turns that sentence into a number.
adaptshipped
Head-only, ground truth untouched
The US support is recalibrated to the target geography's published
totals without ever consulting that geography's held-out microdata
— the adaptation step and the scoring step use disjoint
information, exactly as a transfer-learning benchmark requires.
score oncev1.2 done
Held-out FRS, scored after each adaptation
The UK leg scores the adapted file against native Family Resources
Survey microdata it never saw. Three runs are complete — two of
four survey views clearly distinct from native as of the latest,
down from three of four on the as-found file (section 07) — so
this is the leg that already measures what reweighting alone, and
one declared mapping fix on top of it, can and cannot repair.
deploy anywayshipped
No native microdata, cross-engine check instead
Belgium has no held-out survey microdata to score against, so
deployment is checked the way that is actually available: an
independent tax-benefit engine, EUROMOD, computing the same
quantities on the identical recalibrated population.
04 — beyond the UK and Belgium
Open microsimulation anywhere does not require open microdata anywhere.
The UK leg works because native household microdata exists and can
be held out. Most geographies fail that condition — not because no
one holds the data, but because the data cannot leave the building
it lives in. National statistical institutes and research centres
routinely hold exactly the restricted microdata this page has no
access to. The protocol above does not need that data to travel;
it needs a referee that can.
It requires an open referee. The referee travels even where
the data cannot. Instead of asking a data custodian for
microdata, ship the custodian a sealed, pre-registered evaluation
procedure: the recalibrated support, the population-view harness,
and a frozen scoring configuration. The custodian runs it locally,
against their own restricted ground truth, inside their own
environment. Nothing sensitive leaves the building — only a
low-dimensional scorecard comes back: energy distance, coverage,
classifier AUC, and the tail block, the same four numbers this
page already reports for the UK leg. A scorecard contains no micro
records, so publishing it is safe by construction, not by review.
A protocol-stage schematic (dashed = specified, not run). The evaluation pack — recalibrated support, harness at a pinned version, frozen scoring configuration, input manifests — is pre-registered before any run, so neither side can select thresholds or blocks after seeing a result. It crosses into the custodian's own environment, runs there against restricted ground truth that never leaves, and only the low-dimensional scorecard returns. No custodian run has happened yet; the UK leg above is this pack's dress rehearsal, proving it runs correctly before any third party runs it.
pre-registrationprotocol stage
Frozen before results exist
The harness version, the four scoring blocks, and their thresholds
are committed before any custodian run — the same discipline that
gives a held-out machine-learning test set its credibility.
Freezing first is what stops either side from selecting the
comparison after seeing how it turns out.
the eval packprotocol stage
A pinned, portable bundle
Harness commit, frozen configuration, input manifests, and an
entry-point script travel together as one pinned artifact — built
to run unattended inside an environment its authors never enter,
producing a fixed output schema every time.
what it buysprotocol stage
Ground-truth checks without data access
Micro-level validation becomes available in every geography with
a willing custodian, not only the geographies whose microdata is
already open. The UK leg is this pack's dress rehearsal: proof it
runs correctly end to end before any third party is asked to run
it against data this page will never see.
The dress rehearsal already earned its keep. Run
against the UK's own held-out file — the one geography where this
page can score against native microdata — the pack did not just
confirm it executes. It caught a concrete, mechanical defect in the
transferred file: a missing income stream that zeroed out state
pensions and cascaded into the means-tested benefits that depend on
them (section 07). A referee that surfaces a specific data error on
its first real run is a referee worth shipping to a custodian who
holds data this page will never see. The pre-registered pack, run as
a custodian would run it, did exactly the job it exists to do.
Reweighting can only redistribute mass across the households already
in the support; it cannot manufacture a joint structure the US
population never exhibits. If the target geography has a
household-composition pattern, an income-tax-and-benefit
interaction, or a demographic combination absent from the source
support, no amount of reweighting recovers it — that is a limit of
this rung of the capacity ladder, not a bug in the reweighting
method. Higher rungs (imputed extensions, or a fully adapted
generative model) exist precisely to spend more capacity against
this limit. This page and its paper report where that limit binds,
not just where the operator succeeds.
06 — measured now
Belgium: 21 of 21 targets, one engine cross-check.
The US support, recalibrated to Belgian published totals, is the
first deployment-leg result this page can report in full. Both
numbers below come from a completed populace-be run, not a
benchmark still executing.
Axiom vs. EUROMOD, worker SSC€643apart on €20.9Bidentical recalibrated population, independent engines
[SOURCE: populace-be run <sha/date>] — pending pin.
The identical-population design is what makes the €643 figure a
statement about engine agreement, not data quality: both the Axiom
Belgian engine and EUROMOD compute worker social security
contributions on the exact same recalibrated households, so the
€643 gap on €20.9 billion isolates disagreement between the two
engines' rules from any disagreement about who is in the population.
This is the deployment leg described above — no native Belgian
microdata sits behind either number, so cross-engine agreement is the
check available, not a substitute for the FRS-style holdout the UK
leg runs.
A proportionality caveat this figure needs every time it's cited.
Belgian employee social security contributions are levied at a rate
that's largely proportional to gross pay, so agreement this close on
an identical population is close to what rate-times-base arithmetic on
the same inputs would already deliver. €643 on €20.9 billion certifies
concept-base alignment — the two engines agree on what counts as
contributable pay — and rules out implementation drift between them,
but it's a weak discriminator: a proportional levy has little
nonlinearity (thresholds, bands, means tests, caps) for two engines to
disagree about. A sharper validation ladder is planned, not yet run:
per-record agreement rather than the aggregate alone, extension to
genuinely nonlinear programs, and agreement on reform-induced deltas
rather than baseline levels — tracked at
axiom-rules-engine#77. Until that ladder
runs, read this figure as a necessary but weak check, not as evidence
the two engines agree on the rules generally.
United Kingdom: the referee ran, diagnosed, reran, and priced a declared fix.
The UK leg has now run three times. Its subject is the prior US-to-UK
transfer: an existing recalibrated file built from open US microdata,
reweighted to UK national, regional, and country targets — the same
target set the native Family Resources Survey file calibrates to. The
pre-registered pack scored the as-found v1 file against the held-out
native FRS file before that file was touched, exactly as section 03
requires; scored a repaired-feed v1.1 rebuild under the identical
frozen scoring configuration; then scored a v1.2 rebuild applying one
declared mapping transform on top of that repair, isolated against
four hypotheses committed to the record before the run. All three are
diagnostic passes, not a finished benchmark: v1 caught a concrete feed
defect, the v1.1 repair moved the headline verdict and a continuous
proper score in opposite directions on the same view, and v1.2 closed
that same view's gap to its floor — reported together below, not as a
single improved number.
held-out verdict, four survey views2 / 4clearly distinct; 0/4 indistinguishable — down from 3/4 on v1, unchanged v1.1 to v1.2income intermediate; tax_benefit and full clearly distinct (unchanged since v1.1); net_income moved from clearly distinct (v1) to intermediate (v1.1, v1.2)
net_income energy ratio to floor6.3× → 12.5× → 1.5×doubled, then fell within reach of the floorthe same view whose band verdict never moves off intermediate at v1.1/v1.2 — a declared mapping fix on top of the feed repair closed almost the entire remaining gap
sampling-noise floor, classifier AUC0.49–0.52at chance, all four views, all three runsthe native file's own split scored as a candidate — machinery validated
determinism, two independent runs eachbyte-identicalcanonical subset, v1, v1.1, and v1.2scores, verdict, and config fingerprint reproduce exactly; only the timestamp differs — eval_config hash unchanged across all three versions
[SOURCE: transfer-paper/experiments/uk-transfer/artifacts/scorecard.json, scorecard_v1_1.json, scorecard_v1_2.json, scorecard_three_way.json, decomp_three_way.json — popdgp 402fb235, period 2025, v1 candidate n=28,532 vs v1.1/v1.2 candidate n=63,128, all vs native FRS n=53,508; verdict bands pre-registered in eval-pack/PREREGISTRATION.md].
This is the quantitative test the framing promised, run three times: calibration
forced the marginal targets to move (transferred mean absolute relative
error 0.309 against the full 568-target registry, native 0.438), yet the
transfer remains clearly distinct from native on most held-out survey
views across all three runs. The mean-error comparison itself is symmetric, not a
transfer advantage: native's higher mean comes from 26 small upper-tail
and composition-band targets it overshoots — excluding them drops its
mean to 0.310, matching the transfer's 0.309 — while the transfer has
its own counterpart, 87 targets pinned at exactly 1.0 relative error from
inputs it cannot populate (absent state pension and reported-benefit
leaves), on 35 of which native's imperfect real estimate is better than
the transfer's hard cap. Every residual gap that is not an artifact of
one of these two floors is inherited joint structure that reweighting to
marginal totals could not repair — the populace thesis, measured. The
failure concentrates in the tax-and-benefit view, whose classifier
separates candidate from native at 0.996–0.998 AUC across all three runs, unchanged
since v1.1. One finding cut at the mapping itself, widened under the v1.1 repair,
and was then closed by a declared fix: the transferred file's 99th-percentile
income ratios, near twice native's at v1, reached 59–78× native at v1.1 —
a concept mismatch in directly mapped capital gains that a larger, more
realistic US gains distribution made worse, not better, because US realised
gains are structurally larger than UK CGT-liable gains. Rather than wait
for better source data to fix a mapping problem it cannot fix, v1.2
declares the fix directly: US realised capital gains are excluded from the
transferred file entirely, an explicit, documented mapping decision, not a
second feed repair. Scored against four hypotheses committed to the record
before the run, the same 99th-percentile ratios return to near native
(0.75× and 0.84×) and the net-income view's energy ratio to its floor falls
to 1.5×, while the tax-and-benefit view — never touched by a capital-gains
fix, because capital gains isn't one of its columns — stays exactly where
it was. One of the four hypotheses was refuted outright, not confirmed: a
working theory about why income tax ran high turned out to be wrong, and
that refutation is reported as a finding, not corrected away.
Two defects the referee caught, one repaired feed, one declared fix.
Running the pack as a custodian surfaced specific, mechanical data
errors, not vague quality complaints, twice over. The v1 transferred file's
state-pension total came out at zero against a native total of 112.5
billion pounds — the feed carried survivor, dependant, and disability
social-security streams but no retirement stream at all, so every
household's state pension defaulted to nothing. The cascade was fully
traced: no state pension raised means-tested pension-credit
entitlement nearly tenfold, which in turn drove reported household
benefits far below the native file. A repaired feed and rerun — with
the pre-registered scoring configuration left unchanged — completed as
v1.1: state pension recovered to 91.5 billion pounds against native's
112.5 billion (an 18.6% gap, down from a complete absence), and
universal credit landed within 0.6% of native. But the same corrected
feed carried a far larger US realised-capital-gains distribution into
the transfer's existing one-to-one income mapping, which the UK
market-income surface admits almost none of — income tax rose to 34.2% above
native (was 2.6% below at v1) and household market income to 872%
above native (was 156% above at v1). A repaired feed cannot repair a
wrong mapping; it can only show more clearly what the mapping gets
wrong.
The mapping needed a mapping decision, not another feed repair, so v1.2
declares one: US realised capital gains are excluded from the
transferred UK file entirely — an explicit, documented reclassification,
not a silent patch. Four hypotheses about what that exclusion would do
were committed to the record before the run. Three were confirmed:
household market income flipped sign, landing 31.6% below
native rather than 872% above it — the transfer now understates market
income once the phantom gains are gone, which is itself informative
about what the transfer is still missing. The net-income view's energy
ratio to its sampling-noise floor fell from 12.5× to 1.5×, essentially
closing that view's gap. And the tax-and-benefit view, whose failure
was never about capital gains, stayed exactly where it was: 90.1× its
floor, clearly distinct, because the reported-benefit surface it needs
was never written to the transfer and this fix doesn't touch it. The
fourth hypothesis was refuted: income tax's 34.2% overshoot, which the
v1.1 repair blamed on phantom gains entering taxable income, did not
move at all — UK income tax doesn't tax chargeable gains, and gains
entered no calibration target, so removing them changed nothing on this
aggregate. That refutation opens a new, still-open question: what
actually explains the income-tax gap, if not capital gains. A repaired
feed cannot repair a wrong mapping, and a repaired mapping doesn't
explain every remaining gap — but naming which mechanism did and didn't
move each number, and testing that in the open before looking at the
result, is what a referee is for.
calibration targetsv1 measured
Transferred MARE 0.309
Against the same 568 finite UK targets, the recalibrated transfer
reaches mean absolute relative error 0.309 to the native file's
0.438 — a mean dragged by a small set of outlier band-level
targets (medians: 0.180 versus 0.178) — while the native FRS
holds the tighter central fit: 43.8% of targets within 10%
against the transfer's 29.9%. Marginal attainment
does not certify the joint.
[SOURCE: artifacts/calibration_summary.json]
held-out FRS scorev1 measured
3 of 4 views clearly distinct
Scored on the population-view harness against the untouched native
FRS file: income intermediate, tax-and-benefit, net-income, and the
omnibus view all clearly distinct. The residual is the inherited
joint structure reweighting cannot repair. The repaired-feed rerun
(v1.1, next card) moves this same verdict, on the same view, in two
directions at once.
[SOURCE: artifacts/scorecard.json verdict]
repaired rerunv1.1 measured
2 of 4 clearly distinct — and one energy ratio nearly doubled
The rerun on the repaired feed — retirement stream restored, the
frozen scoring configuration byte-for-byte unchanged so the
comparison stays pre-registered — is complete. The pre-registered
headline improved: net_income's classifier AUC fell from 0.802 to
0.770, crossing back below the clearly-distinct threshold. Its
continuous energy-distance ratio to the sampling-noise floor did the
opposite, rising from 6.3× to 12.5×. Both numbers are real and both
are reported — a band verdict crossing a line is not the same claim
as a continuous score narrowing, and here they disagree. The next
card closes this exact gap with a declared fix.
[SOURCE: artifacts/scorecard_v1_1.json verdict; scorecard_v1_vs_v1_1.json per_view net_income]
declared transformv1.2 measured
Energy ratio 12.5× → 1.5× — the band verdict never moved
One further rerun, isolating a single change against v1.1: US
realised capital gains, previously mapped one-to-one into a UK
income surface that admits almost none of them, are now excluded
from the transferred file entirely — a declared, documented mapping
decision, tested against four hypotheses fixed before the run. Three
were confirmed, one was refuted. net_income's classifier AUC barely
moved (0.770 → 0.772, still intermediate — the pre-registered
headline holds at 2 of 4 throughout), but its continuous
energy-distance ratio to the sampling-noise floor fell from 12.5× to
1.5×, within reach of the floor itself. Read alongside, never apart:
tax_benefit — the view this fix was never going to touch — stays
clearly distinct at 90.1×, unchanged since v1.1, because the absent
reported-benefit surface it needs was never written to the transfer.
[SOURCE: artifacts/scorecard_v1_2.json verdict; scorecard_three_way.json per_view net_income, tax_benefit; policy_outputs_v1_2.json aggregates_bn_gbp]