populace / an application, not a strategy

applied · open microsimulation anywhere

Anywhere.

Most countries' household microdata is restricted. Rules engines are no longer the bottleneck to modeling their tax and benefit systems — data access is. This page is one application built on populace's six strategies: recalibrating populace's openly redistributable US support to another geography's published totals, producing microsimulation inputs there — validated where ground truth exists, deployed where it doesn't.

paper: in progress Read the six strategies
01 — the reframe

Reweighting a fixed sample is adaptation with zero new parameters.

Treat the support — the record set — as a frozen representation, and the weights as a task head fit on top of it. Recalibrating that frozen support to a new geography's published totals is head-only adaptation: every household stays exactly as observed; only its influence on the aggregate changes. Importance weighting of a fixed sample is the zero-parameter-update case of adapting a generative model — the same operation populace already runs to hit a congressional district's targets, run again at a larger domain distance.

A capacity ladder sits above that zero-parameter case: frozen support plus new weights, then frozen support plus imputed extensions for variables the target geography needs and the US support doesn't carry, then a fully adapted generative model that replaces the support itself. Higher rungs spend more capacity and can close gaps reweighting cannot. The mechanism at each rung is swappable — the question this page and its paper ask is whether the adaptation framing survives the climb, not which rung is correct.

02 — one operation, three distances

Reweighting a congressional district, the UK, and Belgium is the same operator.

populace's local-area method already reweights one national file to a US congressional district's published totals — same country, same institutions, the smallest domain shift the operator runs on. Moving the same operator to the UK and to Belgium does not change what it does; it changes how far the source domain sits from the target. The diagram below places three actual runs on that axis.

One reweighting operator at three domain distances A horizontal domain-distance axis from small to large. Three stops sit on it. The first, a US congressional district, reweights one national file to a sub-national geography inside the same institutions — the smallest shift, an established and shipped populace method. The second, the United Kingdom, reweights the US support to UK published totals and is the held-out validation leg: it is scored against native UK survey microdata (the Family Resources Survey). Its first, version-one, evaluation ran three of four survey views clearly distinct from the native file; a repaired version-1.1 rerun moved in two directions at once — the pre-registered headline improved to two of four views clearly distinct, while the same feed upgrade nearly doubled one view's continuous energy-distance gap to its floor; a version-1.2 rerun then applied a single declared mapping fix on top of that repair and closed almost all of that same gap, from twelve and a half times the sampling-noise floor to one and a half times it, while the pre-registered headline held exactly where it was. The third, Belgium, reweights the same US support to Belgian published totals with no comparable native microdata to validate against; it is the zero-shot deployment leg, cross-checked instead against an independent tax-benefit engine, EUROMOD, on an identical population. All three stops are the same operator — reweighting a frozen support to a new geography's totals — differing only in how far the target sits from the source. small domain distance large domain distance same institutions, sub-national shift different institutions, national shift stop 1 · shipped US congressional district local-area method stop 2 · v1.2 measured United Kingdom held-out validation leg 2/4 views clearly distinct scored vs. FRS · energy ratio 12.5x → 1.5x stop 3 · measured Belgium zero-shot deployment leg 21/21 targets within 1.8% no native microdata to score against the operator, unchanged at every stop reweight a frozen support to published totals
Three actual runs of the same reweighting operator, placed by domain distance. The congressional-district leg (populace's local-area method), the UK leg (v1.2 measured against the held-out FRS, below), and the Belgium leg (measured below) all sit on solid ground. The UK leg's held-out verdict — two of four survey views clearly distinct from native microdata, down from three of four on the as-found file and unchanged since the repaired feed — is the paper's thesis made quantitative; the same repair nearly doubled a different view's continuous energy-distance gap to its floor, and a further declared mapping fix then closed nearly all of that gap while the verdict itself never moved — results reported alongside each other, not one in place of the last. Same operator, increasing distance — small and same-institutions at the congressional-district end, large and different-institutions at the Belgium end.
stop 1

Congressional district

Described on the sparsity strategy page and used throughout populace's US releases: one national file reweighted to sub-national published totals. Same country, same tax-benefit institutions, the smallest domain shift this page discusses — referenced here descriptively, not re-validated on this page.

stop 2

United Kingdom

The held-out ground-truth leg. Its first evaluation scored the recalibrated US support against the Family Resources Survey — native UK household microdata withheld from the adaptation step — with three of four survey views clearly distinct from native, and caught a concrete feed defect. A repaired v1.1 rerun moved the headline to two of four views clearly distinct while nearly doubling a different view's energy-distance gap to its floor; a further v1.2 rerun then applied a single declared mapping fix and closed nearly all of that same gap, on the unchanged scoring configuration throughout. Verdict and all three runs below.

stop 3

Belgium

The zero-shot deployment leg, where no comparable native microdata exists to score against. Validated instead by an identical-population cross-engine check: the Axiom Belgian engine against EUROMOD, isolating engine agreement from data quality. Numbers below.

03 — the validation protocol

Validate where truth exists. Deploy where it doesn't.

The protocol borrows directly from held-out benchmark practice: adapt to a target domain without touching its held-out ground truth, then score against that ground truth once. The UK supplies the ground-truth leg because native survey microdata — the Family Resources Survey — exists to hold out and score against. Belgium supplies the deployment leg precisely because that check is unavailable: no comparably granular Belgian household microdata sits behind the recalibration, so deployment is checked the only way available, cross-engine agreement on an identical population, rather than against native microdata.

Calibration forces target-geography totals to match by construction. Reweighting is fit to hit published aggregates, so once it runs, the aggregates match almost by definition. Every residual gap that remains against the FRS microdata — not the totals, the joint distribution underneath them — measures something reweighting cannot touch: the joint structure the recalibrated file inherited from the US support it started from. That gap is the quantitative test of the populace thesis stated on the support strategy page: everything downstream can reweight the support; nothing downstream can repair it. The UK benchmark turns that sentence into a number.

adaptshipped

Head-only, ground truth untouched

The US support is recalibrated to the target geography's published totals without ever consulting that geography's held-out microdata — the adaptation step and the scoring step use disjoint information, exactly as a transfer-learning benchmark requires.

score oncev1.2 done

Held-out FRS, scored after each adaptation

The UK leg scores the adapted file against native Family Resources Survey microdata it never saw. Three runs are complete — two of four survey views clearly distinct from native as of the latest, down from three of four on the as-found file (section 07) — so this is the leg that already measures what reweighting alone, and one declared mapping fix on top of it, can and cannot repair.

deploy anywayshipped

No native microdata, cross-engine check instead

Belgium has no held-out survey microdata to score against, so deployment is checked the way that is actually available: an independent tax-benefit engine, EUROMOD, computing the same quantities on the identical recalibrated population.

04 — beyond the UK and Belgium

Open microsimulation anywhere does not require open microdata anywhere.

The UK leg works because native household microdata exists and can be held out. Most geographies fail that condition — not because no one holds the data, but because the data cannot leave the building it lives in. National statistical institutes and research centres routinely hold exactly the restricted microdata this page has no access to. The protocol above does not need that data to travel; it needs a referee that can.

It requires an open referee. The referee travels even where the data cannot. Instead of asking a data custodian for microdata, ship the custodian a sealed, pre-registered evaluation procedure: the recalibrated support, the population-view harness, and a frozen scoring configuration. The custodian runs it locally, against their own restricted ground truth, inside their own environment. Nothing sensitive leaves the building — only a low-dimensional scorecard comes back: energy distance, coverage, classifier AUC, and the tail block, the same four numbers this page already reports for the UK leg. A scorecard contains no micro records, so publishing it is safe by construction, not by review.

A pre-registered evaluation pack, run inside a data custodian's own environment A protocol-stage schematic, dashed because no custodian run has happened yet. An evaluation pack — the recalibrated support, the population-view harness at a pinned version, a frozen scoring configuration, and input manifests — is pre-registered before any run: its contents are fixed and committed first. The pack then crosses a boundary into a data custodian's own environment, drawn as a barrier neither side crosses, where it runs against restricted ground truth that never leaves that environment. Only a low-dimensional scorecard — energy distance, coverage, classifier AUC, and the tail block — returns across the boundary and is published. No microdata crosses in either direction. protocol stage · not yet run · dashed = not implemented pre-registered · frozen before any run the evaluation pack recalibrated support population-view harness, pinned sha frozen scoring config + thresholds input manifests + entry-point script committed before any run — no result exists yet to steer it boundary neither side crosses inside a data custodian's own environment the pack, run locally scored against restricted ground truth the custodian holds — generic: a national statistical institute, a research centre no microdata crosses the boundary, either direction published the scorecard energy · coverage · classifier AUC · tail no micro records
A protocol-stage schematic (dashed = specified, not run). The evaluation pack — recalibrated support, harness at a pinned version, frozen scoring configuration, input manifests — is pre-registered before any run, so neither side can select thresholds or blocks after seeing a result. It crosses into the custodian's own environment, runs there against restricted ground truth that never leaves, and only the low-dimensional scorecard returns. No custodian run has happened yet; the UK leg above is this pack's dress rehearsal, proving it runs correctly before any third party runs it.
pre-registrationprotocol stage

Frozen before results exist

The harness version, the four scoring blocks, and their thresholds are committed before any custodian run — the same discipline that gives a held-out machine-learning test set its credibility. Freezing first is what stops either side from selecting the comparison after seeing how it turns out.

the eval packprotocol stage

A pinned, portable bundle

Harness commit, frozen configuration, input manifests, and an entry-point script travel together as one pinned artifact — built to run unattended inside an environment its authors never enter, producing a fixed output schema every time.

what it buysprotocol stage

Ground-truth checks without data access

Micro-level validation becomes available in every geography with a willing custodian, not only the geographies whose microdata is already open. The UK leg is this pack's dress rehearsal: proof it runs correctly end to end before any third party is asked to run it against data this page will never see.

The dress rehearsal already earned its keep. Run against the UK's own held-out file — the one geography where this page can score against native microdata — the pack did not just confirm it executes. It caught a concrete, mechanical defect in the transferred file: a missing income stream that zeroed out state pensions and cascaded into the means-tested benefits that depend on them (section 07). A referee that surfaces a specific data error on its first real run is a referee worth shipping to a custodian who holds data this page will never see. The pre-registered pack, run as a custodian would run it, did exactly the job it exists to do.

05 — what this cannot fix

Bounded by what the frozen support contains.

Reweighting can only redistribute mass across the households already in the support; it cannot manufacture a joint structure the US population never exhibits. If the target geography has a household-composition pattern, an income-tax-and-benefit interaction, or a demographic combination absent from the source support, no amount of reweighting recovers it — that is a limit of this rung of the capacity ladder, not a bug in the reweighting method. Higher rungs (imputed extensions, or a fully adapted generative model) exist precisely to spend more capacity against this limit. This page and its paper report where that limit binds, not just where the operator succeeds.

06 — measured now

Belgium: 21 of 21 targets, one engine cross-check.

The US support, recalibrated to Belgian published totals, is the first deployment-leg result this page can report in full. Both numbers below come from a completed populace-be run, not a benchmark still executing.

calibration targets within 1.8% 21/21every target hit Statbel age×sex×NUTS1, SPF fiscal income, ONSS worker SSC, ONEM unemployment
Axiom vs. EUROMOD, worker SSC €643apart on €20.9B identical recalibrated population, independent engines

[SOURCE: populace-be run <sha/date>] — pending pin. The identical-population design is what makes the €643 figure a statement about engine agreement, not data quality: both the Axiom Belgian engine and EUROMOD compute worker social security contributions on the exact same recalibrated households, so the €643 gap on €20.9 billion isolates disagreement between the two engines' rules from any disagreement about who is in the population. This is the deployment leg described above — no native Belgian microdata sits behind either number, so cross-engine agreement is the check available, not a substitute for the FRS-style holdout the UK leg runs.

A proportionality caveat this figure needs every time it's cited. Belgian employee social security contributions are levied at a rate that's largely proportional to gross pay, so agreement this close on an identical population is close to what rate-times-base arithmetic on the same inputs would already deliver. €643 on €20.9 billion certifies concept-base alignment — the two engines agree on what counts as contributable pay — and rules out implementation drift between them, but it's a weak discriminator: a proportional levy has little nonlinearity (thresholds, bands, means tests, caps) for two engines to disagree about. A sharper validation ladder is planned, not yet run: per-record agreement rather than the aggregate alone, extension to genuinely nonlinear programs, and agreement on reform-induced deltas rather than baseline levels — tracked at axiom-rules-engine#77. Until that ladder runs, read this figure as a necessary but weak check, not as evidence the two engines agree on the rules generally.

07 — the held-out leg, v1.2 measured

United Kingdom: the referee ran, diagnosed, reran, and priced a declared fix.

The UK leg has now run three times. Its subject is the prior US-to-UK transfer: an existing recalibrated file built from open US microdata, reweighted to UK national, regional, and country targets — the same target set the native Family Resources Survey file calibrates to. The pre-registered pack scored the as-found v1 file against the held-out native FRS file before that file was touched, exactly as section 03 requires; scored a repaired-feed v1.1 rebuild under the identical frozen scoring configuration; then scored a v1.2 rebuild applying one declared mapping transform on top of that repair, isolated against four hypotheses committed to the record before the run. All three are diagnostic passes, not a finished benchmark: v1 caught a concrete feed defect, the v1.1 repair moved the headline verdict and a continuous proper score in opposite directions on the same view, and v1.2 closed that same view's gap to its floor — reported together below, not as a single improved number.

held-out verdict, four survey views 2 / 4clearly distinct; 0/4 indistinguishable — down from 3/4 on v1, unchanged v1.1 to v1.2 income intermediate; tax_benefit and full clearly distinct (unchanged since v1.1); net_income moved from clearly distinct (v1) to intermediate (v1.1, v1.2)
net_income energy ratio to floor 6.3× → 12.5× → 1.5×doubled, then fell within reach of the floor the same view whose band verdict never moves off intermediate at v1.1/v1.2 — a declared mapping fix on top of the feed repair closed almost the entire remaining gap
sampling-noise floor, classifier AUC 0.49–0.52at chance, all four views, all three runs the native file's own split scored as a candidate — machinery validated
determinism, two independent runs each byte-identicalcanonical subset, v1, v1.1, and v1.2 scores, verdict, and config fingerprint reproduce exactly; only the timestamp differs — eval_config hash unchanged across all three versions

[SOURCE: transfer-paper/experiments/uk-transfer/artifacts/scorecard.json, scorecard_v1_1.json, scorecard_v1_2.json, scorecard_three_way.json, decomp_three_way.json — popdgp 402fb235, period 2025, v1 candidate n=28,532 vs v1.1/v1.2 candidate n=63,128, all vs native FRS n=53,508; verdict bands pre-registered in eval-pack/PREREGISTRATION.md]. This is the quantitative test the framing promised, run three times: calibration forced the marginal targets to move (transferred mean absolute relative error 0.309 against the full 568-target registry, native 0.438), yet the transfer remains clearly distinct from native on most held-out survey views across all three runs. The mean-error comparison itself is symmetric, not a transfer advantage: native's higher mean comes from 26 small upper-tail and composition-band targets it overshoots — excluding them drops its mean to 0.310, matching the transfer's 0.309 — while the transfer has its own counterpart, 87 targets pinned at exactly 1.0 relative error from inputs it cannot populate (absent state pension and reported-benefit leaves), on 35 of which native's imperfect real estimate is better than the transfer's hard cap. Every residual gap that is not an artifact of one of these two floors is inherited joint structure that reweighting to marginal totals could not repair — the populace thesis, measured. The failure concentrates in the tax-and-benefit view, whose classifier separates candidate from native at 0.996–0.998 AUC across all three runs, unchanged since v1.1. One finding cut at the mapping itself, widened under the v1.1 repair, and was then closed by a declared fix: the transferred file's 99th-percentile income ratios, near twice native's at v1, reached 59–78× native at v1.1 — a concept mismatch in directly mapped capital gains that a larger, more realistic US gains distribution made worse, not better, because US realised gains are structurally larger than UK CGT-liable gains. Rather than wait for better source data to fix a mapping problem it cannot fix, v1.2 declares the fix directly: US realised capital gains are excluded from the transferred file entirely, an explicit, documented mapping decision, not a second feed repair. Scored against four hypotheses committed to the record before the run, the same 99th-percentile ratios return to near native (0.75× and 0.84×) and the net-income view's energy ratio to its floor falls to 1.5×, while the tax-and-benefit view — never touched by a capital-gains fix, because capital gains isn't one of its columns — stays exactly where it was. One of the four hypotheses was refuted outright, not confirmed: a working theory about why income tax ran high turned out to be wrong, and that refutation is reported as a finding, not corrected away.

Two defects the referee caught, one repaired feed, one declared fix. Running the pack as a custodian surfaced specific, mechanical data errors, not vague quality complaints, twice over. The v1 transferred file's state-pension total came out at zero against a native total of 112.5 billion pounds — the feed carried survivor, dependant, and disability social-security streams but no retirement stream at all, so every household's state pension defaulted to nothing. The cascade was fully traced: no state pension raised means-tested pension-credit entitlement nearly tenfold, which in turn drove reported household benefits far below the native file. A repaired feed and rerun — with the pre-registered scoring configuration left unchanged — completed as v1.1: state pension recovered to 91.5 billion pounds against native's 112.5 billion (an 18.6% gap, down from a complete absence), and universal credit landed within 0.6% of native. But the same corrected feed carried a far larger US realised-capital-gains distribution into the transfer's existing one-to-one income mapping, which the UK market-income surface admits almost none of — income tax rose to 34.2% above native (was 2.6% below at v1) and household market income to 872% above native (was 156% above at v1). A repaired feed cannot repair a wrong mapping; it can only show more clearly what the mapping gets wrong.

The mapping needed a mapping decision, not another feed repair, so v1.2 declares one: US realised capital gains are excluded from the transferred UK file entirely — an explicit, documented reclassification, not a silent patch. Four hypotheses about what that exclusion would do were committed to the record before the run. Three were confirmed: household market income flipped sign, landing 31.6% below native rather than 872% above it — the transfer now understates market income once the phantom gains are gone, which is itself informative about what the transfer is still missing. The net-income view's energy ratio to its sampling-noise floor fell from 12.5× to 1.5×, essentially closing that view's gap. And the tax-and-benefit view, whose failure was never about capital gains, stayed exactly where it was: 90.1× its floor, clearly distinct, because the reported-benefit surface it needs was never written to the transfer and this fix doesn't touch it. The fourth hypothesis was refuted: income tax's 34.2% overshoot, which the v1.1 repair blamed on phantom gains entering taxable income, did not move at all — UK income tax doesn't tax chargeable gains, and gains entered no calibration target, so removing them changed nothing on this aggregate. That refutation opens a new, still-open question: what actually explains the income-tax gap, if not capital gains. A repaired feed cannot repair a wrong mapping, and a repaired mapping doesn't explain every remaining gap — but naming which mechanism did and didn't move each number, and testing that in the open before looking at the result, is what a referee is for.

calibration targetsv1 measured

Transferred MARE 0.309

Against the same 568 finite UK targets, the recalibrated transfer reaches mean absolute relative error 0.309 to the native file's 0.438 — a mean dragged by a small set of outlier band-level targets (medians: 0.180 versus 0.178) — while the native FRS holds the tighter central fit: 43.8% of targets within 10% against the transfer's 29.9%. Marginal attainment does not certify the joint. [SOURCE: artifacts/calibration_summary.json]

held-out FRS scorev1 measured

3 of 4 views clearly distinct

Scored on the population-view harness against the untouched native FRS file: income intermediate, tax-and-benefit, net-income, and the omnibus view all clearly distinct. The residual is the inherited joint structure reweighting cannot repair. The repaired-feed rerun (v1.1, next card) moves this same verdict, on the same view, in two directions at once. [SOURCE: artifacts/scorecard.json verdict]

repaired rerunv1.1 measured

2 of 4 clearly distinct — and one energy ratio nearly doubled

The rerun on the repaired feed — retirement stream restored, the frozen scoring configuration byte-for-byte unchanged so the comparison stays pre-registered — is complete. The pre-registered headline improved: net_income's classifier AUC fell from 0.802 to 0.770, crossing back below the clearly-distinct threshold. Its continuous energy-distance ratio to the sampling-noise floor did the opposite, rising from 6.3× to 12.5×. Both numbers are real and both are reported — a band verdict crossing a line is not the same claim as a continuous score narrowing, and here they disagree. The next card closes this exact gap with a declared fix. [SOURCE: artifacts/scorecard_v1_1.json verdict; scorecard_v1_vs_v1_1.json per_view net_income]

declared transformv1.2 measured

Energy ratio 12.5× → 1.5× — the band verdict never moved

One further rerun, isolating a single change against v1.1: US realised capital gains, previously mapped one-to-one into a UK income surface that admits almost none of them, are now excluded from the transferred file entirely — a declared, documented mapping decision, tested against four hypotheses fixed before the run. Three were confirmed, one was refuted. net_income's classifier AUC barely moved (0.770 → 0.772, still intermediate — the pre-registered headline holds at 2 of 4 throughout), but its continuous energy-distance ratio to the sampling-noise floor fell from 12.5× to 1.5×, within reach of the floor itself. Read alongside, never apart: tax_benefit — the view this fix was never going to touch — stays clearly distinct at 90.1×, unchanged since v1.1, because the absent reported-benefit surface it needs was never written to the transfer. [SOURCE: artifacts/scorecard_v1_2.json verdict; scorecard_three_way.json per_view net_income, tax_benefit; policy_outputs_v1_2.json aggregates_bn_gbp]

The paper working out the framing.