How the dataset is built
The full US pipeline as implemented in PolicyEngine/populace (derived from main @ 98d990b). Build-step ids match the staging telemetry, so a running build on the Staging runs page walks these exact steps.
Left to right: survey sources feed the base H5; Ledger facts and references compile the target surface; materialization and calibration produce the dataset and its diagnostics; staging streams live; publish ships to Hugging Face.
Imputation stages declared in source_stages.json — each grafts variables from another survey onto the CPS base. They are baked into the base H5 and are NOT re-run by a fiscal-refresh build.
Cross-referenced against source_stages.json at the commit above: variables flagged by the reform/SOI validation are exactly the ones no enrichment stage declares as an output.
qualified_tuition_expenseseducation credits validate ~40% low · issue ↗qualified_passenger_vehicle_loan_interestOBBBA auto-loan deduction is structurally $0 · issue ↗has_valid_ssn / immigration statusSSN- and immigration-conditioned policy is a no-op · issue ↗childcare expenses (CDCC inputs)CDCC validates ~31% low — under investigation