Repository navigation
Carry over only inputs, the latest at or before the requested period - #562
Conversation
Auto-carry-over took the latest-starting stored period of any kind and returned the default if it started after the requested period. Any later stored period therefore hid an earlier input: a later input (inputs for 2012 and 2014 gave the default for 2013), or a later period the simulation had already calculated (2014 calculated first made 2013 the default). And values the simulation calculated carried like inputs: a twelfth cached at a month by calculate_divide (the policyengine-us monthly_age bug #557 fixes), a value masked by defined_for, a default cached where defined_for was false everywhere, or a formula result from before the formula's end. Results depended on which periods were calculated first. The rule now: a period takes the input stored for the latest-starting period that starts no later than it (on a tie, the one ending last), preferring the variable's own definition-period unit; with none, the default. Values the simulation calculates are stored with derived=True: the mark lives in the storage with the value (InMemoryStorage and OnDiskStorage keep it per key, every put sets or clears it, it counts only while the key is stored, and clones copy it), so deletions, direct writes and #556's shared arrays cannot leave it stale. Holder.is_derived(period, branch_name) reports the mark of the value get_array reads, so a value one branch calculated never hides another branch's input. A derived value never replaces an input the branch reads (calculate_add/calculate_divide used to overwrite inputs stored in another unit), calculate_add caches only sums over several sub-periods, and dump/restore keeps the marks. Storages gain has(), which neither reads nor copies an array. Supersedes #557's carry-over change: its 253 tests pass here. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Like-for-like PE-US benchmark: core #557 vs #562. Setup: Enhanced CPS 2024, 3,000-household subsample (fixed seed), policyengine-us main 20ccd5ac with and without policyengine-us#9738. Each cell compares all 24 benchmark arrays bitwise (income tax, itemization, both branch liabilities, CTC, EITC, DE/ID/VA and state income tax, household net income, SNAP and more).
So on this model the two core PRs are interchangeable. Both remove the age/12 error, both leave single-year results bitwise unchanged, and both need PE-US #9738 for the second year to match exactly. They differ in scope:
Scripts and arrays: |
…ut-only carry-over Round-2 review findings: - A default cached for a period with no earlier input became the base the uprating path uprated later periods from; the default is not cached for variables with uprating (as before). - A value calculated inside a set_input helper was written to the input's branch and recorded as a user input, bypassing the guard that keeps inputs; derived writes now stay on the branch they were calculated on. - dump_simulation reads each value's mark from the same branch and period as the value it dumps, so #552's branch dumps keep matching marks. - Carry-over checks candidates latest first and stops at the first input, instead of resolving the storing branch for every known period. - Storages drop the marks of deleted keys. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With no input at or before the period but a later period stored, master returned the default without caching it; caching it changed what formulas that test whether a value is stored see (policyengine-uk's maintenance loan and current_education formulas check get_array(period.last_year)), and gave the uprating path a calculated base. Return it uncached there, as master did; cache it (derived) otherwise, as master cached the value it carried. This replaces the uprating-only special case. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…he caches Round-3 review findings: - A period stored only under a branch this one cannot read was taken for an input (own-unit preference ranked it first) and read back as NaN. Holder.get_input_periods(branch_name) now lists, in one pass over the stored keys, the periods whose readable value is an input, taking for each period the key get_array reads first; carry-over picks from those. This also replaces the per-period walk up the branch chain. - apply_reform's cache wipe kept the storages' marks, and rebuilding a disk index brought a stale mark back onto a new input; the wipe now clears the marks and OnDiskStorage.restore starts without any. - Storages pickled before the marks existed failed to clone, put or delete; __setstate__ now defaults them (and #556's shared-key set). - The public carry-over rule now states the own-unit preference and that an input for the period itself is read back as stored. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ebuilds Closes three mutants the suite did not kill. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two inputs covering the same days in units other than the variable's (month:2013-01:2 and day:2013-01-01:59) resolved to whichever was stored first. Prefer the larger unit, then the period's string form, so the choice depends only on the inputs (round-4 review). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s as stored Round-2 review fixes: - An entry is dropped only when neither storage still holds its value, so an input deleted from memory that survives on disk stays an input. - Twelve months starting on the first of a month are recorded as the year storage keys them under, so deleting that year drops the entry. - Deleting compares the keys each storage holds before and after: it no longer goes through the record or loads files for a disk-backed holder. - The property model is seeded from the situation, checks to_input_dataframe itself, and a second property covers memory and disk storage. - Regression for the carry-over case found in the review of #562. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two earlier inputs in a variable's own unit can start on the same day: Holder.set_input stores a year:2012:2 input to a yearly variable as given (a set_input helper is only called for another unit), beside one for 2012. The uprating source was max(..., key=start), so the tie went to whichever was stored first, or to the one in memory over the one on disk: 2015 came out as [1.115, 1.115] or [5.576, 5.576] by set order. Move #562's carry-over sort key into _latest_input_key and use it for the uprating source too: on a tie, the input that ends last, then the larger unit, then the period's string form. Distinct periods never tie, so the source depends only on the stored inputs. The factor still runs from the source's start, and the candidates are unchanged (own unit, starting before the period, inputs only). Tests: regressions for both set orders with and without a set_input helper, monthly inputs, a 10-year input whose string sorts first, memory versus disk and a branch input; Hypothesis properties for set-order and disk-placement independence and for flat-index uprating agreeing with auto-carry-over; the fixture's reference rule breaks ties the same way. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…a set per storage #578 gave a storage a set of shared keys only while it shares an array: an empty set in every InMemoryStorage cost 216 bytes per variable per simulation, about 1.3 MB per policyengine-us simulation, and ran PE-US's CI batches out of memory (#577). This branch's derived marks added a second empty set to every storage, and its __setstate__ gave every unpickled storage its own _shared set, which #578's pickle tests reject. Resolve by keeping #578's code and giving _derived the same treatment: _NOTHING_DERIVED, a class-level empty frozenset, until a storage holds a derived value; its own set from the first derived put; released with one dict pop when the last derived key is replaced by an input or deleted. The class attributes cover pickles from before either set existed, so __setstate__ goes. A clone copies only the marks of keys it holds. apply_reform's wipe releases the marks instead of assigning an empty set. Tests: #578's differential now also stores derived values and compares the effective marks and is_derived with the eager storage; new test_storage_derived_index.py pins the lazy marks (inputs never need a set, release on replace/delete/wipe, clones, copies, old pickles, simulation and apply_reform). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…h clones Measured on a policyengine-us household (policyengine-us 2.18.3, one household_net_income calculation, 6,185 storages, 30,925 holder clones across 5 branches), master with #578 retains 20.94 MB per simulation. With derived marks as a set per storage that holds a derived value, the merge retained 24.35 MB: 1,514 sets in the simulation and a copy of its source's marks in every holder clone. That is 2.6 times the 1.3 MB per simulation that ran policyengine-us's CI out of memory (#577). 5,201 of the 5,204 stored values are derived, so record the inputs instead: is_derived is "stored and not an input". The 3 storages holding inputs have a set (648 bytes); the rest read _NO_INPUTS. A storage and its clones share one frozenset of inputs until one of them changes its own. Retained memory is now 21.25 MB per simulation, +0.31 MB on master: one more 8-byte attribute slot in each of the 37,110 storage objects. A pickle records that inputs were recorded; a state from before (no marker) counts every stored value as an input, as it did then. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The carry-over sort key evaluated period.stop for every candidate input. A period that ends after 9999-12-31 (day:9999-12-30:3, or day:2012-01-01:4000000) has no stop: computing it raises OverflowError. Master carried such inputs ([10, 20]); this branch raised. Sort by _end_order, which puts such a period after every period that has a stop, so it still ends last. Found by the review of #582, which uses the same key for uprating. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two earlier inputs in a variable's own unit can start on the same day: Holder.set_input stores a year:2012:2 input to a yearly variable as given (a set_input helper is only called for another unit), beside one for 2012. The uprating source was max(..., key=start), so the tie went to whichever was stored first, or to the one in memory over the one on disk: 2015 came out as [1.115, 1.115] or [5.576, 5.576] by set order. Move #562's carry-over sort key into _latest_input_key and use it for the uprating source too: on a tie, the input that ends last, then the larger unit, then the period's string form. Distinct periods never tie, so the source depends only on the stored inputs. The factor still runs from the source's start, and the candidates are unchanged (own unit, starting before the period, inputs only). Tests: regressions for both set orders with and without a set_input helper, monthly inputs, a 10-year input whose string sorts first, memory versus disk and a branch input; Hypothesis properties for set-order and disk-placement independence and for flat-index uprating agreeing with auto-carry-over; the fixture's reference rule breaks ties the same way. Review (GPT-6.1 Sol, f051cc2): the key's period.stop raised OverflowError for a period ending after 9999-12-31 (day:9999-12-30:3), where the start-only key did not. The key now orders ends with #562's _end_order, which puts such a period last; regressions use a daily variable. The flat-uprating differential now skips only the case where the two paths really diverge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
6ddbaf9 sorted every period whose stop cannot be computed after those that have one, but equal to each other, so of two inputs ending after 9999-12-31 the string form decided: day:9999-12-30:9 was carried over day:9999-12-30:10, and 24 months over 800 days from the same day (delta review, P2). Month and year periods past 9999 also raised ValueError, or returned a stop no date can hold, rather than OverflowError. _end_order is now the number of the day after the period's last day, counted as date.toordinal counts, by integer arithmetic on the same calendar. It equals stop + 1 day wherever stop has a value and orders every other period by its true end. Tests: the review's two cases in both set orders; a property that _end_order agrees with Period.stop where it has a value and with numpy's calendar everywhere (1,000 examples, sizes to 5,000,000); a longer period from the same day ends later. Also asserts that a pickled state without a set of shared keys gets none (a surviving mutant). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two earlier inputs in a variable's own unit can start on the same day: Holder.set_input stores a year:2012:2 input to a yearly variable as given (a set_input helper is only called for another unit), beside one for 2012. The uprating source was max(..., key=start), so the tie went to whichever was stored first, or to the one in memory over the one on disk: 2015 came out as [1.115, 1.115] or [5.576, 5.576] by set order. Move #562's carry-over sort key into _latest_input_key and use it for the uprating source too: on a tie, the input that ends last, then the larger unit, then the period's string form. Distinct periods never tie, so the source depends only on the stored inputs. The factor still runs from the source's start, and the candidates are unchanged (own unit, starting before the period, inputs only). Tests: regressions for both set orders with and without a set_input helper, monthly inputs, a 10-year input whose string sorts first, memory versus disk and a branch input; Hypothesis properties for set-order and disk-placement independence and for flat-index uprating agreeing with auto-carry-over; the fixture's reference rule breaks ties the same way. Review (GPT-6.1 Sol, f051cc2): the key's period.stop raised OverflowError for a period ending after 9999-12-31 (day:9999-12-30:3), where the start-only key did not. The key now orders ends with #562's _end_order, which puts such a period last; regressions use a daily variable. The flat-uprating differential now skips only the case where the two paths really diverge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Merged at the reviewed head
🤖 Generated with Claude Code |
Two earlier inputs in a variable's own unit can start on the same day: Holder.set_input stores a year:2012:2 input to a yearly variable as given (a set_input helper is only called for another unit), beside one for 2012. The uprating source was max(..., key=start), so the tie went to whichever was stored first, or to the one in memory over the one on disk: 2015 came out as [1.115, 1.115] or [5.576, 5.576] by set order. Move #562's carry-over sort key into _latest_input_key and use it for the uprating source too: on a tie, the input that ends last, then the larger unit, then the period's string form. Distinct periods never tie, so the source depends only on the stored inputs. The factor still runs from the source's start, and the candidates are unchanged (own unit, starting before the period, inputs only). Tests: regressions for both set orders with and without a set_input helper, monthly inputs, a 10-year input whose string sorts first, memory versus disk and a branch input; Hypothesis properties for set-order and disk-placement independence and for flat-index uprating agreeing with auto-carry-over; the fixture's reference rule breaks ties the same way. Review (GPT-6.1 Sol, f051cc2): the key's period.stop raised OverflowError for a period ending after 9999-12-31 (day:9999-12-30:3), where the start-only key did not. The key now orders ends with #562's _end_order, which puts such a period last; regressions use a daily variable. The flat-uprating differential now skips only the case where the two paths really diverge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Break equal-start uprating ties as auto-carry-over does Two earlier inputs in a variable's own unit can start on the same day: Holder.set_input stores a year:2012:2 input to a yearly variable as given (a set_input helper is only called for another unit), beside one for 2012. The uprating source was max(..., key=start), so the tie went to whichever was stored first, or to the one in memory over the one on disk: 2015 came out as [1.115, 1.115] or [5.576, 5.576] by set order. Move #562's carry-over sort key into _latest_input_key and use it for the uprating source too: on a tie, the input that ends last, then the larger unit, then the period's string form. Distinct periods never tie, so the source depends only on the stored inputs. The factor still runs from the source's start, and the candidates are unchanged (own unit, starting before the period, inputs only). Tests: regressions for both set orders with and without a set_input helper, monthly inputs, a 10-year input whose string sorts first, memory versus disk and a branch input; Hypothesis properties for set-order and disk-placement independence and for flat-index uprating agreeing with auto-carry-over; the fixture's reference rule breaks ties the same way. Review (GPT-6.1 Sol, f051cc2): the key's period.stop raised OverflowError for a period ending after 9999-12-31 (day:9999-12-30:3), where the start-only key did not. The key now orders ends with #562's _end_order, which puts such a period last; regressions use a daily variable. The flat-uprating differential now skips only the case where the two paths really diverge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep multi-unit periods off disk on Windows in the tie tests OnDiskStorage names a value's file "{branch}_{period}.npy". On Windows, numpy.save raises OSError for default_year:2012:2.npy and default_year:2014:3.npy, so storing year:2012:2 or year:2014:3 on disk fails there (#526). The tie tests put such periods on disk, which failed the four Windows test jobs: the "longer" memory-or-disk regression (both systems) and the set-order property. The fixture's can_store_on_disk says whether the tests may put a period on disk on this platform. On Windows the "longer" regression is skipped (the "shorter" one still puts 2012 on disk beside year:2012:2 in memory), and the property drops periods with a colon in their string form from its drawn on-disk set, so memory-against-disk ties are still drawn there, with the winner in memory. Elsewhere nothing changes: the derandomized draws are the same. Checked with a plugin that makes numpy.save and tempfile.mkstemp reject any file name containing ":": without the guard the same 3 tests fail with the same paths as CI; with it 3 pass and 2 skip. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Probe disk storage for the Windows guard; test an input on disk winning a tie Review of 6148851 (APPROVE WITH NITS): - The guard keyed on sys.platform, so it would not lift when #526 lets OnDiskStorage store a period with a colon in its name. The fixture now tries storing year:2012:2 in a temporary directory once, at import, and DISK_STORES_ANY_PERIOD says whether that worked. - The comments said "a period several units long", but the code checks for a colon in the period's string form (year:2012-03 has one; month:2012-01:12 prints as 2012 and has none). They now say so. - No property example had the winning input on disk and a losing one in memory: in the drawn ties split between memory and disk, the winner was always in memory, which get_known_periods lists first. A new example puts year:2012:2 on disk beside 2012 in memory. Where storage cannot store year:2012:2 it is an expected failure with OSError. There, no tie this property draws can have its winner on disk. Two inputs tie when the key ranks them equal up to their start: the same start, and both in the variable's own unit or both not. This property's inputs are years and months, so the two are in the same unit, and the winner, which ends last, is the longer; every such period it draws has a colon in its string form. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Word the Windows guard by what CI showed; require Hypothesis 6.62 Review of 3f222e1 (APPROVE WITH NITS): - The fixture said Windows does not allow ":" in a file name, so numpy.save raises OSError there. CI showed that for default_year:2012:2.npy, a name with two colons. The docstrings and comments now name what the probe stores (year:2012:2) and say what the tests do where it fails. - example(...).xfail, which the disk-winner example uses, needs Hypothesis 6.62.0 or later: under 6.61.3 it raises AttributeError, so the module could not be collected with an older Hypothesis the dev extra allowed. The dev extra now requires hypothesis>=6.62,<7, and uv.lock's specifier matches. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Order inputs ending after 9999-12-31 by their end in the reference rule Reviews of 578911f (Opus 5.5 and GPT-6.1 Sol, both APPROVE WITH NITS): - The reference rule ordered ends by Period.stop, which raises OverflowError for days that end after 9999-12-31, and gave every such input the same end, after any input whose stop has a value. So two of them fell through to string order, and one beat a month input that ends later (month:9999-01:13's stop is 10000-01-31). The simulation orders them with _end_order. For uprated_daily_flat at 9999-12-31, {day:9999-12-30:10, day:9999-12-30:3} gave [10, 20] from the simulation and [1, 2] from the reference; for carried_any_unit at 9999, {month:9999-01:13, day:9999-01-01:370} gave [1, 2] and [3, 4]. The reference's last_day now numbers each period's last day itself: for days, the start's number plus the days after it; for months and years, stop, moved back by whole 400-year Gregorian cycles (146097 days) to a year datetime has. - Tests: both cases as regressions, in either set order, compared with the reference; and a derandomized property that last_day, written separately, plus one equals _end_order, for days, months and years up to 5,000,000 units long. Against the old reference the regressions fail 6 tests; without the 400-year shift, 3; with last_day off by one, the property fails. - The fixture said that where the probe fails the tests keep every period with a colon in memory. There, the "longer" regression is skipped and the disk-winner example expects the OSError; it now says so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Say in the changelog that the tied inputs are in the variable's own unit Review of f081b7c (GPT-6.1 Sol, APPROVE WITH NITS): the fragment said "two earlier inputs start on the same day", but uprating reads only inputs in the variable's own unit. A yearly variable with no set_input helper, given 2012 and month:2012-01:25, which start together, is still uprated from 2012, though the month input ends later. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Brings in #562 (carry over only inputs), #563 and #582 (uprate only from inputs), #578 (shared-key sets), #574 (per-call parameter tracing) and #584. One input record instead of two: the storages' `_inputs` (memory) and `_derived` (disk) from #562 replace this branch's `_input_keys`, so an input is a value stored without `derived=True`, as on master. Sequence numbers stay alongside. Disk storage takes master's `_path_to_write` (a clone never writes over a file it shares) in place of this branch's per-store file names, so the process tokens and restore ordering go. Dumps write master's `derived_periods.txt` (branch-aware) and restore inputs first, then calculated values under one later number. `apply_reform` marks every value set_input did not store as derived and drops it, keeping master's semantics without wiping and replaying. Tests that expected a calculated result to carry over now check that the result is cached, since only inputs carry. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolve simulation_dumper.py: keep master's derived_periods.txt marks (PolicyEngine#562) and dump the values the simulation's branch reads, reading each value's derived mark on that same branch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Conflicts resolved by keeping master's bookkeeping: - holder.py: master's derived put (PolicyEngine#562) and branch helpers; this PR's fast-cache eviction after every storage write and delete. - simulation.py: master's input-only uprating (PolicyEngine#563/PolicyEngine#582); this PR's calculate-time value-type check runs just before an input is uprated. - simulation_dumper.py: both records for now (inputs.txt and master's derived_periods.txt); a value marked derived restores as derived. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brings in PolicyEngine#562, PolicyEngine#563, PolicyEngine#582 (input/derived storage marks, input-only carry-over and uprating), PolicyEngine#574, PolicyEngine#584 (parameter tracing and copy guards), PolicyEngine#578 (shared keys only while sharing) and PolicyEngine#585 (temporary storage directories). Conflict in holders/holder.py, resolved by keeping master's bookkeeping: - _set keeps master's ``derived`` mark (stored with the value, and a derived value is never redirected to the running set_input's branch) and this branch's ``is_input``, which decides what the record names. - put_in_cache keeps master's guard (a derived value never replaces a readable input) and passes ``is_input=False``. - master's ``Holder._stores`` (storage ``has``) replaces this branch's equivalent ``_stores``. simulation.py merged cleanly but copied the record in ``clone`` twice (PolicyEngine#562 added the same copy); kept master's copy and added the ``_user_input_contexts`` reset next to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Reconcile #552 with master after the merge: - Holder.get_known_periods(branch_name) uses master's _readable_branches (added by #562) instead of a second copy of the lookup, and lists each period once. get_array is master's again. - _dump_holder reads the dumped simulation's branch from the holder, as #560 does, and keeps master's derived_periods.txt with each value's mark read on that branch. - The _calculate comment says what scoping still changes now that uprating and carry-over read only inputs (#562/#563): a later period stored under an unreadable branch no longer leaves a default uncached. Tests: the isolation grid now also compares what the branch then reads (periods and derived marks); a period stored under several readable branches is listed once; a restored branch dump calculates what the branch calculates; and a Hypothesis differential checks get_known_periods against get_array and get_input_periods against is_derived for random values on a five-branch lineage, in memory and on disk. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Branch and copy cleanup (findings 1 and 4). A cut in a simulation's calculation no longer deletes its branch-name key from every descendant, which deleted a descendant's own input stored under the same name. A copy (clone or branch) instead drops, when it is made, the values the simulation it copies is waiting to purge: every such value is marked by the time it is cached, and a copy never sees what is cached after. Purges delete only values the simulation calculated (derived), never an input. Swallowed requests (2). Each request to finish a deeper period first is recorded as outstanding when raised. A formula that returns while one is outstanding (a bare except caught it) has its value dropped and the request raised again; an error raised while one is outstanding is treated as the request. requires_computation_after (3). A deferred period keeps the names of the variables that needed it, so its caller still counts while it is calculated with an empty stack. Tracing and storage (6, 7). The flat trace prefers a completed node to an abandoned one. Deferred values are kept for the outermost calculation in its own cache, which _calculate reads, so deferral works the same traced or not and when the holder stores nothing; values calculated under a cut are marked for purging when written to the fast cache too. Termination (8). MAX_SPIRAL_DEFERRALS is gone: it cut long chains depending on what was cached. Each deferred period starts at most once per outermost calculation, and only on the requested simulation or its registered branches (a copy made in a formula is made anew on every retry), so the calculation still ends. Finding 5 is settled by master's #562 (only inputs carry over); its regression test passes on the merge commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Keep the record of set_input values in step with storage Simulation._user_input_keys records each (variable, branch, period) stored through set_input. _invalidate_all_caches (run by apply_reform) keeps the values it names, to_input_dataframe exports them, and country packages read it to tell an entered value from a calculated one. It drifted from storage in two ways (#559): - delete_arrays deleted the values but kept their entries, so a formula result calculated later for the same period counted as an input: it survived apply_reform and was exported. Holder.delete_arrays, which Simulation.delete_arrays calls for each branch it deletes from, now drops the entries for the variable, that branch and the periods in-memory storage deletes (all of them for an eternal variable). Code that deletes through the holder, as country packages do when they move an input to another variable, is covered too. Disk storage deletes only the period asked for, so the entry for a value it still holds is kept. - clone (so also get_branch) shared the record between simulations that store their values separately, so an input set on a clone, on a branch's parent after the branch was made, or on the original after cloning was recorded for both. The copy now gets its own record, and its own empty list of running set_input calls. Tests: example regressions (11 of 12 fail before the fix; the twelfth guards against dropping too much), and a Hypothesis property that runs random set_input / calculate / delete_arrays / clone / get_branch / _invalidate_all_caches sequences against a reference model of each simulation's inputs. Fixes #559 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Prune the record from what storage deleted; record each value once Follow-up to the review of the first commit: - Holder.delete_arrays no longer looks through the whole record. It compares the periods memory stores for the branch before and after the deletion and discards exactly those entries, so its cost does not grow with the record. Country marginal-rate code deletes every variable on a branch: with 9,000 entries and 3,024 variables the loop took 0.73 s with the scan and takes 0.018 s now (0.004 s without any pruning). A holder with disk storage, which cannot list its periods for every branch name, still looks through the record for the variable's entries in the deleted periods and drops those whose value neither storage holds. - Holder._set records the period as storage keys the value: eternity for an eternal variable whatever period it was set for, and a Period for a handler that passes a string. Each entry names one stored value, so a string-period input is exported and deleting an eternal input drops its entry. - put_in_cache stores with is_input=False (the same keyword as #560), so values a custom set_input handler calculates are formula results, which apply_reform recalculates, not inputs. - subsample starts the record again before it rebuilds the simulation, so it records only what the rebuild stores. Tests: a guard that a memory-only deletion never iterates the record; the eternal, string-period, calculating-handler, disk and subsample cases; the property's model now includes a handler that calculates and stores months under string periods. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep the entry of an input a storage still holds; record twelve months as stored Round-2 review fixes: - An entry is dropped only when neither storage still holds its value, so an input deleted from memory that survives on disk stays an input. - Twelve months starting on the first of a month are recorded as the year storage keys them under, so deleting that year drops the entry. - Deleting compares the keys each storage holds before and after: it no longer goes through the record or loads files for a disk-backed holder. - The property model is seeded from the situation, checks to_input_dataframe itself, and a second property covers memory and disk storage. - Regression for the carry-over case found in the review of #562. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Find the keys a delete removed without a second snapshot Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Track actual input provenance across storage tiers and record early periods safely Preserve replayed-input tuples and master carry-over marks; add r3b soundness witnesses and value-aware disk properties. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Preserve registered input tiers and master holder call semantics Seed existing recorded input locations before writes, retain positional derived semantics, and validate period aliases and early-year replay and deletion. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Exercise twelve-month normalization before input handlers Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Clarify custom input handler cache contract * Forget restored inputs deleted in early years * Use portable compound-period disk filenames * Replay only surviving supplied input tiers after reforms * Check live and restored input provenance against independent references --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Problem
With
auto_carry_over_input_variables = True(policyengine-us, -uk, -canada), a variable with no formula result for a period takes a value from another period.Simulation._calculatetook the latest-starting stored period of any kind and returned the default if that period started after the requested one. So the result depended on which periods had been calculated first:calculate_dividecarried into later years (the policyengine-usmonthly_agebug Carry over only values stored at the variable's definition period #557 fixes). So did a value masked bydefined_for(its mask came along), a default cached wheredefined_forwas false everywhere (it replaced the input), and a formula result from before the formula'send.Executed on master b78b0ba vs this branch. Each row compares a simulation that calculated something first with a fresh one asked only the target (one-person YEAR variables; script
explore.pyin the review folder):defined_forfalse for person 0 in 2013 only; ask 2013, then 2014defined_forfalse for everyone in 2013; ask 2013, then 2014A policyengine-uk household from another session shows the same thing on PE-UK main: with
regionset for 2025, asking 2040 and then 2026 gives LONDON (the default) for 2026; asking 2026 first gives SCOTLAND. On this branch both orders give SCOTLAND.Intended rule (from history)
sorted(known_periods)[-1].if last_known_period.start > period.start: return default, with no test or changelog entry. Reading it as "don't carry a later input backwards" fits; but because it only looked at the single latest period, it also blocked forward carry whenever anything later was stored, calculated values included.period.startand passed the branch through; kept the guard.No country package relies on "a later stored period means default" in its tests: the PE-US and PE-UK YAML suites were run with a shadow core that logs every carry-over decision where this rule and master's differ (results below).
Rule now
A period with no formula result takes the input stored for the latest-starting period that starts no later than it (on a tie, the one that ends last, then the larger unit), preferring the variable's own definition-period unit and using another unit only when there is none (as #557 does, and as before 3.24.0); masked by
defined_forat the period; with no such input, the default. An input never carries backwards. Values the simulation calculated never carry. This is documented onTaxBenefitSystem.auto_carry_over_input_variables.Fix
Simulationcaches (formula, carried, uprated and default values;defined_fordefaults;calculate_divide's month twelfth;calculate_add's multi-period sums) is stored withderived=True.InMemoryStorageandOnDiskStoragekeep the mark per storage key: everyputsets or clears it,deletedrops it,clonecopies it (including Share cached arrays with simulation branches and copy them on first read #556's shared-array clones),apply_reform's cache wipe clears it,OnDiskStorage.restorestarts without it, and storages pickled before it existed default to none (__setstate__). A value calculated while an input is being set (by aset_inputhelper that calculates) stays on its own branch and is not recorded as an input.Holder.get_input_periods(branch_name)makes one pass over the stored keys and, for each period, takes the keyget_array(period, branch_name)reads first (the branch, itsparent_branchancestors,default; memory before disk), keeping it if it is not derived. A period stored only under a branch this one cannot read is never taken for an input, and no per-period branch walk is needed.Holder.is_derived(period, branch_name)answers the same question for one period; both read no array, so branches still copy only what they read (Share cached arrays with simulation branches and copy them on first read #556).get_array(period.last_year) is not None) see what they saw before and the uprating path never uprates from a cached default.put_in_cache(..., derived=True)stores nothing when the branch already reads an input at that period.calculate_add/calculate_divideused to overwrite inputs stored in another unit (helper-less variables) or over several years;calculate_addnow caches only sums over several sub-periods (a single sub-period is already stored bycalculate).derived_periods.txtper variable), read in the same loop and on the same branch as the values dumped (so Keep branch period reads, dumps, and disk deletion consistent #552's branch dumps, once merged, keep matching marks: the merge conflicts there on purpose).Relation to open PRs
known_periods, which this filter reads. The only textual conflict is in_dump_holder, where the resolution passes Keep branch period reads, dumps, and disk deletion consistent #552's branch tois_derivedas it does toget_array.set_inputinvalidation): complementary, and it also records provenance in the storages (is_input). The two flags sit side by side (derivedhere: the simulation calculated it;is_inputthere:set_inputstored it), and the merge conflict input/_setis mechanical: keep both keyword arguments and both sets.defined_formasks compound when a later year is uprated from a calculated intermediate year); Uprate only from inputs, so uprated values don't depend on calculation order #563 uprates only from inputs withHolder.is_derived.Invariants
For every input set and every sequence of earlier requests (
calculatein either unit,calculate_add,calculate_divide, on the simulation and on branches forked from it):tests/fixtures/carry_over.py::reference).calculate_add/calculate_divide, every input a branch reads is still stored and still carries.Outside this PR's invariant (pre-existing on master, filed separately): uprating (#563, stacked on this PR);
ADD/DIVIDEresults read back by STOCK variables; the spiral depth cutoff;subsampleexporting calculated values as inputs; on-diskderivativeoverwriting the parent's input file;_fast_cacheeviction on period containment; deleted inputs left in_user_input_keys; on-disk keys with_in the branch name (#552).Tests
tests/core/test_carry_over_order.py: 56 regressions, no Hypothesis import; on master b78b0ba most fail, and those that pass are controls for behaviour that must not change (no backward carry,ADDover an input period, an off-unit input starting with the period, the year-first tie, reform replay of an ancestor input, the copy-on-write read count, defaults as master cached them).tests/core/test_carry_over_order_property.py: the Hypothesis property for invariants 1 and 2, 400 derandomized examples plus an@exampleper regression shape, with branch-side requests and variables with noset_inputhelper. Skipped where Hypothesis is not installed (the smoke job).tests/fixtures/carry_over.py: shared variables and the reference rule.mutation_check10.pyplus targeted reruns, 33 mutants of the carry-over rule,get_input_periods, the derived marks in holder and both storages, the input guard,ADD/DIVIDE, default caching, the cache wipe, index rebuilds, old pickles and dump/restore): all 31 non-equivalent mutants killed, 12 by the property alone (the rest pin branch, storage, dump and caching paths outside its domain). The 2 survivors are equivalent for current callers: caching a single sub-period'sADDresult as derived, andhas()implemented throughget.uv run pytest tests: 1200 passed, 4 skipped, 1 xfailed (local, Python 3.14t). CI runs on Ubuntu and Windows, Python 3.11-3.14, and the smoke job.ADD/DIVIDEoverwriting inputs (multi-year, other-unit, anchored, in branches), per-branch provenance, provenance after deletion, uncached defaults, a cached default becoming an uprating base, ties, dump/restore and its branch, the input-helper branch context, copies of Share cached arrays with simulation branches and copy them on first read #556's shared arrays, the per-decision scan cost, the smoke-job import, the property's weakness, the changelog length, unreadable branches, marks across the cache wipe, index rebuilds and old pickles. Pre-existing issues outside this change are filed separately (listed under Invariants). Round 4 (GPT-6.1 Sol, on the round-3 fixes): APPROVE, with two non-blocking notes. First, a P3 performance tradeoff:_calculatedecodes stored keys twice per decision (get_known_periodsandget_input_periods), which costs extra only for histories beyond the period parser's cache (about 1,024 periods); left as a follow-up. Second, an insertion-order tie between equal-extent inputs, now broken by unit.Measured impact
Every number below is from a real run: policyengine-us 6c5170fd and policyengine-uk 7b9fc379 (both main on 2026-10-01), core master b78b0ba versus this PR (35e20d9; later commits are tests only). Scripts, logs and arrays are in the review folder (
us_bench.py,uk_bench.py,compare.py,full/,round5/).PE-US, full Enhanced CPS, 2026, single year: 23 of 24 benchmark arrays bitwise identical: income tax $2,141.0853bn, state income tax $521.6485bn and household net income $14,462.0966bn in both, and every other tax, credit, SNAP and household array. One differs,
spm_unit_net_income, for 14 SPM units (6,223 weighted), +$3.01m. That is a fix:aca_magi_fractionreads the prior-year guideline,tax_unit_fpgfor 2025, which readsstate_fipsfor 2025. The dataset storesstate_fipsfor 2024, and the run caches 2026 first. On master, the cached 2026 blocks carry-over to 2025, so every household gets the defaultstate_fips6 (California) for 2025.tax_unit_fpgfor 2025 changes for 320 SPM units,aca_magi_fractionfor 292, andaca_ptcfor 21 tax units: +$3.38m, from $40.8246bn to $40.8280bn (7,130 weighted). Medical out-of-pocket spending net of the credit falls $3.01m for 14 SPM units, raising their SPM net income by the same amount.aca_ptcenters household health benefits, not household net income.PE-US, 3,000-household subsample, 2026: 25 of 25 arrays bitwise identical (including marginal tax rates and the itemization branches).
PE-UK, full Enhanced FRS 2024-25, 2026: 36 of 36 arrays bitwise identical (household net income £1,757.970bn, Universal Credit £78.494bn, income tax £313.139bn). The shadow log recorded 6 carry-over decisions, 3 with a different source and none with a different value.
YAML suites under the shadow core:
ssi_lives_in_medical_treatment_facility,ssi_medicaid_pays_majority_of_care,receives_ssi.policy/reformin one process, run out of memory on stock master too (reform/passed a 12 GB footprint within 205 s on b78b0ba; both cores grew about 0.6 GB per completed reform test: 12.0 GB after 21 tests on master, 8.1 GB after 13 on this PR). A separate investigation into the YAML test runner's memory growth owns that; it is not caused by this PR.Multi-year in one simulation (3,000-household subsample):
ParameterNotFoundError ... md.msde.ccs.payment.informal.rates.UNKNOWN): in theitemizing/not_itemizingbranches forked at 2025 and reused at 2027, 27 Maryland households getcountyUNKNOWN.countyformula. It emulates carry-over by returningholder.get_array(sorted(holder.get_known_periods())[-1])from the default branch, which inside a branch reused across years returns None.county_fips, notcounty.Behaviour changes
get_array(2012)is set; master returned the default uncached. With no input to carry, caching is exactly as on master.set_input = None.After #578 (head
2580d43b)#578 merged while this PR was open and conflicted with it in
in_memory_storage.py. The Fix section above describes the head round 4 reviewed (912a4ac5). Three commits follow it. None changes which input carries for a period with a finite end.12445111: merge of master. Give a storage a set of shared keys only while it shares an array #578 gives a storage a set of shared keys only while it shares an array, because an empty set in everyInMemoryStoragecost about 1.3 MB per policyengine-us simulation and ran PE-US CI out of memory (3.32.12 gives every holder's storage an empty set: about 1.3 MB more per policyengine-us simulation #577). This PR had added a second set to every storage and a__setstate__that gave every unpickled storage its own_sharedset, which Give a storage a set of shared keys only while it shares an array #578's tests reject. The merge keeps Give a storage a set of shared keys only while it shares an array #578's code and drops that__setstate__.d89a2ea9: the in-memory storage records its inputs, not its derived values.is_derivedis now "stored and not an input". A simulation calculates far more values than it is given: in a policyengine-us household, 5,201 of 5,204 stored values are derived. A storage has a set of inputs only while it holds one, and it shares one frozenset with its clones until either changes its own. A pickled state records that inputs were recorded; a state from before counts every stored value as an input, as it did then. The on-disk storage keeps its set of derived keys.6ddbaf9f,2580d43b: periods ending after 9999-12-31. The carry-over key calledperiod.stop, which raises for a period that ends after the last datedatetimehas (day:9999-12-30:3), where master carried the input. The key now uses_end_order: the number of the day after the period's last day, by integer arithmetic. It equalsstopplus one day whereverstophas a value, and orders every other period by its true end.Retained memory per simulation. Measured on a policyengine-us 2.18.3 household with one
household_net_incomecalculation: 6,185 storages, and 30,925 holder clones across 5 branches.The remaining 0.31 MB is one more 8-byte attribute slot in each of the 37,110 storage objects.
Results unchanged. policyengine-us, eCPS 2024 (us-data 1.112.3), 3,000-household subsample, 2026: 25 of 25 arrays bitwise identical between master
cbfdedf7and6ddbaf9f.2580d43bchanges only_end_order, which the round-2 review checked against the old key on 440,000 candidate sets of positive-sized periods with a finite end: all matched.Tests.
tests/core/test_storage_input_index.pypins the input index.test_storage_shared_index_differential.py) now also stores derived values and comparesis_derivedwith a storage that always had both sets.Period.stopand against numpy's calendar.2580d43b: 1252 passed, 4 skipped, 1 xfailed. Country-template YAML: 39 passed. CI: 18 of 18.Reviews of the delta (GPT-6.1 Sol, hard tier):
6ddbaf9f: REQUEST CHANGES. Periods ending after 9999-12-31 all sorted equal. Fixed in2580d43b. It found the storage change sound over 60,000 operations against an independent model.2580d43b: APPROVE WITH NITS. It checked 574,173 endpoints against an independent calendar model, and every earlier surviving mutant is now killed.Periodcontract of positive sizes) now ends before a one-day period, wherestoptreated it as one day.axiom: n/a: engine fix to carry-over in policyengine-core; no policy rule changes
🤖 Generated with Claude Code