Repository navigation
Conversation
Contributor
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
cramforce
force-pushed
the
perf/cycle-growth-cap
branch
from
October 11, 2026 03:45
ccc686f to
73a99b2
Compare
Full passes are scheduled by heap growth, and the fraction they wait for widens while growth-triggered passes keep finding nothing. Widening went one doubling per pass after two passes in a row had asked for it, so a heap that grows for a whole run never reached the cap of 16 that #833 introduced: on tsc-ts, runs with seven full passes widened to 9x at the sixth pass, and the seventh still walked 2.7-3.1M objects (294-393 ms) late in the run. The old level now widens two doublings per qualifying pass, never past what the pass's yield asks for, so passes that keep finding nothing reach the cap on the fourth pass instead of the sixth, while the heap is still small. Narrowing is unchanged: a pass that finds old garbage moves the fraction straight to what its yield asks for, and widening still waits for two passes in a row. The mature level keeps a doubling per pass: a mature pass walks only what was promoted since the last one, so its total walking hardly depends on how passes are spaced, and narrower spacing lets less young garbage float. The memory bound is #833's: once passes stop finding garbage, the heap may grow to 17 times its size after the last full pass before the next. It is now reached after four such passes instead of six. Jumping straight to the cap after two was also measured; it reaches the bound at heaps of a few thousand objects and is not better on tsc-ts at --checkers 8. test_cycle pins the schedule from where an old-garbage ring narrows the fraction: the pass after it does not widen, the second pass in a row that finds nothing reaches the cap (one doubling per pass fails here at caps 1 and 16), and a capped pass comes exactly at the cap. cycle.test.ts also runs SCR_CYCLE_GROWTH_CAP=8. Measured with tsc-ts origin/parallel-check e469e33 (--optimization=speed) against the parent commit, interleaved in Linux Vercel Sandboxes (8 vCPU), medians; diagnostics and exit codes identical. The build is the next commit's with its parallel walk off (SCR_CYCLE_PARALLEL=0), which runs this commit's collector. RSS in MB: single-threaded (10 rounds) wall RSS Excalidraw 7.94 -> 7.32 -7.8% 1263 -> 1280 Playwright 7.09 -> 6.90 -2.7% 1457 -> 1470 TypeORM 6.67 -> 6.39 -4.1% 1014 -> 1046 tsc-ts self-check 4.33 -> 4.03 -6.8% 626 -> 622 --checkers 8 (10 rounds, second host) Excalidraw 2.56 -> 2.55 -0.6% 2120 -> 2200 Playwright 2.46 -> 2.39 -3.0% 2649 -> 2686 TypeORM 1.96 -> 1.85 -5.4% 1270 -> 1315 tsc-ts self-check 1.04 -> 1.04 +0.5% 742 -> 744 Pass log (instrumented build, one single-threaded run each): full-pass time 337 -> 89 ms (self-check), 426 -> 130 ms (TypeORM), 323 -> 104 ms (Excalidraw); the largest full pass now walks 1.1-1.3M objects early in the run, and total collector time drops by 27-35%.
cramforce
force-pushed
the
perf/gc-widen-faster
branch
from
October 11, 2026 03:45
686cef0 to
a9d294d
Compare
This was referenced Oct 11, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
#833 raised the growth cap of the cycle collector's full passes to 16, but the fraction still widened only one doubling per pass, after two passes in a row had asked for it. A heap that grows for a whole run gets about seven full passes, so it never reached the cap: on tsc-ts the sixth pass widened to 9x and the seventh still walked 2.7-3.1M objects (294-393 ms) late in the run.
The old level now widens two doublings per qualifying pass, never past what the pass's yield asks for. Passes that keep finding nothing reach the cap on the fourth pass instead of the sixth, while the heap is still small. Narrowing stays immediate, widening still waits for two passes in a row, and the mature level keeps one doubling per pass: its passes walk only what was promoted since the last one, so spacing them out saves little and lets young garbage float.
The memory bound is the one #833 states: once passes stop finding garbage, the heap may grow to 17 times its size after the last full pass before the next. That bound is now reached after four such passes instead of six. Jumping straight to the cap after two passes was also measured. It reaches the bound at heaps of a few thousand objects (it broke an existing test's 4x reclamation window) and was not better at
--checkers 8.Measurements
tsc-ts origin/parallel-check e469e33,
--optimization=speed, Linux Vercel Sandboxes (8 vCPU), interleaved, medians. Diagnostics and exit codes are identical. The build is PR 2's withSCR_CYCLE_PARALLEL=0, which runs this PR's collector.--checkers 8, 10 rounds (second host)Pass log (instrumented build, one single-threaded run each):
Runtime benchmark suite (release,
--layouts=4 --runs=16): sizes +0.00% for every workload, peak RSS geomean -0.17%, time geomean +0.8%. Every workload is neutral except regex-logs +8.0% (CI +6.3..+10.0) and ast-interp +1.4% (slower), and alloc-trees -2.1% (faster). regex-logs spends no time in the collector. On one binary, toggling the widening step (1 vs 2) gives 136.2 vs 136.3 ms in a quiet hyperfine block. The direct prev-vs-PR binaries differ by about +1.5% in hyperfine, and other agents have measured 8% layout-only swings on this workload. I read it as layout.Tests
test_cyclepins the schedule from where an old-garbage ring narrows the fraction: the next pass that finds nothing does not widen, the second in a row reaches the cap, and a capped pass comes exactly at the cap. Widening one doubling per pass fails this at caps 1 and 16.cycle.test.tsaddsSCR_CYCLE_GROWTH_CAP=8. All 20 threshold × cap combinations pass under ASan+UBSan (macOS).Stacked on #833.