Skip to content

Run large full cycle-collector passes on idle CPUs in speed builds - #850

Draft
cramforce wants to merge 1 commit into
perf/gc-widen-fasterfrom
perf/gc-parallel-passes
Draft

cramforce wants to merge 1 commit into
perf/gc-widen-fasterfrom
perf/gc-parallel-passes

Conversation

@cramforce

Copy link
Copy Markdown
Contributor

Summary

In executables built with --optimization=speed on Linux and Darwin, a full pass over at least 100,000 live cycle-headered objects runs markRoots, markGray and scan on helper threads together with its own thread. Settling (white filtering, re-buffering, teardowns, re-arming) stays on the owner, so free_fn, weak-reference disposal and FinalizationRegistry notifications run on the same thread at the same point as before. The result is the sequential black/white partition; only teardown order within a pass can differ.

  • Idle CPUs only: the CPUs available to the process minus its running or runnable threads. Linux reads /proc/self/task/*/stat; Darwin uses task_threads + thread_info (new). The count is checked once per large pass. Helpers are created lazily, block all signals, and are capped at 15. One parallel pass runs at a time per process.
  • SCR_CYCLE_PARALLEL caps the helpers (0 turns the walk off); SCR_CYCLE_PARALLEL_MIN sets the floor.
  • Speed builds only: compiled into every executable, the walk added 6-10 KB of x86-64 code, +10.0% on a 99 KB release program (log-lines). The runtime pack's speed flavor now defines SCR_SPEED. Release, dev and library builds compile the walk out.

Measurements

Where large full passes remain, it removes most of their time. With SCR_CYCLE_GROWTH_CAP=1 (5 rounds, parallel vs sequential):

Excalidraw Playwright TypeORM self-check
single-threaded -7.1% -4.6% -8.7% -5.7%
--checkers 8 -1.9% -3.9% -1.9% +3.7%

On the default schedule after PR 1, the remaining full passes are small: full-pass time drops 89 -> 32 ms on the self-check, 130 -> 39 on TypeORM and 104 -> 36 on Excalidraw. The wall-time effect is within noise:

run Excalidraw Playwright TypeORM self-check
single-threaded, 10 rounds +1.4% -2.2% -1.3% -0.7%
single-threaded, 5 rounds (final build) +5.0% -1.2% -1.5% -3.2%
--checkers 8, 10 rounds -1.5% +1.7% +0.5% -1.5%

RSS +0.3% to +1.3%; user CPU +2-3% single-threaded. On #833's widening alone (without PR 1), the prototype gave -1.2% to -6.2% single-threaded (5 rounds).

Sizes: tsc-ts 23,676,592 -> 23,685,776 bytes (+9,184). Runtime suite in speed builds: +1.1% to +9.7% (geomean +4.0%), time geomean +0.1% (all neutral), peak RSS +0.3%. Release builds are unchanged.

Tests

  • test_cycle: randomized graphs (sparse, hubs of 1000-5000 children, long chains cut open; 3 seeds × 20,000 nodes) checked against the test's own reachability and expected counts.
  • cycle.test.ts: the whole suite with every full pass forced parallel. That covers a plain speed build at 3 helpers and a worker build at 1/3/7 helpers (and 0) under ASan+UBSan, plus the worker build under ThreadSanitizer at 1/3/7 helpers. It passes on macOS arm64 and Linux x64.
  • Mutations caught: double blackening, seeding scan from every gray node, whitening black nodes, gray claim by store instead of exchange, and a non-atomic decrement (also flagged by TSan).
  • Docs: NEXT_DIST_DIR=.next-check pnpm check passes (23/23 route tests).

Open question for review

On tsc-ts's default schedule after PR 1, this PR's wall-time effect is within noise. It pays where large passes remain: memory-capped runs and larger heaps. It costs about 430 lines of concurrent code and 4% of speed-build size. Whether to land it is a judgment call.

Draft: after the faster widening (previous PR) this adds little to tsc-ts by default (−2.2% to +1.7%); it pays where large passes remain (−5 to −9% with SCR_CYCLE_GROWTH_CAP=1). Kept ready for memory-capped runs and larger heaps; to be revisited after the nursery-pass work.

@vercel

vercel Bot commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
scriptc Ready Ready Preview, v0 Oct 11, 2026 3:46am UTC

A full pass over a large heap walks every live headered object twice
(markGray, then scan restoring the counts) and is bound by cache misses. In
executables built with --optimization=speed on Linux and Darwin, a full pass
over at least 100,000 live headered objects now runs markRoots, markGray and
scan on helper threads together with its own thread. Settling (white
filtering, re-buffering, teardowns, re-arming) stays on the owner, so free_fn,
weak-reference disposal and FinalizationRegistry notifications run on the
same thread at the same point of the program as before.

The walk computes the sequential partition. Each node is claimed once per
phase by an atomic exchange of its color (read first, so already-claimed
nodes are not dirtied) and each internal edge is trial-deleted once by an
atomic decrement. Scan starts after a phase barrier and is black propagation
from the gray nodes whose counts stayed positive, each node blackened once
(restoring its edges once); whatever is still gray is white. Full passes
skip no generation, so survivors are promoted as they are blackened. Only the
order of teardowns within a pass can differ, and that order is unspecified.
Trace functions run on helpers; they are reads (scr_dyn_trace_v also records
its per-entry edge decision, on the one worker tracing that object).

Helpers come from idle CPUs only: the CPUs available to the process minus
its running or runnable threads, read once per large pass from
/proc/self/task on Linux and from task_threads/thread_info on Darwin, so a
process whose script threads occupy every core collects sequentially.
Helpers are created on first use, sleep between passes, block every signal
and number at most 15; one parallel pass runs at a time per process.
SCR_CYCLE_PARALLEL caps the helpers (0 turns the walk off) and
SCR_CYCLE_PARALLEL_MIN sets the floor.

Why speed builds only: compiled into every executable, the walk added 6-10
KB of x86-64 code (+10.0% on log-lines, a 99 KB release program) to programs
whose heaps never reach it. The runtime pack's speed flavor now defines
SCR_SPEED; release, dev and library builds compile the walk out. Release
builds of the runtime benchmark suite are byte-identical in size to the
parent's. In speed builds the suite grows 1.1-9.7% (geomean +4.0%) with
unchanged run times (geomean +0.1%, every workload neutral) and peak RSS
(+0.3%); tsc-ts grows 9,184 bytes to 23.7 MB. --optimization=speed is
documented to trade size for speed, and the CLI page now names the walk.

This lands the exp/gc-research prototype (88e75d90, 7d527206, cb471cfd,
b1b34a67, 36dbc67e) with a Darwin idle-CPU estimate (the prototype assumed
every CPU idle there), the configuration read once under the busy flag
instead of racily, the pool's idle count updated atomically, signals blocked
in helpers, the worker array moved to zero-initialized data, and the
speed-only gate.

test_cycle gains randomized graphs (sparse, hubs of 1000-5000 children,
long chains cut open; three seeds, 20,000 nodes): after the sweep exactly the
nodes unreachable from owned ones are gone, and every survivor holds exactly
its reachable parents' and owners' count, black and unbuffered; dropping the
owners frees the rest. cycle.test.ts runs the whole suite with every full
pass forced parallel (SCR_CYCLE_PARALLEL_MIN=1) in a plain speed build at
3 helpers and a worker build at 1, 3 and 7 helpers (and 0) under
ASan+UBSan, and the worker build under ThreadSanitizer at 1, 3 and 7
helpers. Mutations each fail: blackening a node twice, seeding scan from
every gray node, whitening black nodes, claiming gray with a store instead
of an exchange, and a non-atomic decrement (also reported by TSan).

Measured with tsc-ts origin/parallel-check e469e33 (--optimization=speed)
against the parent commit, interleaved in Linux Vercel Sandboxes (8 vCPU),
medians; diagnostics and exit codes identical.

Where large full passes remain the walk removes most of their time. With
SCR_CYCLE_GROWTH_CAP=1 (5 rounds), parallel vs sequential:

  single-threaded   Excalidraw -7.1%  Playwright -4.6%  TypeORM -8.7%  self-check -5.7%
  --checkers 8      Excalidraw -1.9%  Playwright -3.9%  TypeORM -1.9%  self-check +3.7%

On the default schedule the previous commit leaves only small full passes
(pass log: full-pass time 89 -> 32 ms on the self-check, 130 -> 39 ms on
TypeORM, 104 -> 36 ms on Excalidraw), and the wall-time effect is within
noise:

  single-threaded, 10 rounds   +1.4%  -2.2%  -1.3%  -0.7%
  single-threaded, 5 rounds    +5.0%  -1.2%  -1.5%  -3.2%
  --checkers 8, 10 rounds      -1.5%  +1.7%  +0.5%  -1.5%
  (Excalidraw, Playwright, TypeORM, self-check)

RSS +0.3% to +1.3%; user CPU +2-3% single-threaded.

This branch was successfully deployed

1 active deployment
Preview — 596cceee Deployed Oct 11, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant