Skip to content

PTY snapshot suite: rare non-repeating case failures under parallel load (pass in isolation) #2565

Description

@simulacre7

Summary

Running the full migration* PTY snapshot suite in parallel occasionally fails a small, non-repeating set of cases (~0.4% of case-executions). Every affected case passes deterministically when run in isolation, and the failing names differ between runs — so this looks like load-sensitivity in the PTY runner environment rather than any individual fixture or product bug. Filing per m0g3r's suggestion in #2483 after it showed up there first.

Observations (macOS arm64, M-series, just snapshot-test migration)

Five full-suite runs across two commits, all with a fresh packages/cli/dist (freshness guard green):

run commit result failed cases
A 31163c5 (PR #2483 branch) 151/154 migration_dynamic_oxc_configs, migration_framework_shim_vue, migration_standalone_yarn4_idempotent
B cc20535d (main) 159/161 migration_not_supported_vitest3, migration_from_tsup_monorepo_success
C cc20535d (main) 161/161
D cc20535d (main) 161/161
E 31163c5, earlier same-day 148/154 different set again (incl. migration_husky_or_prepare, migration_standalone_bun_install)
  • No case name repeats across runs. Union of failures over five runs: 10+ distinct fixtures, each failing exactly once.
  • Every one of them passes in isolation (just snapshot-test <name>), immediately after the failing run, unchanged tree.
  • The failure-heavy runs were the first suite run after a fresh build / cold caches; warm re-runs (C, D) were fully green. Cold-state work on first touch (managed runtime provisioning, registry-bridge first hits) overlapping with ~14 parallel PTY cases is my best guess at the mechanism, but I have not isolated it.
  • m0g3r additionally verified in feat(migrate): preserve dynamic Oxlint and Oxfmt configs #2483 that the flaked fixtures are structurally unremarkable (step counts mid-pack; one of 21 double-vp migrate fixtures).

Why it may matter

Locally it's a shrug; in CI a 0.4% per-case flake across ~650 cases makes a meaningful fraction of runs red for reasons unrelated to the diff under test, and the changing names make it hard for contributors to tell signal from noise (it cost a round of triage on #2483).

Environment

  • macOS arm64 (Darwin 25.2), 14 threads reported by the runner
  • just snapshot-test migration (both flavors), toolchain per rust-toolchain.toml

Happy to re-run with any added instrumentation (timings, per-case retries, runner verbosity) if that helps narrow it — I can reproduce roughly one flaky run per two or three cold full-suite runs.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    Priority

    None yet

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions