cranelift: per-table mutability tracking + call_indirect elisions on immutable funcref tables#13445
Conversation
213baca to
a40c9b3
Compare
|
@matthargett could I request that (per our AI policy) you rewrite the PR description here? In particular, there are a bunch of phrases here that seem to be a locally-evolved jargon and are not very comprehensible:
These are just examples; in general I'd like to see a description of the work that is aimed to actually communicate to a human, not dump a bunch of out-of-context details, especially before diving in to a 2k-line PR. (And, to double-check per our AI policy: have you reviewed this whole PR by hand before posting it?) |
I did edit the PR description, by my own typing hands, and it folds in the feedback given about both the verbosity and the AI policy in the chat.
sorry, this is a term we used in the code for my product called BugScan back in 2003. in that and this context, analyzing code to derive predicates that can be chained together for solving/elision purposes. here it's bits being set and not structs/objects, so "factory" wasn't a good metaphor choice. This PR does the full chain: it does the analysis, sets the bits, and then the cranelift modifications uses those bits to do some elision of opcodes based on the proof of analysis that the predicates. The reason there's a split between the PRs is that while there's some low-level CPU counters and profiler stats that improve just with this PR, but wallclock/e2e results on my devices didn't show an above-noise uplift.
the elision floor is when I wasn't getting any movement, even on CPU counter improvements, from trying to elide even more instructions. I'm not trying to use confusing metaphors on purpose, and I'm sorry my phrasing wasn't clearer.
the "discarded" metric comes from the XCode profiling tools, specificallt xctrace's template for CPU bottlenecks. it's how many branch prediction mis-predicts happen, and its relevant in a bunch of interpreter performance work I've done over the years (starting with Tcl on the DEC Alpha and PowerPC).
sorry, I really did try to make it more straightforward and not repeat things the diff already communicates.
I re-reviewed diffs inbetween each on-device benchmark pass, which took 30-40 minutes across the cross-section of physical devices I have at my house. I know I'm not the world's best programmer or communicator, even after a few decades of practice. If you feel like a video/audio chat would help, I'm up for it. If it's just too much overhead for you all, that's okay: I can keep a clean patch stack in my fork and we can revisit (or not) at your leisure. |
|
I did just notice that when cleaning up my fork's branch and splitting the PR that I lost some cleanup commits from my local clone. I'm fixing that now. |
a40c9b3 to
b2ab608
Compare
|
Would you be amenable to splitting up this PR into separate PRs for each optimization applied here? The optimizations here look sound to me and reasonable to implement, but I find it a bit difficult to consider them all at once vs one-at-a-time. Some high-level comments I have on this are applicable to this area, for example:
At a high-level, as well, can you describe the shapes of modules that would benefit from these sorts of optimizations? For example I would expect constant-index dispatches to be optimized by the frontend, the entire table sharing one signature to be relatively rare, and the null-check not mattering much in practice since it's something where we just catch the fault. For the bounds-check I believe we already optimized fixed-size tables (statically fixed-size at least via the type) and the lazy-init changes I'm not totally sold on yet myself (e.g. would want to benchmark/analyze more). Overall it feels like a pretty slim shape of module that would fit within these constraints, but that doesn't meant that they're not important. The optimizations here are pretty easy to read and reason about, so that's why I'm curious to understand more about this use case. |
b2ab608 to
8ca462d
Compare
eb215db to
1a13468
Compare
if you're comfortable with merging individual pieces that may not have e2e benchmark uplift individually, that's fine by me :) based on the changes and my fresh reading since i originally proposed this, it seems like dropping the constant-index to direct-call rewrite entirely is a good idea. you're right that frontends get there first. I checked my whole benchmark corpus (18 modules: Rust/LLVM, AssemblyScript, Porffor, Emscripten-built sqlite3) — of 912 call_indirect sites, zero have a constant index. they're all fed by loads (vtable slots) or locals. Binaryen's Directize pass already does this transform toolchain-side where it's already provable.
ok, I see it, and it's slightly more than a refactor by my eye. the mutability bit currently lives on ModuleTranslation (compile-only), so consolidating means promoting it to a serialized field on Module next to table_initialization, with helpers like table_is_immutable(TableIndex) / static_funcref_image(TableIndex) replacing the three hand-written predicate chains in func_environ.rs. is that what you were thinking?
no problem, I just found that in the devirtualization optimization that I helped shepherd into GCC in ~2010, by the time an e2e test failed, the individual pieces of plumbing that slowly drifted in multiple areas represented a nested problem where each fix risked other regressions. (GCC ended up disabling our test cases one by one, refusing to revert changes that verifiably caused regressions, and LLVM eventually caught up around ~2016 and blew past durably.) looking closer at the coverage turned up real gaps: nothing in the tree (pre-existing or added here) runtime-tests table.grow-then-call_indirect into the grown region, or table.set of a null/wrong-signature entry followed by a dispatch that must trap. I've written a *.wast covering those must-not-fire shapes (including an exported table clobbered by a second module through the import, which is the wast-expressible analog of host mutation) — it passes on all four engine configs against this branch with the last commit I pushed. lmk if you have other test scenarios in mind. something else I noticed: wasmparser already defines the shared-everything-threads table.atomic.set / table.atomic.rmw.* operators. They can't reach translation today (wasm_unsupported!), but the analysis will match them anyway so it can't silently go stale if that proposal lands. I was thinking/reaching ahead a bit, but a nice effect of how things are organized.
ok.
your suspicion is right and I have som first numbers. As written (serial OperatorsReader::read over every body), the scan costs ~83% of a full serial validate_all on Emscripten-built sqlite3.wasm (it's 833 KB and 4.8 ms scan vs 5.8 ms validate e2e). I agree that's too much overhead. If I try using the VisitOperator API and running bodies in parallel (same shape as validation) brings it to 0.57 ms (e2e). and there are free bail-outs: skip the walk when the module has no defined funcref tables (10 of my 18 benchmark modules), when every table is already conservatively marked (imported/exported), and early-exit once all tables are marked. lmk if this is agreeable and I can make the change (here or in a separate PR slice).
the target is Pulley on iOS/watchOS/tvOS/visionOS, where the App Store rules out JIT so everything is interpreted, and each retained check in the call_indirect sequence is one-plus extra interpreter dispatch per call rather than a folded native instruction. That also reframes two of your points: the null check is nearly free on native because the fault is the check, but Pulley (and any signals_based_traps(false) config) emits it explicitly; and uniform-signature tables are rare in big C/C++ modules (sqlite3 has ~25 signatures across ~620 elems) but common in the single-language modules this targets. 5 of the 8 table-bearing modules in my test corpus (both AssemblyScript builds, both Porffor builds, plus a Rust library called xmrsplayer) have exactly one signature across the whole table. Porffor funnels every indirect call through one giant calling-convention signature by design. The common shape across all 8: exactly one defined funcref table, one active element segment, and zero table mutation anywhere in the code section. btw, I looked at Porffor nd AssemblyScript based on the questions/recommendations in the chat. |
|
Separate PRs are fine, yeah, and is something we try to do where possible. For testing I'm basically asking you to write If you're curious, the main driver for me to request separate PRs is that (a) it makes it much easier to review and (b) any mistake here is likely a CVE-in-waiting for Wasmtime. Correctness here is paramount and is much easier to review in smaller chunks to ensure that all corner cases are well-tested, everything's given proper thought, etc. |
no worries! I'll reframe this PR as the first slice and stack the next steps as drafts.
we're in total alignment, protecting against malicious/pathological programs matters in my interpreter-only deployment scenario as well. I did another adversarial review with the CVE framing, and it found a real soundness hole that my fuzzing before didn't find. finalize_table_init folds element segments into the precomputed image only up to the first non-foldable segment (dynamic offset or expressions-form); the rest are applied at instantiation. The elision predicates treat the image as the table's complete contents, so a deferred segment installing a wrong-signature function dodgess an elided sig check (type confusion), and a deferred null write dodges an elided null check. (I'd say "derp", but it's actually a pretty complex scenario of interactions so I don't feel too embarassed.) The rebuilt PR stack fixes it (once I push) — static_funcref_image refuses tables with deferred segments — with regression coverage at three levels (wast modules that would miscall/crash on the current #13445 branch, environ unit tests, and the guard living inside the helper so no caller can forget it). |
4ac698c to
a5d6180
Compare
Add a serialized `mutated_tables` set on `Module` populated during translation: imported and exported tables are pre-marked (the embedder, or another module importing the export, can write them through the public API), then one decode pass over the code section marks the destination of every table-writing instruction — table.set, table.fill, table.grow, table.copy (destination only), table.init, and the shared-everything-threads atomic table writes. No module containing the atomic forms can currently be compiled (the cranelift translator rejects that proposal), but the scan decodes every operator, so matching them now keeps the analysis from going silently stale if the proposal is implemented. Active element segments are initial state, not mutation; elem.drop writes no table. The pass is skipped when a module has no tables (10 of the 18 modules in the benchmark corpus that motivated this) or when every table is pre-marked (e.g. Emscripten output, which exports its function table). Bodies are scanned with the `VisitOperator` API; with the new `parallel` feature (wired into wasmtime's `parallel-compilation`) they run on the rayon pool, and the serial fallback stops early once every table is marked. Forcing the walk over an 833 KiB Emscripten sqlite3 module measures ~4 ms serial / ~0.6 ms parallel; in a default build that module takes the skip path instead. Consumers arrive separately, starting with a constant table bound for tables that can never grow. `Module::table_is_immutable` documents the contract they may rely on; because they use it for correctness, a false positive here is a soundness bug, not a missed optimization. Bodies are unvalidated when the scan runs, so out-of-range destination indices are dropped rather than sizing any per-table structure from unvalidated input (the module is rejected later by real validation).
`make_table_base_bound` already treats `min == max` tables as static. Extend that to tables translation proved immutable: their size is `min` forever, so the per-call `current_elements` load disappears and the bounds check compares against a constant. Imported and exported tables never qualify through the new path (the host can grow them); they are pre-marked mutated by the analysis, so only the `min == max` rule can apply to them, as before. The new disas test pins the shape for a `min < max` table that nothing grows. The three regenerated goldens show the same folding kicking in for table-init startup functions: their bound loads and spectre guards constant-fold away since the tables they initialize can't have been resized before startup runs.
…tables Two wast files pinning runtime behavior for the table shapes the immutable-table optimizations care about, on both sides of the line. Immutable side: mixed-signature table with a null hole, uniform table fully initialized, uniform table with a hole, element segments covering only a prefix, a populated min<max table nothing grows (indices between the covered prefix and min hit null; indices in [min, max) are out of bounds), an empty declared-growable table, and tables whose element segments are only partially applicable as a static image (an expressions-form segment defers to instantiation) — a deferred entry with a different signature must still trip the signature check, and a deferred null write must still trip the null check, even though neither is visible in the compile-time image. Mutated side (where the optimizations must not fire): table.grow then dispatch into the grown region succeeds and the bound tracks the grown size; table.set of a null entry then dispatch traps; and an exported table written by a second module through its import — the value change is observed at a constant-index call site, and wrong-signature and null writes still trap. These are pure behavior tests: they hold on any correct configuration, with or without the optimizations, on every engine config the wast runner covers.
a5d6180 to
4a46653
Compare
TL;DR
Adds a
ModuleTranslation::tables_mutatedbit (set during translation bytable.set/table.fill/table.copy-dest /table.grow/table.initopcodes, or any passive elem segment that could land at runtime, or any leftover-segment shape that crosses a runtime resize) and uses it to elide six redundant runtime checks oncall_indirectagainst provably-immutable funcref tables:call FVMFuncRef *directly at instantiation)Each elision is gated on a stronger predicate than the previous; passive-segment + leftover-segment soundness corners are covered by integration tests.
Why
Each elision saves a small per-
call_indirectcost, but the combined predicate (is_eagerly_initialized_funcref_table) is what lets the downstream Pulley opcode-fusion stack (a follow-up PR) collapse the dispatch tail. Real-world graphql-js validation pipelines compile to ~98 call_indirect sites all dispatching through a single immutable funcref table — every site qualifies. The original impetus for this optimization of the Rust xmrsplayer compiled to wasm, needs later PRs in this chain to see uplift. The Porffor and AssemblyScript concerns I was asked to also solve for don't get uplift from this PR, their uplift comes at the end of the PR stack.Soundness
Three corners had to be tightened during development:
elem.initagainst a slot the predicate said was immutable would slip through).vmctxfrom the precomputedVMFuncRef, not the caller's.1, produced bytable.fill(null)on a tagged table; excluded by the immutability half of the predicate).tests/all/leftover_elem_segment_soundness.rs+ 4 disas filetests cover the soundness shapes.crates/environ/tests/table_mutability.rshas 12 cases for the predicate itself.Learnings / Caveats
An attempt at fully eliding the lazy-init brif (commits 1-8: egraph-folds it to
trapz) showed ~14 % branch mis-prediction increase on iPhone 12 E-core profiler across 3+ runs without a wallclock improvement. I kept the form from commits 1-7 (brif retained, mask + tagged-pointer test); commitdisable the commits 1-8 brif elision based on PMU evidencedocuments this. I'm looking beneath pure wallclock time into lower-level CPU counters to make sure we aren't silently layering things in a way that surprises us later.Tests
crates/environ/tests/table_mutability.rsintegrationtests/all/leftover_elem_segment_soundness.rsStacks under a follow-up PR for Pulley opcode fusion at the call_indirect lazy-init site.