diff --git a/doc/rfc/index.md b/doc/rfc/index.md index d3319e66..9e64d486 100644 --- a/doc/rfc/index.md +++ b/doc/rfc/index.md @@ -19,6 +19,7 @@ Design documents and technical proposals, grouped by scope. Shared/cross-cutting - [Extension Contract](submitqueue/extension-contract.md) - When extensions take orchestrator identity (request/batch) and resolve granular content themselves vs. take controller-resolved data; revises the BuildRunner base/head contract - [Gateway Status and List APIs](submitqueue/status-list-api.md) - Gateway-owned request context, materialized current status, sqid or change-URI status lookup, and queue admission listing - [Speculation](submitqueue/speculation.md) - Why SubmitQueue speculates, the path/tree model, and the two pluggable seams: speculation-tree enumeration and path selection +- [Outcome Predictor](submitqueue/outcome-predictor.md) - How likely a batch is to reach Succeeded: the scorer prices the change, the predictor revises that price from the path set and batch state (`pathPassed`, `pathFailed`, `merging`, `cancelling`) - [Best-First Speculation Path Generation](submitqueue/speculation-generator-best-first.md) - The default Generator: per-head lazy streams of flip subsets merged best-first across heads, log-probability ranking, and the strict snapshot contract - [Modular Queue Wiring](submitqueue/modular-queue-wiring.md) - Declare-don't-assemble engine (`pipeline.Construct`) that unifies topic registry, controller registration, DLQ pairing, and lifecycle ordering into one typed call; services self-declare via Deps struct + Stages slice, hosts own per-queue profiles and transport diff --git a/doc/rfc/submitqueue/outcome-predictor.md b/doc/rfc/submitqueue/outcome-predictor.md new file mode 100644 index 00000000..d4402a10 --- /dev/null +++ b/doc/rfc/submitqueue/outcome-predictor.md @@ -0,0 +1,109 @@ +# Outcome Predictor + +How likely a batch is to reach Succeeded, given the scorer's price for the change plus what this speculate run has already observed. + +See [speculation.md](speculation.md) for batches, paths, heads, and the Speculator. This document specifies the dependency probability the default Generator uses to rank paths. + +## The idea + +The **scorer** prices the change from content signals such as its size. That number does not move after the batch is admitted. + +The **predictor** prices the situation. It starts from the scorer's price and revises it with facts the speculate run already holds: a path *passed*, a path *failed*, the batch is *merging*, the batch is *cancelling*. + +`bestfirst` ranks a path by the probability that every unresolved dependency assumption holds. For each dependency, the predictor returns the probability of *succeeds* and `bestfirst` uses its complement for *fails*. Equal scorer prices can therefore produce different rankings for later heads once one dependency has a *passed* build. + +They are two contracts because they answer different questions. Putting path-set evidence on `Score` was tried: every content scorer took a parameter it discarded. + +**Default is a no-op.** Every factor starts at `1`, so the predictor returns the scorer's price until someone sets a factor. + +## What a factor is + +A factor revises the scorer's price. It is not itself a probability: `10` does not mean `0.10`, and `0.3` does not mean the batch is 30% likely to succeed. + +| Value | Meaning | +| --- | --- | +| `1` | Leave the scorer's price alone (the default if the key is omitted) | +| greater than `1` | More likely to reach Succeeded | +| between `0` and `1` (exclusive) | Less likely to reach Succeeded | + +Config accepts any finite value greater than `0`; there is no finite upper cap. + +The unconfigured scorer prices every batch at `0.5`. From that price, one factor `f` produces: + +| Factor | Price | +| --- | --- | +| `1` | 0.50 | +| `10` | ~0.91 | +| `12` | ~0.92 | +| `0.3` | ~0.23 | +| `0.25` | 0.20 | + +A scorer price of `0.6` with `pathPassed: 10` becomes about `0.94`. `merging: 12` on top of that becomes about `0.995`. + +`pathFailed: 0.3` from `0.5` becomes about `0.23`. It applies at most once because the path set has one current entry for the all-*succeeds* path; retry attempts replace that entry rather than adding evidence. `0` is rejected because it would pin matching batches at probability 0. + +When a factor revises the price, the arithmetic multiplies odds (`p / (1-p)`) and converts back, so the result stays in `(0, 1)` and the same factor means the same thing at any scorer price. Neutral factors return the scorer's price unchanged, including `0` or `1`. Adding to the probability provides neither property. + +YAML: + +```yaml +predictor: + type: evidence + factors: + pathPassed: 10 + pathFailed: 0.3 + merging: 12 + cancelling: 0.1 +``` + +The example values above are guesses, for reading the tables. The shipped default is to omit `factors` (every factor `1`). + +Profiles may set factors under `defaults.predictor` and revise them per queue. An omitted key keeps the inherited value — from defaults, or `1` when neither side named it. A queue `predictor` block overlays only the keys it names; it does not replace the whole map. An omitted `predictor` block on a queue inherits the defaults entirely, so every factor stays `1` until someone sets one. + +## Evidence + +| YAML key | When it applies | Typical direction | +| --- | --- | --- | +| `pathPassed` | Once, if a path that assumes every dependency *succeeds* has *passed* | Up | +| `pathFailed` | Once, if the path that assumes every dependency *succeeds* has *failed* | Down | +| `merging` | While the batch is *merging* | Up | +| `cancelling` | While the batch is *cancelling* | Down | + +`bestfirst` already treats a terminal batch as a fact (*Succeeded*, *Failed*, *Cancelled*). The predictor is not asked. *Merging* is not terminal: a merge can still fail, so how much it is worth stays a price. + +### Only the *succeeds* path counts + +The batch being priced is itself a head, so the run may have built it more than once under different assumptions about *its* dependencies. Only one of those builds is evidence. + +Take `C` depending on `B`, and `B` depending on `A`. Ranking `C`'s candidates needs the probability that `B` reaches Succeeded, so the Generator calls `Predict` with `B` and `B`'s path set. That set can hold two finished builds: + +| `B`'s path | What was compiled | +| --- | --- | +| `B` with `A` *succeeds* | `B` on top of `A`'s changes | +| `B` with `A` *fails* | `B` without them | + +`B` merges after `A` does, so the first build is a build of the code that will actually land: if it *passed*, `B` is likely to merge, and `pathPassed` applies. + +The second is a different set of changes. `B` may call something `A` introduces and fail to compile on its own — a *failed* result that says nothing about `B` merging in the normal case. Counting it would push `B` down the ranking over a build it was never going to need, while a green build of the real combination sits in the same set. + +So `pathPassed` and `pathFailed` both look only at paths that assume every dependency *succeeds*. Results on any other path are skipped. This is a filter on which results are evidence, not a check on whether an assumption came true — nothing here revisits that. + +## Rejected alternatives + +Design choices a reader might suggest after the sections above. Each names the alternative, why it fails here, and what this RFC does instead. + +### Fold path evidence into `Score` + +Give `Score` the speculate run's path sets so one call returns a situation-aware price. We tried it: content scorers took the parameter and discarded it; the composite forwarded evidence it never read. **Instead:** keep `Score` for the change; add `Predictor` for the situation (see [The idea](#the-idea)). + +### One model for content and evidence + +Train or tune a single estimate over diff shape and build outcomes together. Content signals and situation signals change at different rates, need different amounts of data, and would force every queue onto the same content scorer. **Instead:** scorer stays per-queue; evidence factors layer on in YAML. + +### Treat *merging* and *cancelling* as settled in the Generator + +Rank a *merging* batch like Succeeded and a *cancelling* batch like Cancelled. We tried and reverted: a merge can still fail, so the rank was wrong once outcomes diverged. **Instead:** only terminal states short-circuit in the Generator; *merging* and *cancelling* are predictor factors (see [Evidence](#evidence)). + +### Let the predictor read the path-set store + +`Predict` loads path sets from storage on each call — smaller API, fewer parameters. Each call can see a different snapshot mid-run (stale or split-brain relative to the rank the Generator is building). **Instead:** the speculate run reads all path sets as one snapshot and passes the matching set for each dependency. diff --git a/doc/rfc/submitqueue/speculation-generator-best-first.md b/doc/rfc/submitqueue/speculation-generator-best-first.md index 458e5e00..c0359d17 100644 --- a/doc/rfc/submitqueue/speculation-generator-best-first.md +++ b/doc/rfc/submitqueue/speculation-generator-best-first.md @@ -36,7 +36,7 @@ D.Dependencies = [] C contains A because C directly conflicts with A, not because B depends on A. The generator uses these stored direct dependencies and does not compute a transitive closure. -The scorer estimates: +The predictor estimates: | Dependency | Success | Failure | | --- | ---: | ---: | @@ -47,13 +47,13 @@ The batch being built is written before its assumptions. For example, `C [A succ ## The snapshot is a caller precondition -`Generate` receives the queue's live batches as a snapshot and takes it as given. A well-formed snapshot carries unique, non-empty batch IDs, includes every batch a head's direct dependencies reference, and gives no head an empty, duplicate, or self dependency. Those are preconditions the caller owns, established where the snapshot is assembled. The generator does not re-check them: it is on the hot path of every run, the checks it could make are the ones an assembled-correctly snapshot can never fail, and paying for them here only spreads the same contract across two places. A malformed snapshot yields undefined candidates rather than an error. +`Generate` receives the queue's live batches and path sets as one snapshot and takes it as given. Path sets are what each batch's builds have done so far — at most one per head, none for a batch nothing has speculated on. The speculate controller assembles both halves once per run and never re-reads mid-run; see [outcome-predictor.md](outcome-predictor.md) for why the predictor does not load them itself. A well-formed snapshot carries unique, non-empty batch IDs, includes every batch a head's direct dependencies reference, and gives no head an empty, duplicate, or self dependency. Those are preconditions the caller owns, established where the snapshot is assembled. The generator does not re-check them: it is on the hot path of every run, the checks it could make are the ones an assembled-correctly snapshot can never fail, and paying for them here only spreads the same contract across two places. A malformed snapshot yields undefined candidates rather than an error. -A dependency the generator cannot price is the one bad input it absorbs, because it arrives from the injected scorer rather than from the caller and there is no earlier point that could catch it. Three cases take the same 0.95 default: a score outside `[0, 1]` or `NaN`, a scorer call that returned an error, and a dependency the snapshot never carried. The default is optimistic on purpose, so a dependency nobody could estimate keeps its head's preferred path near the front instead of burying it or failing the whole run on one number — and failing the whole run is the real hazard, because `Generate` seeds the heap for every head at once, so one unpriceable dependency would otherwise cost the queue every candidate it had. A batch absent from the snapshot is never passed to the scorer at all: it would resolve to a zero-valued batch belonging to no queue, so scoring it would price some other batch entirely or fail on the empty queue name. Any deliberate defaulting still belongs to the scorer implementation, which knows what information it does and does not have; this is only the floor under it. +A dependency the generator cannot price is the one bad input it absorbs, because it arrives from the injected predictor rather than from the caller and there is no earlier point that could catch it. Three cases take the same 0.95 default: a probability outside `[0, 1]` or `NaN`, a predictor call that returned an error, and a dependency the snapshot never carried. The default is optimistic on purpose, so a dependency nobody could estimate keeps its head's preferred path near the front instead of burying it or failing the whole run on one number — and failing the whole run is the real hazard, because `Generate` seeds the heap for every head at once, so one unpriceable dependency would otherwise cost the queue every candidate it had. A batch absent from the snapshot is never passed to the predictor at all: it would resolve to a zero-valued batch belonging to no queue, so predicting it would price some other batch entirely or fail on the empty queue name. Any deliberate defaulting still belongs to the predictor implementation, which knows what information it does and does not have; this is only the floor under it. ## Step 1: `Generate` prepares each head -`Generate` scores each unique unresolved dependency once, however many heads wait on it. If A appears in both B's and C's dependency lists, the scorer is still called for A only once. +`Generate` predicts each unique unresolved dependency once, however many heads wait on it. If A appears in both B's and C's dependency lists, the predictor is still called for A only once. For each unresolved direct dependency, `Generate` records: @@ -420,7 +420,7 @@ A and D tie at 1.0, so batch ID puts A first. Other exact ties prefer fewer flip `Generate` must eagerly: -- score every unique unresolved direct dependency needed by an eligible head, substituting the default for any score that is not a probability; +- predict every unique unresolved direct dependency needed by an eligible head, substituting the default for an unusable probability; - choose each unresolved dependency's preferred assumption and calculate its `flipCost`; and - total the best score for every head. @@ -447,9 +447,9 @@ The ordering stays the same. `CandidatePath.RankingScore` contains this logarith - `Succeeded` fixes an assumption to succeeds. - `Failed` or `Cancelled` fixes an assumption to fails. - `Cancelling` remains undecided because cancellation may lose a race with completion. -- `Merging` also remains undecided, because a merge can fail. It is tempting to treat it as committed to landing and skip the scorer call, but that puts a state-specific policy inside the search: whether a path betting against a merging batch is worth funding is a question of price, and price belongs to the scorer. The allocator draws the same line — "no batch state enters this decision" — and the generator holds it too. Nothing is lost by staying open: a single passed path still waits for the merge result, while passed paths covering every outcome let the controller bypass the dependency (see [speculation.md](speculation.md)). Funding the unlikely side spends budget, which is the allocator's to ration. +- `Merging` also remains undecided, because a merge can fail. It is tempting to treat it as committed to landing and skip prediction, but that would turn an uncertain state into a fact inside the search. How much *merging* changes the probability belongs to the predictor; the Generator only consumes that price, and the Allocator still does not interpret batch state. Nothing is lost by staying open: a single passed path still waits for the merge result, while passed paths covering every outcome let the controller bypass the dependency (see [speculation.md](speculation.md)). Funding the unlikely side spends budget, which is the Allocator's to ration. - A fixed assumption stays in the returned path but contributes probability 1 and has no flip. -- A shared dependency is scored once per run. +- A shared dependency is predicted once per run. ## Why the algorithm works diff --git a/doc/rfc/submitqueue/speculation.md b/doc/rfc/submitqueue/speculation.md index 1887cff7..d026eb71 100644 --- a/doc/rfc/submitqueue/speculation.md +++ b/doc/rfc/submitqueue/speculation.md @@ -105,9 +105,9 @@ The one extension. It decides *which paths to build and which running ones to ca ### The default Speculator -The default Speculator is composed from two swappable interfaces — a **Generator** and an **Allocator** — so scoring and preemption policy can vary independently. They are composition points inside the default implementation, not controller-facing extensions: the controller depends only on the Speculator contract, and an alternate Speculator need not use or expose this split. The default opens the Generator's candidate stream over the batches, then hands that stream and the path sets to the Allocator. +The default Speculator is composed from two swappable interfaces — a **Generator** and an **Allocator** — so ranking and preemption policy can vary independently. They are composition points inside the default implementation, not controller-facing extensions: the controller depends only on the Speculator contract, and an alternate Speculator need not use or expose this split. The default opens the Generator's candidate stream over the batches, then hands that stream and the path sets to the Allocator. -- **Generator** — yields the queue's candidate paths as one iterator across heads in `BatchStateSpeculating`. *Contract:* every candidate has a Speculating head and is coherent; none repeats or contradicts a resolved fact. Ranking is implementation-defined — the Generator may compute it directly, call an injected scorer extension, or use other injected data — and the score it carries is meaningful only within the run. The Allocator consumes the iterator in the order the Generator yields it and does not interpret the score. *Default:* `bestfirst` ranks best-first by the probability that a path's assumptions all hold. +- **Generator** — yields the queue's candidate paths as one iterator across heads in `BatchStateSpeculating`. *Contract:* every candidate has a Speculating head and is coherent; none repeats or contradicts a resolved fact. Ranking is implementation-defined — the Generator may compute it directly, call an injected pricing extension, or use other injected data — and the score it carries is meaningful only within the run. The Allocator consumes the iterator in the order the Generator yields it and does not interpret the score. *Default:* `bestfirst` asks the predictor for each unresolved dependency's probability of reaching Succeeded, then ranks paths by the probability that all their assumptions hold. - **Allocator** — spends the build budget (the queue's cap on concurrent builds) over the iterator. *Contract:* it pulls in order until the budget fills and matches candidates to existing paths by ID, so a pending or building path keeps the slot it already holds rather than starting a second attempt, and a candidate whose path is already terminal in the path sets is skipped rather than rebuilt; pending dispatches are replayed by the controller as described above. Pending, building, and cancelling paths charge the budget (a cancelling build holds CI until terminal), while terminal ones charge none. Cancellation is best-effort, so the Allocator does not spend capacity it merely expects a cancel to release and risk exceeding the hard CI cap. *Default:* the sticky policy fills only free slots and leaves in-flight builds running; a preempting policy cancels in-flight paths below the funded set. Budget is the only rationing lever — there is no ranking-score floor. A build cancelled to make room still charges budget until its cancel reaches terminal and publishes dirty, so the next run funds the released slot — the queue converges over successive ticks rather than oversubscribing in a single pass. ### Extension APIs @@ -118,4 +118,8 @@ Signatures live in code and are not copied here, so they cannot drift. This sect **Speculator** — [`submitqueue/extension/speculation/speculator`](../../../submitqueue/extension/speculation/speculator/README.md). `Speculate` takes one queue snapshot (the batches and their path sets) and returns the build and cancel actions it proposes; a path it wants left alone has no entry. Actions must target Speculating heads. Verdicts stay controller-owned, so there is no merge or fail action. -**Generator and Allocator** — [`generator`](../../../submitqueue/extension/speculation/generator/README.md) and [`allocator`](../../../submitqueue/extension/speculation/allocator/README.md), the two composition points inside the default Speculator. The Generator opens a pull-based stream of candidate paths over the batches; the Allocator spends the build budget over that stream, reconciling it against the path sets. Both abort on a cancelled context. +**Scorer** — [`submitqueue/extension/speculation/scorer`](../../../submitqueue/extension/speculation/scorer/README.md). Prices a batch's change from content signals. The default pipeline does not rank on it directly; the queue's predictor is built over it. + +**Predictor** — [`submitqueue/extension/speculation/predictor`](../../../submitqueue/extension/speculation/predictor/README.md). Revises the scorer's price with path-set evidence and batch state. The default `bestfirst` generator is built over it. See [outcome-predictor.md](outcome-predictor.md). + +**Generator and Allocator** — [`generator`](../../../submitqueue/extension/speculation/generator/README.md) and [`allocator`](../../../submitqueue/extension/speculation/allocator/README.md), the two composition points inside the default Speculator. The Generator opens a pull-based stream of candidate paths over the batches and path sets; the Allocator spends the build budget over that stream, reconciling it against the path sets. Both abort on a cancelled context.