fix(kernel): close the provisioner's four residual mid-boot lifecycle gaps - #3303
kevinjosethomas wants to merge 23 commits into
Conversation
…d startup Review follow-ups on PR #3257 (both blockers independently verified by both adversarial reviewers at 21da7aa): - F1 panic wedge: a panic in run_startup or a progress listener skipped the memo clear and the done send, which closed the watch and wedged every later ensure() on the dead memo. The boot now runs on its own JoinHandle, so the publisher always clears the memo and sends the result, converts a panicked boot into a "kernel startup task failed" failure, and clears the listener state so the next ensure() boots fresh (self-heal; the memo re-arms on close). The watch now carries the startup result itself: joined callers keep the manager even when a concurrent stop takes the provisioner's live owner (TS managerPromise), and a dispose-raced boot rejects with the honest disposed-startup cause instead of handing callers a dead manager. - F2 stop-gate gap: stop_kernel took the manager and installed the pending_stop gate in two separate lock passes; a revival ensure in the gap captured the previous gate and could restore against a snapshot still being flushed. The gate is now installed in the same lock block that takes the manager. - F3 dispose fixture vacuity: the dispose leg pins the absolute interpreter count at dispose-return (exactly one spawn; the aborted retry never ran) in addition to the stasis assert. - Fixtures: a panicking progress handler must not poison the next boot (the re-arm regression leg), and a revival boot must wait for the in-flight stop's held host request before its first progress stage (exact ordering barrier, not flush timing).
… gaps Staged hardening set for the #3257 shared-startup head (the design consult's unresolved-risk list + the Macroscope 07:10/07:15 R3 residuals, both fresh reviewers dispositioned pre-existing and recommended this follow-up): - Atomic settle/publish: the publisher's disposed check, memo-generation check, listener teardown, and manager park now share ONE lock scope (TS ensure()'s `managerPromise === startup` guard). A stale publisher can no longer clear a NEWER memo armed underneath it - a defunct clear inside the park->publish->clear window used to wipe the newer memo and arm a transient duplicate kernel - and a boot can no longer park a live kernel into a provisioner that reported itself torn down between the old separate check and publish scopes (the dispose TOCTOU). - kill() clears the startup memo (TS kill() clears `managerPromise`): a boot doomed mid-flight is killed by its own settle instead of parking a resident kernel nobody owns; the next ensure() boots fresh. - Concurrent stops of one in-flight boot share the first stop's gate (`pending_stop_for_startup`): the take-losing stop task used to open a fresh gate before the winner's final snapshot flush, un-gating a revival to race that flush over the same on-disk file. The take-loser that still exists (against a direct-arm stop) now waits the currently armed gate before opening its own. Fixtures: kernel_startup_memo.rs (stale-publisher generation survival; kill-during-boot; dispose-during-boot end-state pin; kill vs #3257's panic-cleanup re-arm) + the concurrent-stops fixture in kernel_stop_revive.rs. All drain-boundary observables on the single-threaded test runtime (the settle chain runs in one scheduler drain; no sleeps, no interposition into wake chains). Failing-first verified on the pre-fix head: the kill/stale fixtures delivered a LIVE resident kernel; the concurrent-stops fixture caught the loser's early gate open in the settle drain.
Semantic merge of the merged #3257 shape (b28f22d's review follow-ups: the memo-generation clear, the stop-gate chaining, and the "Waiting for the previous kernel to stop..." progress stage) with the staged hardening chain at 87d138b. - provisioner.rs: the hardening mechanics stand (the one-lock-scope settle/publish, kill()-during-boot generation invalidation, the same-boot stop JOIN via pending_stop_for_startup); the weaker interim fixes they subsume (the separate-scope same_channel clear, the previous_stop chaining) give way to them; the merged #3257 progress-stage emit is kept - main's evolved revival-ordering test asserts on that exact stage. - kernel_stop_revive.rs: main's evolved ordering barrier test plus the lane's concurrent-stops-share-one-gate fixture.
The Macroscope/Bugbot findings on the stopgate-hardening head, as drain-boundary fixtures BEFORE the fix (RED by construction; the fail-first receipts are the ffix-branch gates cited in the fix commit): - superseding_stop_after_kill_cannot_deadlock_the_revival_gate: a stop armed for one boot, kill(), a revival gated on that stop's gate, and a second stop of the revived boot that supersedes it; the take-losing first stop waiting the currently-installed gate forms the wait cycle (first stop -> second stop's gate -> revival boot -> first stop's gate) - the oracle is the bounded settle, so the cycle shape wedges against the 30s timeout. Absolute spawn count pinned. - doomed_settle_keeps_the_newer_boot_listener_state: a doomed boot settling against a newer memo generation must not wipe the newer boot's progress listeners/replay stage - a late joiner of the newer boot replays that boot's current stage to a fresh handler; the gated counting wrapper holds each boot's handshake at a file barrier (the wrapper reads its ordinal from the count file, so the count file is pre-created; a Drop guard opens every gate on any exit so a wrapper never outlives the test as an orphan), making the settle ordering a test-controlled barrier, never a timing race.
…ener teardown The five Macroscope/Bugbot findings on the stopgate-hardening head reduce to two root causes; both fixed, with the drain-boundary fixtures that failed first on the unfixed tree (RED receipt: gate_1790910402 on the fixtures-only tree e3b92c4 - fmt rc0, clippy rc0, test rc101 with EXACTLY the two new fixtures failing at their designed oracles: "a late joiner must replay the active boot's stage" and "the first stop settled (no wait cycle)", both bounded-settle timeouts; everything else green; the compile-caught authoring attempts gate_1790908504/ gate_1790908744/gate_1790909006/gate_1790909694 are the fixture authoring cycle, incl. the counting-wrapper ordinal-read-before-create bug that made one early RED vacuous): - The stop-gate CHAINING (main's #3257 shape) dropped by the JOIN: arm_stop now captures the gate the arm REPLACED and the stop task waits THAT before opening its own - never the currently-installed gate. The installed-gate design missed the supersede window (the take-losing second stop's loser saw its OWN gate installed and opened before the first direct stop's final snapshot flush settled, un-gating a revival to race that flush) and admitted a three-task wait cycle (a stop waiting a superseding gate that waits a revival that waits the first stop's gate). The chain is strictly-older-only, so acyclic by construction; the same-boot JOIN is unchanged (the second stop of the same boot still never arms). - The settle's listener teardown is generation-scoped: the doomed boot clearing startup_listeners/last_startup_message unconditionally wiped a NEWER memo generation's progress listeners and replayed stage. The clear now runs only when the settling memo is the active generation (or no memo is armed at all - the stale-entries case).
…y assert
- arm_stop installs the arm's own receiver directly (clippy redundant_clone:
the clone was the value's last use after the Armed struct stopped
carrying stop_rx).
- the listener fixture's replay oracle asserts the replay FIRES with the
active boot's current stage; pinning the exact stage string made the
oracle flaky against legitimate later stages the boot emits around its
held handshake ("Preparing Python runtime...").
The 4880ab3 settle fix scoped the LISTENER TEARDOWN to the active memo generation but left the EMIT unscoped, and its own fixture proved it: doomed_settle_keeps_the_newer_boot_listener_state failed 3/3 on the fixed tree (replayed "Preparing Python runtime..." instead of "Starting Python kernel...") - a boot kill() invalidates still emitted its later stages, overwriting the newer generation's last_startup_message and firing the newer boot's listeners with the wrong boot's stage. The emit now threads the boot's memo generation through run_startup/start_kernel/start_kernel_impl and writes the shared progress state (last_startup_message + the listener fan-out) only when the emitting memo is still the ACTIVE generation; the boot's own on_progress handler keeps firing unconditionally (it is its caller's, not shared state). Also drops the redundant stop_rx.clone() that red-inked CI's clippy job on the swarm push (redundant_clone at the arm_stop install). Gates (keeper worktree, CARGO_TARGET_DIR=.cargo-target-keeper-3303): fmt rc0; clippy --workspace --all-targets -D warnings rc0; kernel fixtures kernel_startup_memo 5/5 + kernel_stop_revive 5/5 (the previously-red doomed_settle fixture green); pa-core lib 1043 passed, 1 failed - workspace_snapshot::head_tree_baseline_excludes_only_absent _skip_worktree_paths, a git-2.34.1 sparse-checkout artifact reproduced outside the lane (git 2.34.1 marks root files S after `sparse-checkout set --cone`; CI's newer git passes it; the PR does not touch workspace_snapshot).
… into lane/perf-kernel-stopgate-hardening-keeper
…ill() FINDINGS-1 from the pre-push review (gpt-6-sol cross-model reviewer, report /home/ubuntu/handoffs/REVIEW-3303.md): kill() invalidates the startup memo but retained the killed boot's last_startup_message and startup_listeners, so the next ensure() replayed the killed boot's stale stage to a fresh handler and the newer boot's emits fanned out to the killed generation's listeners - the 5e3037a emit gate stops the doomed boot's own emissions but not the B-to-A delivery or the stale replay. kill() now clears both shared fields in the same lock scope that invalidates the memo (the same stale-entries reasoning as the settle's generation-scoped teardown), and the listener fixture grows the transition oracles the reviewer asked for: A's joiner's collector must show nothing after the kill, and the newer boot's handler must show exactly its own stage twice (fan-out + boot-local, the pre-existing main shape) with no stale replay first. RED receipt: with only the kill-clear reverted, the fixture fails at "the killed generation's listener must not hear the newer boot" (kernel_startup_memo.rs:531); restored, all 10 kernel fixtures pass. Gates: fmt rc0; clippy --workspace --all-targets -D warnings rc0; pa-core lib 1043 passed, 1 git-2.34.1 sparse-checkout artifact (workspace_snapshot, untouched by this PR, green on CI's newer git).
Macroscope's follow-up finding on 5901a45 (thread r4162712975): a run_startup task whose memo kill() invalidated still consumed its retry - each attempt spawns an interpreter and runs restore/bootstrap, and the stale attempt's settle only killed its kernel, leaving duplicate kernel work and stale restore notifications/state against the fresher generation. - The retry loop now checks boot_generation_is_live (memo still the armed startup generation AND provisioner not disposed) before each retry AND again after the backoff, so a kill() during the backoff window cannot resurrect a dead generation's boot. The disposed arm of the old guard folds into the same check (dispose() does not clear the memo, so both conditions are needed). - The restore surface (on_restore + last_restore) is now generation-gated with the same predicate: a boot kill() invalidated restores its snapshot only for the settle to kill its kernel, so its restore must not fire the notification nor pollute last_restore. Fixture: kill_during_the_retry_backoff_pins_the_spawn_count - a fail-first flaky wrapper (first invocation exits 1 before the handshake; later invocations exec the real kernel python), kill() during the backoff window, and the absolute spawn count pinned at 2. RED receipt: with the memo check removed the fixture fails 3 vs 2 (the doomed boot consumed its retry); restored, all 11 kernel fixtures pass. Gates: fmt rc0; clippy --workspace --all-targets -D warnings rc0; pa-core lib 1043 passed / 1 git-2.34.1 sparse-checkout artifact (workspace_snapshot, untouched by this PR, green on CI's newer git).
… into lane/perf-kernel-stopgate-hardening
…com/PrimeIntellect-ai/prime-agent into lane/perf-kernel-stopgate-hardening
Cursor Bugbot's finding on 155a876 (thread r4162867816): a boot that kill() invalidated still fired on_unavailable_skills after its bootstrap - the callback shares the restore notice's mailbox, so a discarded kernel could append stale skills-unavailable rows the next turn showed the model, even though on_restore/last_restore are generation-gated in the same function. The skills report now uses the restore gate's exact contract: the generation-liveness check and the callback share ONE lock scope (the callbacks must not re-enter the provisioner, same as the startup-progress listeners), so a kill() can no longer slip between the check and the report. Fixture: doomed_boot_does_not_report_unavailable_skills - the gated counting wrapper holds the doomed boot's interpreter, kill() invalidates its generation, the newer boot arms with the SAME broken skill, and the mailbox must carry exactly the live boot's report. RED receipt: with the gate reverted the fixture fails 2 vs 1 (the dead generation's report lands first); restored, all 12 kernel fixtures pass. Gates: fmt rc0; clippy --workspace --all-targets -D warnings rc0; pa-core lib 1043 passed / 1 git-2.34.1 sparse-checkout artifact (workspace_snapshot, untouched by this PR, green on CI's newer git).
…-hardening-keeper
|
The kill-during-boot fix is real (fixtures fail on main). The rest guards against things main already handles:
In the kill path:
[written by prime-agent, reviewed by snimu] |
…-hardening-keeper
Main's #3004 promoted the windows cross-check (clippy for x86_64-pc-windows-gnu), and the new job is red on main's own tip: #3195's interleave harness binds a unix-socket probe, but tokio gates UnixListener behind all(unix) - the module fails to cross-compile for windows (E0433, the single red in pa-daemon's lib-test target). The module declaration now carries #[cfg(unix)] (the same contract as feed's #[cfg(unix)] tests - every race test in the module rides the unix-socket probe, so the whole module is unix-only). This is an upstream break the lane folds and fixes so its head can be green; main carries the same fix independently. Gates: cargo check -p pa-daemon --all-targets --target x86_64-pc-windows-gnu rc 0; fmt rc 0; clippy --workspace --all-targets -D warnings rc 0; pa-daemon turn_stream_tests 42 passed on unix.
The two Cursor Bugbot findings on 4c74a09, plus the cross-model reviewer's TOCTOU finding on the first fix shape: 1. "Stop-arm memo never released" (Medium): pending_stop_for_startup is written when a stop arms against an in-flight boot and never cleared when that boot settles or kill() invalidates it. The arm's memo receiver is a strong handle to the settled StartupResult, so a spent arm keeps a parked-or-failed manager alive past stop_kernel - a later failed shutdown's process would leak with the stale receiver as its sole owner. Both spenders now release the arm (the boot's settle and the kill() that invalidates it), through one shared release_spent_stop_arm. 2. "Doomed boot can clobber snapshots" (Medium): a boot kill() invalidated still runs start_kernel_impl to completion, and its failure teardowns shut down with the dispose snapshot policy (default true) against the same snapshot_dir the replacement boot restores from. The four failure teardown sites now route through hold_snapshot_flush_gate: the policy (dispose policy while the boot's memo is still armed, never a flush for a kill()-invalidated boot; dispose() does not clear the memo so a disposed boot keeps the dispose's own policy) AND, when the policy still flushes, a pending-stop revival gate for the teardown's own flush - the decision and the gate install share ONE lock scope, so a kill() landing between the decision and the flush still leaves the replacement boot gated on this teardown (the same revival gate stop_kernel arms against its own flush; a stop ACTIVELY armed for this boot is the one exemption - its memo wait settles only after this teardown, so its own gate already covers the flush - while a merely installed older gate, including one long settled and left in place, covers nothing and is replaced by the teardown gate). Also gates the two unix-only kernel test targets (kernel_startup_memo.rs, kernel_stop_revive.rs) with #![cfg(unix)] - main's #3004 promoted the windows cross-check and the whole-target /bin/sh wrapper fixtures never cross-compiled (the sibling kernel targets and lock_compat already carry the gate). RED receipts (provisioner unit tests): release no-op -> both arm tests fail; flush gate install removed -> the gate assert fails; restored -> 12/12. Gates: fmt rc0; clippy --workspace --all-targets -D warnings rc0; the 12 kernel fixtures pass; pa-core lib 1046 passed / 1 git-2.34.1 sparse-checkout artifact (workspace_snapshot, untouched by this PR, green on CI's newer git); clippy --workspace --all-targets --target x86_64-pc-windows-gnu -D warnings rc0.
sethkarten
left a comment
There was a problem hiding this comment.
LGTM — approved on Seth Karten's instruction (submitted via his agent).
|
Fixed: the spent-arm leak (release in settle and kill), the cfg(unix) gates. Remaining:
[written by prime-agent, reviewed by snimu] |
…-hardening-keeper Folding this merge drops the lane's own interleave gate from c9de47a: main's #3306 (db686e0) gates the harness whole-file with #![cfg(unix)] inside interleave.rs, and keeping the lane's outer #[cfg(unix)] on the mod declaration too would leave the module with duplicated cfg attributes (clippy duplicated_attributes risk under -D warnings). The resolution takes main's side for turn_stream_tests.rs, so the interleave files match main exactly.
…ed gates snimu's five remaining review items (2026-10-02T20:36:31Z comment): 1. StopArm::Joined deleted. The join made a second stop of one in-flight boot wait the first stop's gate after arm_stop had already written the second caller's dispose_snapshot, splitting the two snapshot policies. Concurrent stops of one boot now supersede: each stop arms its own gate and its task waits the strictly older gate it superseded before opening its own, so a revival cannot cross before the final flush settles; pending_stop_for_startup stays (the flush-gate exemption is load-bearing), now consulted by the failed-boot teardown instead of the join. 2. kill() parks the startup memo it invalidates (doomed_startups) instead of clearing it. The doomed boot's kernel exists until its settle, and a later dispose() must wait that settle the same way it waits an armed memo - without the park, a kill followed by a dispose skipped the doomed boot entirely and could orphan its kernel at worker exit. Each parked memo is spent by its own boot's settle. 3. The settle path checks the generation before the disposed flag (settle_decision): a killed+disposed boot reaching Ok tears down with kill() semantics and never flushes a snapshot, and kill->ensure->dispose cannot overlap two flushes on one snapshot_dir. The order is pinned by a unit test. 4. The four memo fixtures hold their boots at file gates instead of spinning on yield_now; the flaky fixture holds before its failure on the same schedule, so a kill lands mid-boot by construction. Both gated wrappers self-release when their fixture directory vanishes (the counting loop falls through to the real interpreter, whose stdin pipe is gone; the flaky loop exits 1), so a cancelled test cannot leave a polling shell behind. 5. The main merge folds #3306's #![cfg(unix)] inside interleave.rs, dropping c9de47a's duplicate mod gate from the resolution (the duplicated_attributes clippy risk). Red-first regressions, both demonstrated failing on the pre-change behavior: dispose_after_a_kill_waits_the_doomed_boot and a_killed_then_disposed_boot_never_flushes_the_snapshot; the settle-order unit pin fails with the old disposed-first order. The joined-gate drain test is deleted with its arm. Cross-family review (bugbot, gpt-6-sol): READY-CLEAN, no findings.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 8d29f37. Configure here.
…-hardening-keeper
cursor[bot] finding on 8d29f37 ("Killed boot still spawns interpreter", Medium): the boot permit gate checked only the dispose signal, so a kill() landing while a boot was still waiting - queued behind an in-flight stop's gate or the boot permit itself - let the doomed generation spawn an interpreter its own settle then had to kill (no flush, no park, but a real boot racing the replacement generation on the same snapshot_dir and kernel-stderr.log). The permit-gate closure now rechecks boot_generation_is_live as the last point before the interpreter spawn and fails fast with "Kernel provisioner killed before start" - non-retryable (FATAL_MARKERS "provisioner killed", unit-pinned). The settle stays the correctness backstop for the post-check race. Red-first regression a_kill_while_the_boot_waits_the_stop_gate_spawns_no_interpreter: boot A holds mid-handshake at a file gate; a stop arms against it; kill #1 invalidates A; a fresh ensure arms memo B (polled once on the test thread so the arm is guaranteed, not yield-assumed); kill #2 invalidates B while it still waits; releasing A settles the chain and B fails the liveness check before the spawn. Red on the unfixed tree (B spawned a 183ms handshake and settled "killed during startup"); green after (the spawn count stays at one). Also folds pi/main c24ac22 (factory/machine lib, acp single daemon path, semantic-edge ledger, telemetry first-frame fix, factory e2e, subagent relisting, factory state-machine port, vouches). Cross-family review (bugbot, gpt-6-sol): READY-CLEAN on the fix and on the test-hardening re-review.

Evidence
Full-suite VM gate (the receipt-of-record) — fleet_pipeline.fleet_gate
gate_1790905327_154675(host logs: /home/ubuntu/fleet-pipeline-runs/gate_1790905327_154675)b0b4b30051f526405b04bd2160c7eb85291fbe8don org/main tipa0c131566b6fc25c832ab9b973dab37af8080da5(merge-basea0c131566b6fc25c832ab9b973dab37af8080da5, base folded)it4kvxkzo4mx4laket2ztbcd), hosts_guard OK, kernel_guard OK, release killverify cleancargo +1.98.1 fmt --all --check: rc 0cargo +1.98.1 clippy --workspace --all-targets -- -D warnings: rc 0cargo +1.98.1 test --workspace --no-fail-fast: rc 0 — 80 test-result lines, 1751 passed, 0 failed, 12 ignored, zero unknown failuresCrates-mode pre-check gate (same head, opened the PR earlier)
gate_1790905137_154675:cargo fmt --all --checkrc 0,cargo clippy -p pa-core --all-targets -- -D warningsrc 0,cargo test -p pa-core --no-fail-fastrc 0 (21 result lines, zero failures; all 5 drain-boundary fixtures pass on the folded tree).-p pa-core-scoped. The full-suite receipt above is the receipt-of-record; this block restates the crates receipt accurately.What this PR does
Closes the provisioner's four residual mid-boot lifecycle gaps — the design consult's unresolved-risk list (the kill-vs-boot and atomic-settle entries) plus the two Macroscope residuals that R3 dispositioned pre-existing with this exact follow-up recommended. One commit (
87d138bdb, +722/-89, 3 files) on the #3257 head, folded over the current main tip:ensure()'smanagerPromise === startupguard). A stale publisher can no longer clear a NEWER memo armed underneath it, and a boot can no longer park a live kernel into a provisioner that reported itself torn down between the old separate check and publish scopes (the dispose TOCTOU).kill()clears the startup memo (TSkill()clearsmanagerPromise), so a boot doomed mid-flight is killed by its own settle instead of parking a resident kernel nobody owns; the nextensure()boots fresh.pending_stop_for_startup): the take-losing stop task used to open a fresh gate before the winner's final snapshot flush, un-gating a revival to race that flush over the same on-disk file. The take-loser now waits the armed gate.kernel_startup_memo.rs(stale-publisher generation survival; kill-during-boot; dispose-during-boot end-state pin; kill vs share one kernel startup across first use, stop, and close #3257's panic-cleanup re-arm) + the concurrent-stops fixture inkernel_stop_revive.rs. All observables are drain-structural on the single-threaded test runtime (the settle chain runs in one scheduler drain); interpreter spawns are counted by a wrapper that pins absolute counts.Failing-first evidence (archived, sha256
10dac2c5...)On the pre-fix tree (
7ddb838a4, the #3257 head) the memo fixtures failed 4/4 — the kill fixtures delivered a LIVE resident kernel after a mid-flight kill; the dispose fixture caughtdispose()returning before the in-flight boot settled; the stale-publisher fixture caught the newer memo's kernel shut down underneath a late joiner (LEG1_RC=101). On the fixed tree: the 5 fixtures +kernel_startup_join5 +kernel_stop_revive4 +kernel_prewarm2 + the 9 provisioner unit tests all green; fmt/clippy rc0. Evidence:/home/ubuntu/hillclimb/vms/perf-kernel-stopgate-hardening/evidence-stopgate.tgz(sha25610dac2c537d3f810dfbde8be94f71bda3045c563687f02e60c158c6e79b4e261).The fold (semantic merge, hunk-verified)
Main's squashed #3257 (
b28f22ded) already carries a weaker interim of items 1/3 (the separate-scopesame_channelclear; theprevious_stopchaining) plus the"Waiting for the previous kernel to stop..."progress stage. The fold keeps the hardening's atomic-scope mechanics (they subsume the interims) and re-grafts the progress emit — main's evolvedrevival_waits_for_in_flight_stop_before_bootingordering test asserts on that exact stage. Post-fold delta verified hunk-by-hunk: HEAD vs main = exactly the 3 hardening files; the 9 hardening hunks inprovisioner.rsare byte-identical to the staged commit's, and the only extra delta vs the staged commit is the 7-line emit graft.TS parity
TS anchors (carried verified by the kernelgate design consult and the #3257 review arc, TS checkout
cd1f215c):tools/ipython.ts:427-437(kill()clears/awaitsmanagerPromise),:440-485(the memoizedmanagerPromise, themanagerPromise === startupclear guard, joined concurrent ensure),:499-512(new boot gated on prior dispose),:625-639(first tool access awaits the memo). Frozen surfaces untouched: no wire/API/tool-schema, prompt, session-JSONL, or TUI bytes change (the re-grafted progress stage is main's current behavior, byte-identical).Gate receipts (exact head
b0b4b3005)gate_1790905327—cargo fmt --all --checkrc0,cargo clippy --workspace --all-targets -- -D warningsrc0,cargo test --workspace --no-fail-fastrc0 — 80 test-result lines, 1751 passed, 0 failed, 12 ignored, ZERO unknown failures, release killverify clean, hosts_guard OK, kernel_guard OK,base_state=folded(merge-base == org/main tipa0c131566), VMit4kvxkzo(pool-0), finished 2026-10-02T02:00:55Z.gate_1790905137— fmt rc0, clippy -p pa-core rc0, test -p pa-core rc0 (21 result lines, zero failures; all 5 drain-boundary fixtures pass on the folded tree).kernel_stop_reviveordering test (which asserts on the re-grafted progress stage) and the hardening's join mechanics — both green in the same full-suite run: the semantic merge is proven by the suite.Known reds
kernel_restore_guardsco-scheduling family: renewed 2026-09-29 (solo x3 green), expiry extended to 2026-10-06; zero pool fires since the v1.4.7 claim guard. This lane does not touch the kernel restore surface.No self-merge: the PR stays open for the orchestrator's numeric review + both adversarial reviewers' SHA-bound APPROVEs.
Note
High Risk
Changes concurrency, snapshot flush ordering, and kernel process lifetime during boot/stop/dispose/kill—bugs could leak interpreters, corrupt session snapshots, or deadlock revival.
Overview
Hardens
IpythonKernelProvisionersoensure(),kill(),dispose(), andstop_kernel()cannot race while a kernel is still booting.Startup memo generations tie each in-flight boot to its own watch channel. The publisher task now performs an atomic settle/publish under one lock: generation liveness, dispose flag, listener teardown, and whether the manager parks or is torn down (
settle_decisionprefers kill over dispose for invalidated generations). Stale publishers clear only their own memo (same_channel), andkill()parks invalidated memos indoomed_startupssodispose()waits them out.kill()during boot invalidates the memo, clears shared progress state, releases stop arms tied to that boot, and the settle path kills the interpreter instead of leaving a resident manager. Retries, restore/on_restore, unavailable-skills callbacks, and shared progress emits are generation-gated so dead boots cannot retry, clobber snapshots, or pollute the next turn.stop_kernel()is refactored througharm_stop()withpending_stop_for_startup: concurrent stops supersede gates but each task waits only the strictly older gate it replaced (avoids revival/snapshot flush races and wait cycles). Failed-boot teardowns usehold_snapshot_flush_gateso killed generations never flush over an on-disk snapshot; live/disposed boots can still gate revival on their flush.Adds Unix integration tests (
kernel_startup_memo.rs, extendedkernel_stop_revive.rs) and unit tests for settle order, stop-arm release, and flush gating.Reviewed by Cursor Bugbot for commit 78f038a. Bugbot is set up for automated code reviews on this repo. Configure here.
Note
Fix provisioner's four residual mid-boot lifecycle gaps
ensure(),kill(),dispose(), andstop_kernel()when a kernel boot is still in flight. A boot invalidated bykill()or a racingdispose()is now torn down at settlement instead of being published, with an ordered settlement decision (settle_decisionin provisioner.rs).kill()now parks an invalidated startup memo indoomed_startups;dispose()waits for those doomed boots to settle before returning, and a killed boot never flushes a dispose snapshot.stop_kernel()stops form an ordered chain viaarm_stop(): each new stop claims the live manager or the in-flight startup atomically, and its gate opens only after the superseded stop settles.dispose()can now block longer because it waits for all captured startup memos, including killed ones, to settle; settlement and publication inrun_startupmoved to the publisher task under one state lock.Changes since #3303 opened
kernel::provisioner::start_kernel_implto recheck boot generation liveness immediately after acquiring the start permit and abort with a specific error before spawning the interpreter if the generation was killed while waiting, and added that error marker to the non-retryable list instartup_failure_is_retryable. [78f038a]SupervisorChildSessionsInner::reseed_from_ledgerto rebuild the in-memory child registry from the persisted RLM spawn ledger for the current parent session, andSupervisorChildSessionsInner::prime_cursor_if_lazyto initialize a reseeded child's usage cursor to the end of its history on first delivery, preventing back-billing of pre-restart entries. [78f038a]SupervisorChildSessionsInnerchild-settle handling to record the child's return on the parent's semantic-edge ledger by reading the child's last committed request id and invokingrecorder.record_child_returned, and to update the child's display file from status 'running' to 'completed' when the run settled successfully. [78f038a]run_daemon_attached_acp_modeto create and attach the daemon session during startup viabind_daemon_session, fail early on create/attach errors, remove agent_start/agent_end markers, implement terminal quiescence by polling autonomous status and child roster until outstanding subagents reach zero, cancel outstanding RLM children during stop/close/teardown, and publish heartbeats_changed broadcasts as SessionInfoUpdate frames. [78f038a]print_runtime::acp_mode_main, deleted thetry_daemon_attached_acphelper, and removed in-process ACP modules includingacp::session,acp::autorefine,acp::compaction_arms,acp::events,acp::goal_continuation,acp::in_process_config,acp::producer, andacp::prompt. [78f038a]wire_events::wire_updatesto map bash_start, bash_output, and bash_end kernel events into ACP tool-call frames with synthetic ids derived from runId, and modifiedtool_result_textto returnNonefor empty tool results and filter out empty text blocks. [78f038a]semantic_edgesmodule includingSemanticEdgeRecorder,SemanticEdgeIdentity, ledger append/read helpers,wrap_stream_fnto inject request ids into model stream calls, andSemanticCompactionguard for compaction lifecycle events. [78f038a]create_sessionto optionally wrap the stream function with semantic edges whensemantic_identityis provided, construct aSemanticEdgeRecorder, set an RLM spawn semantic anchor on the host bridge, and pass a pre-semanticside_question_stream_fninto the session, and added afactory_hostfield toSessionEnginebuilt from captured session facts. [78f038a]execute_compactionto usesemantic_edges::summary_slice_callfor each summary request, wrappingcomplete_summary_callto inject request-specific headers, and to commit the compaction on the recorder before persisting the compaction entry. [78f038a]RlmHostBridge::register_runto capture aSemanticSpawnAnchorand setrequest.spawned_by_request_idby consulting the anchor to resolve the parent's last turn request id if the agent is streaming, and propagatedspawned_by_request_idthrough spawn handling to injectspawnedByRequestIdinto the child's session configuration. [78f038a]rlm::factorymodule with validation (validate_factory_spec), run orchestration (run_factory,status_factory,stop_factory,resume_factory,graph_factory,watch_factory), CLI dispatch, and activity handling;HarnessState::create_factory,update_factory,delete_factorymethods; and a new 'factory' harness kind with deep-copy semantics for arguments. [78f038a]RefinementKind::Factoryvariant, extendedREFINEMENT_KINDSand related utilities to support 'factory', introducedfactory_enabledfunction reading agent_dir/settings.json, and modifiedapply_refinement_proposalto refuse factory create/update edits whenfactory_enabledis false, recording them as planned but unapplied withFACTORY_DISABLED_MESSAGE. [78f038a]SessionEngine::factory_activitymethod parsing factory activity requests, performing a blocking preflight for 'run' actions viaFactoryHost(model allowlist/auth/catalog checks), and delegating to the kernel manager'sfactory_activity, and introducedFactoryHostbridge withFactoryHostConfigcapturing session facts for preflight. [78f038a]/factoryCLI command with subcommands list, import, export, addedFactoryCommandenum,parse_factory_commandparser validating subcommands and flags, andrun_factory_commandentrypoint delegating to the kernel's factory CLI dispatcher via a Python runner script. [78f038a]DaemonCommand::FactoryActivityvariant with fields id, active_session_id, action, run_id, spec_id, timeout_ms, and flattened rest; extendedcommand_active_session_id,command_type_name,default_server_capabilities, and protocol tests to support factory_activity. [78f038a]factory_activitymodule implementingfactory_lane_enabledreading settings,advertised_server_capabilitiesconditionally excluding 'factory_activity', andWorker::handle_factory_activityhandler validating payload and delegating toagent_engine.factory_activity, and integrated factory settings into handshake and worker command dispatcher. [78f038a]FactoryViewcomponent,factory_viewmodule with diagram rendering,SessionUi::open_factory_page,handle_factory_view_key, andfactory_controlmethods, integrated factory updates viaFactoryUpdatechannel, and added Factory activity group to the activity dock gated byfactory_activity_supported. [78f038a]Settingswithfactory: Option<FactorySettings>field, addedFactorySettingsstruct with optional enabled bool, and implementedSettingsManager::get_factory_enabledandset_factory_enabledmethods reading/writing the global scope only. [78f038a]/factorybuiltin slash command toCANONICAL_BUILTIN_SLASH_COMMANDSwith client-side execution, argument hint '[on|off|status]', and implementedSessionUi::handle_commandbranch for '/factory' with subcommands 'on', 'off', and informational paths, guarding 'off' against live factory runs. [78f038a]biased;directive and reorder arms to poll targeted events first, ensuring session events published before a response are observed before the reply on the wire, and extended reader task to clear all pending replies inreader_resident.pendingon connection termination, causing in-flight routes to fail immediately. [78f038a]LONG_ROUTE_TIMEOUT_MSwithWORKER_REQUEST_TIMEOUT_MSset to 24 hours, introducedclient_route_timeoutfunction mapping DaemonCommand variants to eitherWORKER_REQUEST_TIMEOUT_MSfor turn-long operations orROUTE_TIMEOUT_MSfor control routes, and updated route timeout usage inSupervisor::route_client_command, handshake, and worker lifecycle methods. [78f038a]interactive_mode::run_interactive_modeto spawn an async flush of startup telemetry viaflush_startup_telemetry, defer joining the flush handle with a timeout after the interactive run completes, and updatedbuild_tui_optionsto usecrate::mode::create_telemetry_disabledfor the telemetry_disabled option. [78f038a]spawn_supervisor_detachedininteractive_mode::daemonandupdate_flow::spawn_supervisorandrun_update_commandto callset_new_sessionon Unix targets before spawning, detaching the supervisor and coordinator processes into new sessions. [78f038a]HeadlessStep::SubmitAndSettleandUiInput::SubmitAndSettlevariants, implementedsettled_sequencehelper to derive the settled event sequence after a prompt, and integrated a submit-and-settle barrier inrun_interactive_surfacethat delays progression until the session processes events up to the settled sequence or timeout. [78f038a]custom_message::custom_message_entriesto map RLM child failure/terminal-notice custom types toAgentMessageentries viarlm_child_status_rowinstead of injected-prompt rows, removedRlmChildStatusandRlmChildOutcomefrominjected_promptmodule, and deleted rendering helpers for child status rows. [78f038a]print_runtime::acp_mode_mainto always resolve daemon socket, ensure daemon running, construct a daemon session create command viadaemon_acp_create, and run daemon-attached ACP mode; deletedtry_daemon_attached_acp; updatedbuild_headless_engine_withto ensure SessionManager is present, constructSemanticEdgeIdentity, and passsemantic_edgesinto engine build options; extractedselect_headless_session,explicit_cwd_override,stored_session_cwdhelpers; and refactoredbuild_session_manager_with_leaseto useselect_headless_session. [78f038a]AgentSessionEnginelifecycle initialization to no longer fall back to PRIME_AGENT_MODEL_PROVIDER/PRIME_AGENT_MODEL env vars for model selection, wirechildren.set_semantic_edgesto record returned children's committed requests, passsemantic_edgesinto a later build step with clonedsemantic_identity, and addedfactory_activitymethod forwarding to the running kernel. [78f038a]Worker::createhandler to constructSemanticSpawnOriginfrom runtimeMetadata when kind is 'subagent', passsemantic_spawninto bind/rebind path, and invokereseed_rlm_childrenafter creation; addedWorker::reseed_rlm_childrenmethod callingreseed_from_ledgeron the children handle; and changedrefresh_replaced_session_stateto async, passingsemantic_spawn: Noneand awaitingreseed_rlm_childrenon success. [78f038a]spawned_by_request_id: NoneinRlmSpawnRequeststructs, and updated engine/options construction in tests to passsemantic_edges: None. [78f038a]a_daemon_restart_relists_and_wakes_the_parents_childanda_daemon_restart_marks_a_still_running_child_as_failedto verify child relisting and status after daemon restart,sigkill_closes_the_spawned_child_and_passivates_the_rowto assert child row appears with status 'error' post-restart,a_spawned_child_ledger_names_its_spawning_request_and_the_parent_records_the_returnto verify semantic linkage and parent recording, andinteractive_launcher_detaches_supervisor_from_client_sessionto validate supervisor session ownership and detachment. [78f038a]hash_runtime_sourceto include packaged machine library files undersrc/rlm/machinesviacollect_package_data_filesin the runtime identity hash, and added a test verifying identity changes when machine files change. [78f038a]review-sweep/MACHINE.md,pr-manager/MACHINE.md, andbuilder/MACHINE.mdwith JSON machine-spec definitions and documentation. [78f038a]Macroscope summarized 8d29f37.
Follow-up heads (the Macroscope round, 2026-10-02)
The five Macroscope findings on this head reduce to two root causes, fixed across four commits:
4880ab3f8— chain superseding stop gates (arm_stopcaptures the gate its arm REPLACED; the stop task opens its own only after that strictly-older gate settles, so a revival gated on the new gate cannot cross before the superseded stop's final snapshot flush, and the strictly-older chain cannot form wait cycles) and generation-scope the settle's listener teardown (the doomed boot's settle clearsstartup_listeners/last_startup_messageonly when its memo is still the active generation). Two new drain-boundary fixtures, fail-first (RED receiptgate_1790910402on the fixtures-only treee3b92c48d).24531b57a— drop the redundant gate clone (the clippy red on the CI run for4880ab3f8).5e3037aef— generation-scope the startup progress EMIT: the settle fix left the emit unscoped, and the listener fixture proved it (the doomed boot'sPreparing Python runtime...clobbered the newer generation's replay stage afterkill()). The emit now writes the sharedlast_startup_messageand fires the listener fan-out only while the emitting memo is still the ACTIVE generation (the boot's own handler still fires).5901a45a7—kill()clears the killed generation's shared progress state (startup_listeners+last_startup_message) in the same lock scope that invalidates the memo, so the nextensure()neither replays the killed boot's stale stage to a fresh handler nor fans the newer boot's stages out to the killed generation's listeners. The listener fixture gains the transition oracles (A's joiner hears nothing after the kill; the newer boot's handler shows exactly its own stage, no stale replay first) with its own RED receipt.4777baaaf— dead generations do not retry and do not surface restores (Macroscope's follow-up finding on5901a45a7): the retry loop checks the boot's generation (boot_generation_is_live: memo still the armed startup generation and not disposed) before each retry and again after the backoff, so akill()during the backoff window cannot resurrect a dead generation's boot; and the restore surface —last_restoreplus theon_restorecallback — is generation-gated in one lock scope (check, write, and callback together, matching the listener contract), so a doomed boot can neither overwrite the fresher generation's restore result nor append a stale restore notice. Fail-first fixturekill_during_the_retry_backoff_pins_the_spawn_count(flaky-first wrapper, kill during the backoff, spawn count pinned at 2; RED receipt: the memo check removed fails 3 vs 2 — the doomed boot consumed its retry).cfcd93af0— a dead generation does not report unavailable skills (Cursor Bugbot's finding on155a87697):on_unavailable_skillsnow uses the restore gate's exact one-lock-scope contract (generation-liveness check and callback share the state lock), so a bootkill()invalidated can no longer append stale skills-unavailable rows to the notice mailbox. Fail-first fixturedoomed_boot_does_not_report_unavailable_skills(RED receipt: the gate reverted fails 2 vs 1 — the dead generation's report lands first).f43fcffc5— empty retrigger commit only, no code change: CI's first run forcfcd93af0failed one unit,pa-cli'sacp_threshold_auto_compaction_publishes_the_compaction_meta(stopReasonnull vsend_turn), a known flaky ACP e2e — the identical test binary was green on this branch at155a87697and fails intermittently on main's own runs atcf285dce(main CI shows both success and failure for that commit). The retrigged run is green.8884c3262— plain fold of main'se7e27c24c(the/settingsmenu search fix, fix keys dying in the /settings menu search #3309; pa-core untouched by the fold). Local gates on the folded tree: fmt rc 0, clippy-D warningsrc 0, kernel fixtures 12/12, pa-tui lib 1341/1341; CI green on the head.c9de47ae0— gate the interleave race harness to unix: main's promote the windows jobs and add the contributor trust gate #3004 promoted the windows cross-check (clippy forx86_64-pc-windows-gnu), red on main's own tip — add race tests for idle session stops and fix two bugs #3195's interleave harness binds a unix-socket probe and tokio gatesUnixListenerbehindall(unix). The module declaration carries#[cfg(unix)](same contract as feed's gated tests; verified green withcargo check -p pa-daemon --all-targets --target x86_64-pc-windows-gnu).3e6868044— spent stop arms release and doomed teardowns never flush (Cursor Bugbot's two Mediums on4c74a09ef, plus two cross-model review findings in the fix shape):pending_stop_for_startupis released by both spenders (the settle and thekill()), so a spent arm's memo receiver cannot keep a parked-or-failed manager alive paststop_kernel; the four failure-teardown sites route throughhold_snapshot_flush_gate— the flush decision (dispose policy while the boot's memo is armed, never a flush for akill()-invalidated boot) and, when flushing, a pending-stop revival-gate install share ONE lock scope with the skip only for a stop actively armed for the boot, so akill()racing the teardown leaves the replacement boot gated on the flush. The two unix-only kernel test targets (kernel_startup_memo.rs,kernel_stop_revive.rs) also carry#![cfg(unix)]— their/bin/shwrapper fixtures never cross-compiled for the new windows job (the sibling kernel targets already carry the gate). RED receipts in the provisioner unit tests (12/12).Cross-model pre-push review (a different model family than the lane's authors, per the repo's review bar): verdict READY-CLEAN after each fix round, including the reviewer's own two findings in the fix commits (a kill-time shared-state leak and a restore-callback race, both fixed with RED→GREEN oracles); report at
/home/ubuntu/handoffs/REVIEW-3303.md. Keeper gates on the pushed tree:cargo fmt --checkrc 0;cargo clippy --workspace --all-targets -- -D warningsrc 0; kernel fixtures 12/12 (kernel_startup_memo7 +kernel_stop_revive5);cargo test -p pa-core --lib1043 passed / 1 git-2.34.1 sparse-checkout artifact inworkspace_snapshot(root files get the skip-worktree bit aftersparse-checkout set --coneon git 2.34; untouched by this PR and green on CI's newer git).