diff --git a/.github/workflows/rust.yml b/.github/workflows/rust.yml index 0ef4e135..b2eec427 100644 --- a/.github/workflows/rust.yml +++ b/.github/workflows/rust.yml @@ -88,9 +88,11 @@ jobs: --locked -- --ignored --exact --nocapture done - name: Test fleet journal and public controller models - # Each case runs concurrent native runtimes and SQL workers. Bound whole - # fixtures on the hosted runner; retain each case's deadlines and races. - run: cargo test -p cellule-host --example fleet_operations --all-features --locked -- --test-threads=2 + # Each case runs concurrent native runtimes and SQL workers. Isolate + # unrelated fixtures on the hosted runner; retain each case's deadlines, + # worker concurrency, protocol races and required evidence. + # Preserve earlier assertion details even if a later native case aborts. + run: cargo test -p cellule-host --example fleet_operations --all-features --locked -- --test-threads=1 --nocapture - name: Verify Axum HTTP publication, retries and cold recovery against RustFS env: AWS_ACCESS_KEY_ID: cellule-ci diff --git a/Cargo.lock b/Cargo.lock index c0109284..0739f8ec 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -268,6 +268,7 @@ dependencies = [ name = "cellule-host" version = "0.1.0" dependencies = [ + "async-trait", "blake3", "bytes", "cellule-app", diff --git a/crates/cellule-host/Cargo.toml b/crates/cellule-host/Cargo.toml index 9ab6951a..d5bd5002 100644 --- a/crates/cellule-host/Cargo.toml +++ b/crates/cellule-host/Cargo.toml @@ -18,6 +18,7 @@ tracing.workspace = true uuid.workspace = true [dev-dependencies] +async-trait.workspace = true bytes.workspace = true cellule-store.workspace = true ed25519-dalek.workspace = true diff --git a/crates/cellule-host/docs/lifecycle.md b/crates/cellule-host/docs/lifecycle.md index 3465442b..e2f1aed8 100644 --- a/crates/cellule-host/docs/lifecycle.md +++ b/crates/cellule-host/docs/lifecycle.md @@ -98,6 +98,36 @@ still reports that original source and keeps lease maintenance and Draining. Dropping the task group aborts its remaining owned tasks. Stopped still requires successful joins through the canonical node drain lane. +## Blob artifact ownership + +During startup, declare `BLOB_ARTIFACT_STORE_COMPONENT` in the required component +set, install the task group, then call `install_blob_artifact_store` with the +product-configured provider. Pass the returned clone into the typed client's +existing Blob configuration. Duplicate installation, closed original admission +and installation after startup are refused before retaining another facility. + +The store uses the existing reverse-order facility drain. Every clone closes +admission and the drain joins original namespace/GC operations before runtime +shutdown. Cancelling a shutdown waiter preserves both the original host drain +and accepted Blob work. A deadline may return an incomplete drain; it cannot +turn a still-running operation into closed-plus-zero or Stopped. The next drain +joins the retained original work. An unrelated store has independent admission. + +The store's lifecycle observation distinguishes local joining from result +success. Native source-bearing failures remain available through a retained +clone. A returned prepared command holds no running job until execution through +the configured client acquires this original store admission; closure refuses +new dispatch from its clones and reconstructed snapshots. Already accepted +commands still require their original result evidence. Complete uploads, +external streams, read/backup +pins, global references and unknown remote effects still require their original +owners and fleet barriers. Installation does not remove `BlobInventory` or +authorize Blob Cell release, global GC or physical-node finalization. +An original supervisor lost during forced runtime teardown retains unjoined +work and closes admission. Its later provider completion does not repair that +join. Store drain returns an error, so the host cannot use a zero task count to +claim Stopped or withdrawal. + ## Requested node-log rotation Install the existing durability provider during startup, then call @@ -594,8 +624,9 @@ in either attachment order; matching registry revisions alone are insufficient. This closure settles that boot's enrollment. It does not convert a recovered tombstone into planned withdrawal or prove replacement policy, affected-writer relocation, operation completion or permission to stop the physical node. -`SettleRoles`/`Finalize` still require those additional barriers and the existing -native shutdown handoff. The [focused example cases](../minion/README.md#failed-boot-process-evidence) +`SettleRoles` and the complete role/movement evidence needed to enter Closing +remain separate barriers. Finalize runs through the existing native shutdown +handoff only after Closing is committed. The [focused example cases](../minion/README.md#failed-boot-process-evidence) exercise a joined child lifetime and independently reconstructed evidence; they do not qualify a complete multi-process fleet or provider deployment. @@ -723,6 +754,48 @@ For failed-source recovery, `recovery_inputs` supplies an existing canonical `NodeTakeoverProof` and recovery manifest store. Ordinary node recovery first fences the failed session and seals/pins its required tail; the provider lookup performs none of those effects. +For a process-closed receiver that claimed ownership after a proven release, +`receiver_recovery_inputs` resolves prerequisites for the exact failed control +and accepted route. Its default returns no proof. The host confirms a separate +`ReceiverRecoveryBasis` before native takeover and `ReceiverRecoveryEvidence` +before actor admission. The journal must persist both atomically against the +original acceptance and reject mixing this basis with an Idle acquisition basis. +The successful outcome remains Activated: the original clean release is retained. +Fresh serving verifies canonical acquisition history, any pinned recovered +suffix, and native derivation from the release root. Missing provider/journal +evidence keeps the attempt charged. Deploy readers for record kinds 30 and 31 +before enabling this path; old source-recovery records keep their exact binding. + +An interrupted receiver evidence write may leave a safe Idle rollback root. +Replay of the original accepted effect reconfirms the immutable takeover input +and materialization from canonical acquisition history, then records the missing +evidence. Before resuming ordinary admitted acquisition, it verifies that Idle +root against the recovery and clean release prefixes. The original basis and +capture times remain unchanged; no second movement permit or Idle basis replaces +them. If ordinary acquisition already won, replay verifies the current actor +against the same history. Missing, corrupt, or substituted history leaves Unknown +and the full charge. Fresh inspection performs no evidence writes or acquisition. + +Failed-source Recover uses the same ordering after a failed evidence write or +lost reply: confirm the original accepted `RecoveryBasis`, reconstruct missing +`RecoveryEvidence` only from its exact immutable native acquisition, and retain +the original materialization. Effect replay verifies the selected Idle root +against that materialization and any original sealed suffix before ordinary +admitted acquisition. An ordinary acquisition winner must pass the same native +history, origin and actor checks. Missing or substituted history remains a +charged blocker even when the current writer can read its local SQLite state. +The suffix verifier for this pre-acquisition path requires an unowned Idle +control with cleared overlay and rechecks its complete value before and after +origin I/O; it grants no serving or settlement proof. + +If an interrupted takeover still owns a Recovering control with its pinned +overlay, replay supplies the original retained control and failed-session proof +to `resume_takeover_restored_observed`. That shared native path reconfirms the +original basis, materializes the same suffix and records the result before actor +admission, without another ownership epoch. Failed-source Recover and routed +receiver Activate use this continuation. A current claim that does not derive +exactly from the original checked input remains unresolved. + The application authenticates management requests. The journal atomically checks the current controller epoch, head, permit, intent, endpoint, and deadline when first accepting an action. `AcceptedFleetAction` validates its @@ -744,14 +817,16 @@ complete primitive maintenance coverage and role finalization remain under imple | Source preflight refusal | The actor returns `Error::CellReleaseRefused` only before this request begins canonical deactivation. Record Rejected with its original source error; the reconciler then cancels and joins unused receiver credit. Independent local eviction grants no release proof for this attempt. | | Prepare receiver | Validate catalog/Cell/incarnation and current release/schema support; reserve actual runtime/LTX resources before returning Reserved. A refusal preserves the original error and leaves the source serving. | | Maintenance Cordon | Apply retained Draining intent through the shared admission gate. Replays retain the exact acceptance/result; a dropped waiter leaves publication owned. This preserves existing obligations and uses no shutdown lane. | -| Other maintenance effects | Refuse role settlement and finalization until their inventory and host barriers are implemented. Fresh maintenance inspection reads the exact local admission gate: only Draining reports Cordoned, while Active/Cordoned stays blocked. This inspection cannot establish role settlement, Stopped, or withdrawal. | +| SettleRoles | Require the opaque complete role and replacement-policy proof through `apply_fleet_role_settlement`. A reconstructed controller refreshes the proof at its current request/head while retaining the original acceptance and stable action key. Publication compares the proof's head and registry atomically; an old proof cannot authorize Closing. | +| Finalize | Require committed Closing evidence. Close finite-action admission and join accepted work through the canonical node drain; publish Stopped only after shutdown and exact boot withdrawal/retirement. | +| Maintenance inspection | Read the exact local admission gate: only Draining reports Cordoned, while Active/Cordoned stays blocked. This inspection cannot establish role settlement, Stopped, or withdrawal. | | Activate receiver | Confirm the exact Idle acquisition basis is durably retained before ordinary ownership CAS; consume prepared credit through canonical restore and actor activation. | | Acquisition-basis reply lost | Keep authority untouched and the prepared credit charged. Repeating the same accepted action confirms the original basis and capture time before takeover. | | Unused credit after release | Journal CleaningReceiver, cancel/join that credit, and retain the source position and fleet permit. Cleanup returns to Released, with no claim of pre-release cancellation. After confirmed cleanup, ordinary admitted acquisition may resume. | | Ordinary acquisition on the preferred session | Free the unused preparation without closing the ordinary writer. Verify current authority, actor readiness, and the required release position. Matching session alone never proves prepared credit was consumed. | | Recover unresolved source release | Journal Recovering, cancel proved-unused preparation, validate canonical failed-session proof, and confirm `RecoveryBasis` before ownership CAS. Confirm `RecoveryEvidence` after exact recovery publication and before actor admission. Return Recovered only with current actor/authority proof. | | Recovery input reply lost | No ownership CAS starts. Retain the original input/time; an unchanged full canonical control permits confirming that basis and continuing the accepted action. | -| Recovery position reply lost | No actor is admitted. Canonical rollback preserves the materialized root; the attempt stays charged and Unknown. Ordinary acquisition can restore serving, after which inspection verifies it against the retained recovery evidence. | +| Recovery position write fails or reply is lost | No actor is admitted. Canonical rollback preserves the materialized root; the attempt stays charged and Unknown. Accepted effect replay reconfirms exact native history and retained evidence before admitted Idle reacquisition, or verifies an ordinary acquisition winner. Fresh inspection cannot repair metadata or acquire. | | Duplicate | Compare the full immutable specification, including cost and physical identities. Join current work or return its committed result. | | Dropped caller | Retain and finish execution plus journal publication independently of the caller. | | Result publication failure | Retain the original checked result. Result delivery joins the owned task, so an immediate subsequent dispatch can retry publication. Drain also retries; neither repeats source release. | @@ -777,6 +852,21 @@ exact-root restoration, and rollback path. The journal atomically binds basis and evidence to original acceptance, rejects changed inputs, and returns the original time for identical writes. +`SettleRoles` validates read-only evidence and starts no physical effect. If +publication of its joined retained proof fails again, the executor returns the +original publication error and removes that local observation. The durable +acceptance and any already-committed historical result stay in the journal. +The failed pass cannot authorize Closing: another pass must collect all native +and foreign roles and replacement policies afresh, then commit a proof at the +current full head/registry. Movement receipts keep their native evidence across +every failed publication retry. + +A cancelled settlement waiter leaves its original native task owned. A changed +head does not replace that running job; a competing request remains blocked with +its original Conflict. After joining, a failed publication requires fresh +capture as above. Neither local receipt removal nor an empty action bank proves +physical role settlement or node closure. + The executor uses at most two retained action jobs and charges their envelopes, results, and acquisition inputs to the existing node retained-byte ledger. Unknown work retains its fleet permit in the journal. A missing local receipt, @@ -903,6 +993,20 @@ retained original requests; Pending requests without a current reference still require their original nonexecution or closure evidence. Applications account the bounded copied buffer, with at most 10,000 rows and 128 rows per page. +For several physical followers, `FleetFollowerReferences::collect_all` shares +one fresh canonical directory traversal per continuation round. Requests are +unique, intent-bound and ordered; their combined page limit is 128 rows, with +one lookahead per window. Each member keeps the same complete cursor, topology, +exact rows and original interval as individual collection. Every new traversal +authenticates all records before filtering, including expired obligations. + +`FleetFollowerReferences::recheck_all` first invalidates every earlier +confirmation, then checks every member against a fresh traversal and the full +roster. Only a wholly matching set advances confirmation intervals. Failure or +cancellation preserves all original rows/times and leaves the whole set +unconfirmed. A retry must freshly confirm every member before role coverage is +valid again. Applications account the complete per-member copied buffers. + After native and policy collection, `references.recheck(...)` traverses every page again and compares exact rows. Native topology fingerprints intentionally omit volatile coverage and leader liveness, so a first-page fingerprint cannot @@ -975,8 +1079,9 @@ follower inventories. Persisted history, enrollment and transport codecs are unc `reader_evacuations()` and `follower_evacuations()` expose the retained originals; each check's `record()` supplies its immutable durable history. A supplied subset does not establish complete policy coverage or upgrade an incomplete observation. -All remaining roles, failed-owner recovery, original accepted native/external work -and terminal drain handoff are still required before SettleRoles/Finalize. +All remaining roles, failed-owner recovery and original accepted native/external +work must be settled before SettleRoles and before the operation can enter +Closing. Finalize then owns the terminal drain handoff and exact withdrawal. ### Retain the complete original maintenance role set diff --git a/crates/cellule-host/minion/README.md b/crates/cellule-host/minion/README.md index 8562c51d..ccf3fd84 100644 --- a/crates/cellule-host/minion/README.md +++ b/crates/cellule-host/minion/README.md @@ -1,21 +1,202 @@ # Minion fleet operations reference -Status: durable journal and public-driver models, plus a finite real three-node -admission-overload and controller-restart scenarios with receipt readback. -The `balance` command exercises real count convergence. Complete maintenance, -receiver-loss, production observations and -qualification remain required by the [fleet plan](../../../docs/fleet-operations-plan.md). +Status: durable journal and public-driver models, plus finite real-node +overload, maintenance, controller-restart, receiver-loss, and count-balance +scenarios with receipt readback. Reader and follower maintenance use three or +four managed boots with native role evacuation. Role-enabled production +observations and qualification remain required by the [fleet plan](../../../docs/fleet-operations-plan.md). ## Run +Maintenance restart tests lose the reply after a real follower donor commits +`SettleRoles`, fail its result write, or lose acknowledgement after that write +commits, then reconstruct the driver over an independent SQLite client. The original acceptance/key remains unchanged; a fresh complete proof +must cover the current head before Closing. Six cases renew the same claimant +and replace it after the real controller lease expires. They check final +withdrawal, canonical-root restoration, the original command result, and joined +resources: + +```sh +cargo test -p cellule-host --example fleet_operations --all-features --locked \ + scenario::follower_tests::persisted::restart:: -- --test-threads=1 +``` + +The expiry case keeps the original boots live with actual signed heartbeats, +then checks the new claimant/epoch and the old controller's fenced refusal. +The result-write cases require the old proof's Conflict, joined local retirement +and a fresh complete capture before Closing. The original acceptance, native +error and any committed historical receipt remain unchanged. + +Cancellation cases pause the actual result transaction/reply, cancel its +controller waiter, and join an exact duplicate to the retained native owner: + +```sh +cargo test -p cellule-host --example fleet_operations --all-features --locked \ + scenario::follower_tests::persisted::cancelled_settlement:: -- --test-threads=1 +``` + +A new head cannot replace the running original or authorize closure with its old +proof. External process/provider failures and the complete W7 fault matrix still +require separate evidence. + +The focused receiver-loss tests cover both unchanged Idle authority and a failed +receiver that reached Recovering or Serving after clean source release: + +```sh +cargo test -p cellule-host --example fleet_operations successor_tests:: --locked +``` + +They use real nodes, canonical directory takeover proof, immutable receiver +recovery records, lost reply replay and an independently reopened journal after +controller lease expiry. No configured takeover proof keeps the failed claim +unresolved. The stopped-process witness is an explicit in-process reference +provider; process/provider qualification remains required. + +The `recovery_faults::` cases pause actual receiver basis/evidence transactions +before commit or after commit before reply. They check cancellation on both +sides of takeover, repeated evidence-write failure, safe Idle resumption and an +ordinary acquisition winner. Missing, corrupt and valid substituted native +history cannot authorize continuation. Each completed case verifies the original +receipt/value, immutable records through an independent journal client, and +joined node resources. These selected cases retain two-worker native races; +CI runs unrelated minion fixtures serially to isolate their finite deadlines. + +Run receiver loss after clean source release through the public reconciler: + +```sh +cargo run -p cellule-host --example fleet_operations --locked -- receiver-loss +``` + +This reserves a receiver, releases one Cell, joins the receiver's runtime, and +publishes typed closure of its exact enrolled boot. The reconciler retains the +original release and routes activation and cleanup to the third node. The +command loses a committed activation reply, replays the exact accepted action, +reopens an independent journal client after controller lease expiry, and checks +that the previous controller is fenced. It verifies all twelve command receipts +and SQL values, fences the moved source handle, and checks final Cell counts +`[11, 0, 1]`. Exit joins all three nodes, retires their boots, closes both journal +clients, and checks empty resource ledgers. Failures preserve their original +error through the same cleanup. This finite command uses joined in-process +lifecycle evidence; it does not qualify OS crash or external-job supervision. + Run the real-node scenario from the workspace root: ```sh cargo run -p cellule-host --example fleet_operations --locked -- overload ``` -It creates three independently leased CellNodes over shared strict in-memory -object storage, initializes twelve real SQLite Cells, and acknowledges a command +Run planned maintenance through the public reconciler and native node actions: + +```sh +cargo run -p cellule-host --example fleet_operations --locked -- maintenance +``` + +This cordons the first node, moves all twelve Cells to the two eligible nodes, +settles the complete writer-only role inventory, finalizes and withdraws the +exact boot, and verifies every command receipt on its destination. The example +proves an empty reader/follower obligation set for its closed writer-only +constructor; role-enabled deployments and provider/process failures remain +separate qualification requirements. + +Run planned maintenance under continuous SQL command admission: + +```sh +cargo run -p cellule-host --example fleet_operations --locked -- maintenance-busy +``` + +Two client lanes continuously mutate one original donor Cell while the same +public reconciler cordons the node, prepares receiver reservations and dispatches +native busy release. The actor's immutable owner fence remains observable while +publication owns its publisher. Root advancement invalidates complete counts; +explicit maintenance can retain demand for that independently verified writer. +Source quiescence follows receiver preparation and joins accepted publication +before canonical release. No extra application quiescence call is used. + +Each lane must receive an admission refusal before its finite 512-command bound; +exhausting that load is failure. The command resolves every acknowledged request +on its canonical successor and compares the exact digest, sequence and stored +result. An audit table must contain exactly one row per acknowledgement. All +twelve original receipts and SQL values also survive movement. Output reports +accepted commands, both admission refusals and exact audit rows. Success and +failure join every client task before node/journal cleanup; all three boots +retire and runtime resource ledgers empty. This finite in-process workload does +not qualify external primitive owners, process crashes or measured provider load. + +Run reader-only maintenance with a live foreign writer: + +```sh +cargo run -p cellule-host --example fleet_operations --locked -- maintenance-reader +``` + +This registers three managed boots and a real reader, acknowledges a mutation +to 29, then submits maintenance for the reader's physical node through the same +public driver. Missing replacement policy keeps the operation Evacuating and +the original reader usable. After signed reader selection observes the cordon, +the command opens a replacement, performs canonical reader evacuation and +publishes immutable policy evidence. Fresh complete observation authorizes +SettleRoles and native Finalize. It checks Completed, Stopped, exact withdrawal, +fenced original reader, receipt-bound replacement readback and the original +writer's stored mutation result. Cleanup joins all three nodes, both reader +enrollments and boot rows, and checks the existing resource ledgers. Output +keeps zero writer releases/activations separate from two receipt checks and +reader replacement. This finite in-process example does not qualify external +process failure or the remaining follower/primitive combinations. + +Run live-follower maintenance with two-member durability: + +```sh +cargo run -p cellule-host --example fleet_operations --locked -- maintenance-follower +``` + +Four managed boots provide a writer, two original follower stores and one spare. +The donor has no writers but retains a foreign follower lane. A complete pass +without replacement policy stays Evacuating and retains its original enrollment. +The command enables the spare through a real signed heartbeat, covers the old +acknowledged tail through canonical writer drain, and requests rotation from the +existing durability supervisor. Both original member retirements and exact old +coverage must be confirmed before epoch 2 with the two eligible members can +settle the donor obligation. Canonical acquisition resumes the writer on its +original node. A new command is acknowledged through the existing durability +gate while the installed replacement ensemble is rechecked. Object proof may win before the +directory activation CAS; Active alone is not required to prove installation. +Fresh durable policy and complete role observation authorize native Finalize. +The command checks Completed, Stopped, permanent withdrawal, original retirement +history and canonical-root SQL readback of 29, covering both acknowledgements. +The original request keeps its digest, sequence, expiry and stored outcome. Cleanup joins all four nodes and +retires both ensembles and every boot; final writer counts include the spare: +`[1, 0, 0, 0]`. Release/activation counts describe fleet writer movement and remain +zero. Five service/receipt checks and two replacement members are reported +separately. + +The finite live-owner adapter retains exactly two prepared epochs and exact +original close barriers in memory. It uses fresh canonical append/retire +authorization and offers no failed-owner seal/tail capability. External process +restart, remote authentication and the complete primitive maintenance matrix +remain separate qualification requirements. + +Run combined reader and live-follower maintenance on one physical donor: + +```sh +cargo run -p cellule-host --example fleet_operations --locked -- maintenance-roles +``` + +This extends the same four-boot follower scenario with managed readers installed +before readiness. The donor owns a reader and a foreign follower tail. A durable +follower evacuation policy alone must leave maintenance Evacuating. The native +reader evacuation then refuses while one selected replacement has no Established +reader, preserving the original enrollment and open view. After both eligible +receivers establish real readers, native evacuation closes and joins the donor +view and publishes immutable reader policy evidence. Complete fresh observation +must prove both roles before SettleRoles and Finalize. The command checks both +replacement readers against the captured receipt, preserves the original +mutation and renewed writer acknowledgement, and joins all four nodes, eleven +enrollment rows, and every resource ledger. Writer counts remain `[1, 0, 0, 0]`; +eight service/receipt checks are reported separately from zero fleet writer +moves. This qualifies an in-process reader/follower combination; busy primitive +owners, process/provider failures and reboot remain required by the fleet plan. + +The overload command creates three independently leased CellNodes over shared +strict in-memory object storage, initializes twelve real SQLite Cells, and acknowledges a command on each. A held seven-GiB disk admission token causes the existing actor's measured ledger classifier to enter Shedding after its normal dwell. The public reconciler allocates its bounded two-move batch; the scenario releases that @@ -312,8 +493,10 @@ BeginEvacuation binds it to the manifest. Session adoption cannot erase that anchor, and a missing anchor refuses capture. Phase replay cannot repair a missing manifest by inserting zero; missing or corrupt pages retain their original errors. Newly accepted remote replacements after capture remain visible in the current roster -and native graph. Complete policy/work coverage and terminal joining remain -required; this historical set does not enable SettleRoles or Finalize. +and native graph. Complete policy/work coverage remains required before +SettleRoles. This historical set alone grants neither SettleRoles nor Finalize; +Finalize requires a separately committed Closing operation and the node's bound +managed-boot withdrawal. ## Exact original reader joining @@ -366,9 +549,10 @@ phase publication before dispatch, fresh activation/retirement, independent cleanup, retained permits after lost replies/timeouts, competing controllers, stop-new-moves behavior, and pressure relief using remaining shared budget. The separate `scenario::tests` cases invoke the same real-node implementations -as the `overload` and `controller-restart` commands. The model tests do not create Cell actors. Planned -maintenance and receiver-loss scenarios still require their complete barriers; -controller replacement here retains all three live node sessions. +as the `overload`, `maintenance`, and `controller-restart` commands. The model +tests do not create Cell actors. Foreign follower evacuation, receiver loss, +and provider/process qualification remain incomplete; controller replacement +here retains all three live node sessions. The host now owns each whole closing attempt after a shutdown waiter disappears. It retains the same lane through facility/runtime/lease cleanup and exposes local @@ -379,8 +563,9 @@ version and SQLite enrollment row into that owner before readiness. Native shutdown checks canonical withdrawal and committed boot retirement before Stopped; a lost retirement reply retains Draining until an exact replay confirms the original evidence. Partial startup still uses the retained application's -boot cleanup. These local observations supply no relocation or redundancy proof -and do not enable the still-refused fleet Finalize action. +boot cleanup. These local observations supply no relocation or redundancy +proof. Fleet Finalize consumes separately committed Closing evidence and uses +this same retained shutdown and withdrawal owner. Publish native rotation/reader proofs before starting terminal shutdown: weak native handles can disappear when an autonomous successful stop clears owners. @@ -579,8 +764,23 @@ cargo test -p cellule-host --example fleet_operations --all-features --locked sc ``` Complete authenticated lookup, source/failed-owner successor policy, checked -nonexecution, accepted-work joining and SettleRoles/Finalize remain delivery -requirements. See the [host matching recipe](../docs/lifecycle.md#match-every-required-maintenance-policy). +nonexecution and complete SettleRoles behavior remain delivery requirements. +Finalize runs after the journal commits Closing; it joins accepted action work +and publishes Stopped only after exact withdrawal. +See the [host matching recipe](../docs/lifecycle.md#match-every-required-maintenance-policy). + +## Native source reader succession + +The [source reader recipe](../docs/source-readers.md) composes exact retained +native retirement with an actual managed writer successor and current ready +reader policy. It consumes the complete original/current request set and is +retained through the public observer and matcher. Missing evidence remains an +explicit source obligation. A checked local request does not finish maintenance. + +```sh +cargo test -p cellule-host --example fleet_operations --all-features --locked \ + scenario::reader_tests::evacuation::source:: -- --test-threads=2 +``` ## Native source reader succession diff --git a/crates/cellule-host/minion/journal/action_result_faults.rs b/crates/cellule-host/minion/journal/action_result_faults.rs new file mode 100644 index 00000000..63c97a3f --- /dev/null +++ b/crates/cellule-host/minion/journal/action_result_faults.rs @@ -0,0 +1,85 @@ +//! Failures and cancelled waiters at the actual role-result write/reply. +use super::*; + +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub(crate) enum ResultWriteBoundary { + BeforeCommit, + AfterCommit, +} + +pub(super) struct RoleResultFault { + boundary: ResultWriteBoundary, + pause: Option, +} + +struct RoleResultPause { + entered: tokio::sync::oneshot::Sender<()>, + resume: tokio::sync::oneshot::Receiver<()>, +} + +impl SqliteJournal { + pub(crate) fn fail_next_role_result(&self, boundary: ResultWriteBoundary) { + self.arm_role_result(RoleResultFault { + boundary, + pause: None, + }); + } + + pub(crate) fn pause_next_role_result( + &self, + boundary: ResultWriteBoundary, + ) -> ( + tokio::sync::oneshot::Receiver<()>, + tokio::sync::oneshot::Sender<()>, + ) { + let (entered, captured) = tokio::sync::oneshot::channel(); + let (resume, paused) = tokio::sync::oneshot::channel(); + self.arm_role_result(RoleResultFault { + boundary, + pause: Some(RoleResultPause { + entered, + resume: paused, + }), + }); + (captured, resume) + } + + fn arm_role_result(&self, fault: RoleResultFault) { + let mut slot = self.inner.role_result_fault.lock().unwrap(); + assert!( + slot.replace(fault).is_none(), + "role result fault already armed" + ); + } + + pub(super) async fn role_result_boundary( + &self, + role_settlement: bool, + boundary: ResultWriteBoundary, + ) -> JournalResult<()> { + let fault = { + let mut slot = self.inner.role_result_fault.lock().unwrap(); + if role_settlement + && slot + .as_ref() + .is_some_and(|fault| fault.boundary == boundary) + { + slot.take() + } else { + None + } + }; + if let Some(fault) = fault { + if let Some(pause) = fault.pause { + let _ = pause.entered.send(()); + let _ = pause.resume.await; + } + return Err(std::io::Error::other(match boundary { + ResultWriteBoundary::BeforeCommit => "injected role result before commit", + ResultWriteBoundary::AfterCommit => "injected role result after commit", + }) + .into()); + } + Ok(()) + } +} diff --git a/crates/cellule-host/minion/journal/actions.rs b/crates/cellule-host/minion/journal/actions.rs index 957bd455..2932f48a 100644 --- a/crates/cellule-host/minion/journal/actions.rs +++ b/crates/cellule-host/minion/journal/actions.rs @@ -1,7 +1,16 @@ use super::*; +#[derive(Clone)] +struct ClosedReceiverProof { + snapshot: FleetJournalSnapshot, + node: NodeId, + session: SessionId, + digest: Digest, + retired: bool, +} + impl Db<'_> { - fn accepted( + pub(super) fn accepted( &self, key: Digest, node: NodeId, @@ -45,6 +54,171 @@ impl Db<'_> { }) .transpose() } + + fn check_receiver_route(&self, action: &FleetAction) -> JournalResult<()> { + let Some(route) = action.receiver_route() else { + return Ok(()); + }; + let FleetActionKind::Movement { attempt, .. } = action.kind() else { + return Err(OperationError::Conflict.into()); + }; + let spec = attempt.spec(); + let rows = { + let mut statement = self.tx.prepare( + "SELECT key,node,session FROM actions WHERE operation=?1 AND sequence=?2 AND length(sequence)=8 LIMIT 65", + )?; + statement + .query_map( + params![ + spec.id.operation.as_bytes().as_slice(), + spec.id.sequence.to_be_bytes().as_slice() + ], + |row| Ok((blob(row, 0, 32)?, blob(row, 1, 16)?, blob(row, 2, 16)?)), + )? + .collect::>>()? + }; + if rows.len() > 64 { + return Err(OperationError::Conflict.into()); + } + let mut latest: Option = None; + for (key, node, session) in rows { + let key = Digest::from_bytes( + key.as_slice() + .try_into() + .map_err(|_| OperationError::Invalid("action index width"))?, + ); + let node = NodeId::from_bytes( + node.as_slice() + .try_into() + .map_err(|_| OperationError::Invalid("node identity width"))?, + ); + let session = SessionId::from_bytes( + session + .as_slice() + .try_into() + .map_err(|_| OperationError::Invalid("session identity width"))?, + ); + let (accepted, _) = self + .accepted(key, node, session)? + .ok_or(OperationError::NotFound)?; + let FleetActionKind::Movement { + attempt: recorded, .. + } = accepted.action().kind() + else { + return Err(OperationError::Conflict.into()); + }; + if accepted.action().scope() != action.scope() + || recorded.spec() != spec + || accepted.node() != node + || accepted.session() != session + { + return Err(OperationError::Conflict.into()); + } + let Some(recorded_route) = accepted.action().receiver_route() else { + continue; + }; + if let Some(previous) = &latest { + if recorded_route.hop_count() == previous.hop_count() { + if recorded_route != previous { + return Err(OperationError::Conflict.into()); + } + } else if recorded_route.hop_count() > previous.hop_count() { + if !recorded_route.follows(previous) { + return Err(OperationError::Conflict.into()); + } + latest = Some(recorded_route.clone()); + } + } else { + latest = Some(recorded_route.clone()); + } + } + match latest { + Some(previous) if route.follows(&previous) => Ok(()), + None if route.hop_count() == 1 => Ok(()), + _ => Err(OperationError::Conflict.into()), + } + } + + fn accept_action( + &self, + action: FleetAction, + node: NodeId, + session: SessionId, + now_ms: i64, + closed: Option, + ) -> JournalResult { + self.check_scope(action.scope())?; + let key = action.key()?; + if let Some((accepted, result)) = self.accepted(key, node, session)? { + accepted.validate_replay(&action, node, session)?; + return Ok(FleetActionAcceptance::Existing { + accepted, + result: result.map(Box::new), + }); + } + let snapshot = self.snapshot()?; + match (action.receiver_route(), closed) { + (Some(route), Some(closed)) => { + let hop = route + .latest_handoff() + .ok_or(OperationError::Invalid("receiver route lacks final hop"))?; + if !closed.retired + || hop.previous() != (closed.node, closed.session) + || hop.process_closure() != closed.digest + || hop.registry() != snapshot.registry() + || closed.snapshot.head() != snapshot.head() + || closed.snapshot.registry() != snapshot.registry() + || action.receiver_endpoint() != Some((node, session)) + { + return Err(OperationError::Conflict.into()); + } + } + (Some(_), None) => { + return Err(OperationError::Fenced.into()); + } + (None, Some(_)) => { + return Err(OperationError::Invalid( + "closed receiver proof supplied for an ordinary action", + ) + .into()); + } + (None, None) => {} + } + let accepted = AcceptedFleetAction::new_with_registry( + action.clone(), + snapshot.head(), + snapshot.registry(), + node, + session, + now_ms, + )?; + self.check_intents(&action, node, session)?; + self.check_receiver_route(&action)?; + let (operation, sequence, effect) = match action.kind() { + FleetActionKind::Movement { action, attempt } => ( + attempt.spec().id.operation, + attempt.spec().id.sequence.to_be_bytes().to_vec(), + *action as u8, + ), + FleetActionKind::Maintenance { action, operation } => { + (operation.id(), vec![], *action as u8) + } + }; + self.tx.execute( + "INSERT INTO actions(key,node,session,operation,sequence,effect,accepted) VALUES (?1,?2,?3,?4,?5,?6,?7)", + params![ + key.as_bytes().as_slice(), + node.as_bytes().as_slice(), + session.as_bytes().as_slice(), + operation.as_bytes().as_slice(), + sequence, + effect, + accepted.to_bytes()? + ], + )?; + Ok(FleetActionAcceptance::New(accepted)) + } + fn original(&self, accepted: &AcceptedFleetAction) -> JournalResult<()> { let (original, _) = self .accepted( @@ -58,7 +232,11 @@ impl Db<'_> { } Ok(()) } - fn basis(&self, accepted: &AcceptedFleetAction, kind: u8) -> JournalResult>> { + pub(super) fn basis( + &self, + accepted: &AcceptedFleetAction, + kind: u8, + ) -> JournalResult>> { self.original(accepted)?; Ok(self .tx @@ -125,8 +303,10 @@ impl Db<'_> { | MovementAction::Activate | MovementAction::Recover ) { - let receiver = self.required_intent(spec.destination_node)?; - if receiver.session() != spec.destination || receiver.mode() != NodeMode::Active + let (receiver_node, receiver_session) = + action.receiver_endpoint().ok_or(OperationError::Conflict)?; + let receiver = self.required_intent(receiver_node)?; + if receiver.session() != receiver_session || receiver.mode() != NodeMode::Active { return Err(OperationError::Conflict.into()); } @@ -201,22 +381,26 @@ impl FleetActionJournal for SqliteJournal { now_ms: i64, ) -> FleetAdapterFuture<'a, FleetActionAcceptance> { let action = action.clone(); - Box::pin(self.run(move |db| { - db.check_scope(action.scope())?; - let key = action.key()?; - if let Some((accepted, result)) = db.accepted(key, node, session)? { - accepted.validate_replay(&action, node, session)?; - return Ok(FleetActionAcceptance::Existing { accepted, result: result.map(Box::new) }); - } - let accepted = AcceptedFleetAction::new(action.clone(), db.snapshot()?.head(), node, session, now_ms)?; - db.check_intents(&action, node, session)?; - let (operation, sequence, effect) = match action.kind() { - FleetActionKind::Movement { action, attempt } => (attempt.spec().id.operation, attempt.spec().id.sequence.to_be_bytes().to_vec(), *action as u8), - FleetActionKind::Maintenance { action, operation } => (operation.id(), vec![], *action as u8), - }; - db.tx.execute("INSERT INTO actions(key,node,session,operation,sequence,effect,accepted) VALUES (?1,?2,?3,?4,?5,?6,?7)", params![key.as_bytes().as_slice(), node.as_bytes().as_slice(), session.as_bytes().as_slice(), operation.as_bytes().as_slice(), sequence, effect, accepted.to_bytes()?])?; - Ok(FleetActionAcceptance::New(accepted)) - })) + Box::pin(self.run(move |db| db.accept_action(action, node, session, now_ms, None))) + } + + fn accept_closed_receiver_action<'a>( + &'a self, + action: &'a FleetAction, + node: NodeId, + session: SessionId, + now_ms: i64, + closure: &'a FleetFailedBootClosure, + ) -> FleetAdapterFuture<'a, FleetActionAcceptance> { + let action = action.clone(); + let closed = ClosedReceiverProof { + snapshot: closure.snapshot().clone(), + node: closure.canonical().node(), + session: closure.canonical().session(), + digest: closure.digest(), + retired: closure.boot().status() == EnrollmentStatus::Retired, + }; + Box::pin(self.run(move |db| db.accept_action(action, node, session, now_ms, Some(closed)))) } fn publish_action_result<'a>( &'a self, @@ -225,19 +409,47 @@ impl FleetActionJournal for SqliteJournal { ) -> FleetAdapterFuture<'a, ()> { let accepted = accepted.clone(); let result = result.clone(); - Box::pin(self.run(move |db| { + #[cfg(test)] + let role_settlement = matches!(&result.outcome, FleetOutcome::RolesSettledAt { .. }); + let write = self.run(move |db| { db.original(&accepted)?; accepted.validate_result(&result)?; + if let FleetOutcome::RolesSettledAt { + registry, + head_revision, + .. + } = &result.outcome + { + let current = db.snapshot()?; + if current.registry() != *registry || current.head().revision() != *head_revision { + return Err(OperationError::Conflict.into()); + } + } if let FleetActionKind::Movement { action: MovementAction::Activate, .. } = accepted.action().kind() && matches!(result.outcome, FleetOutcome::Activated(_)) { - let basis = AcquisitionBasis::from_bytes( - &db.basis(&accepted, 1)?.ok_or(OperationError::NotFound)?, - )?; - basis.validate_result(&result)?; + if let Some(bytes) = db.basis(&accepted, 1)? { + if db.receiver_basis(&accepted, 1)?.is_some() { + return Err(OperationError::Conflict.into()); + } + AcquisitionBasis::from_bytes(&bytes)?.validate_result(&result)?; + } else { + let evidence = ReceiverRecoveryEvidence::from_bytes( + &db.receiver_basis(&accepted, 2)? + .ok_or(OperationError::NotFound)?, + )?; + let basis = ReceiverRecoveryBasis::from_bytes( + &db.receiver_basis(&accepted, 1)? + .ok_or(OperationError::NotFound)?, + )?; + if basis.accepted() != &accepted || evidence.basis() != &basis { + return Err(OperationError::Conflict.into()); + } + evidence.validate_result(&result)?; + } } if let FleetOutcome::Recovered(recovered) = &result.outcome && let FleetActionKind::Movement { @@ -263,7 +475,42 @@ impl FleetActionJournal for SqliteJournal { if previous == result { return Ok(()); } - if !matches!(previous.outcome, FleetOutcome::Unknown) { + let refreshable_roles = matches!( + accepted.action().kind(), + FleetActionKind::Maintenance { + action: MaintenanceAction::SettleRoles, + .. + } + ) && matches!( + (&previous.outcome, &result.outcome), + ( + FleetOutcome::RolesSettled { .. } | FleetOutcome::RolesSettledAt { .. }, + FleetOutcome::RolesSettledAt { .. } + ) + ) && result.observed_at_ms > previous.observed_at_ms + && match (&previous.outcome, &result.outcome) { + ( + FleetOutcome::RolesSettledAt { + registry: previous, + head_revision: previous_head, + .. + }, + FleetOutcome::RolesSettledAt { + registry: current, + head_revision: current_head, + .. + }, + ) => { + current.revision() >= previous.revision() + && current_head >= previous_head + } + ( + FleetOutcome::RolesSettled { .. }, + FleetOutcome::RolesSettledAt { .. }, + ) => true, + _ => false, + }; + if !matches!(previous.outcome, FleetOutcome::Unknown) && !refreshable_roles { return Err(OperationError::Conflict.into()); } if matches!(result.outcome, FleetOutcome::Unknown) { @@ -280,7 +527,17 @@ impl FleetActionJournal for SqliteJournal { ], )?; Ok(()) - })) + }); + Box::pin(async move { + #[cfg(test)] + self.role_result_boundary(role_settlement, ResultWriteBoundary::BeforeCommit) + .await?; + write.await?; + #[cfg(test)] + self.role_result_boundary(role_settlement, ResultWriteBoundary::AfterCommit) + .await?; + Ok(()) + }) } fn load_movement_action<'a>( &'a self, @@ -304,6 +561,87 @@ impl FleetActionJournal for SqliteJournal { }).transpose() })) } + + fn load_movement_actions<'a>( + &'a self, + scope: FleetScope, + attempt: &'a MoveAttempt, + effect: MovementAction, + ) -> FleetAdapterFuture<'a, Vec> { + let attempt = attempt.clone(); + Box::pin(self.run(move |db| { + db.check_scope(scope)?; + let rows = { + let mut statement = db.tx.prepare( + "SELECT key,node,session FROM actions WHERE operation=?1 AND sequence=?2 AND effect=?3 ORDER BY node,session LIMIT 4", + )?; + statement + .query_map( + params![ + attempt.spec().id.operation.as_bytes().as_slice(), + attempt.spec().id.sequence.to_be_bytes().as_slice(), + effect as u8 + ], + |row| { + Ok(( + blob(row, 0, 32)?, + blob(row, 1, 16)?, + blob(row, 2, 16)?, + )) + }, + )? + .collect::>>()? + }; + if rows.len() > MAX_RECEIVER_HANDOFFS + 1 { + return Err(OperationError::Conflict.into()); + } + rows.into_iter() + .map(|(key, node, session)| { + let key = Digest::from_bytes( + key.as_slice() + .try_into() + .map_err(|_| OperationError::Invalid("action index width"))?, + ); + let node = NodeId::from_bytes( + node.as_slice() + .try_into() + .map_err(|_| OperationError::Invalid("node identity width"))?, + ); + let session = SessionId::from_bytes( + session + .as_slice() + .try_into() + .map_err(|_| OperationError::Invalid("session identity width"))?, + ); + let (accepted, result) = db + .accepted(key, node, session)? + .ok_or(OperationError::NotFound)?; + let FleetActionKind::Movement { + action, + attempt: original, + } = accepted.action().kind() + else { + return Err(OperationError::Conflict.into()); + }; + if accepted.action().scope() != scope + || accepted.action().key()? != key + || *action != effect + || original.spec() != attempt.spec() + || accepted.node() != node + || accepted.session() != session + { + return Err(OperationError::Conflict.into()); + } + accepted.validate_replay(accepted.action(), node, session)?; + Ok(FleetActionAcceptance::Existing { + accepted, + result: result.map(Box::new), + }) + }) + .collect() + })) + } + fn record_acquisition_basis<'a>( &'a self, basis: &'a AcquisitionBasis, @@ -311,6 +649,9 @@ impl FleetActionJournal for SqliteJournal { let basis = basis.clone(); Box::pin(self.run(move |db| { let bytes = basis.to_bytes()?; + if db.receiver_basis(basis.accepted(), 1)?.is_some() { + return Err(OperationError::Conflict.into()); + } if let Some(bytes) = db.basis(basis.accepted(), 1)? { let original = AcquisitionBasis::from_bytes(&bytes)?; if original.accepted() != basis.accepted() || original.control() != basis.control() @@ -421,4 +762,29 @@ impl FleetActionJournal for SqliteJournal { .transpose() })) } + + fn record_receiver_recovery_basis<'a>( + &'a self, + basis: &'a ReceiverRecoveryBasis, + ) -> FleetAdapterFuture<'a, ReceiverRecoveryBasis> { + self.receiver_recovery_basis(basis) + } + fn load_receiver_recovery_basis<'a>( + &'a self, + accepted: &'a AcceptedFleetAction, + ) -> FleetAdapterFuture<'a, Option> { + self.receiver_recovery_input(accepted) + } + fn record_receiver_recovery_evidence<'a>( + &'a self, + evidence: &'a ReceiverRecoveryEvidence, + ) -> FleetAdapterFuture<'a, ReceiverRecoveryEvidence> { + self.receiver_recovery_evidence(evidence) + } + fn load_receiver_recovery_evidence<'a>( + &'a self, + accepted: &'a AcceptedFleetAction, + ) -> FleetAdapterFuture<'a, Option> { + self.receiver_recovery_result(accepted) + } } diff --git a/crates/cellule-host/minion/journal/controller.rs b/crates/cellule-host/minion/journal/controller.rs index f2a86ff5..68df9326 100644 --- a/crates/cellule-host/minion/journal/controller.rs +++ b/crates/cellule-host/minion/journal/controller.rs @@ -61,6 +61,81 @@ impl FleetJournal for SqliteJournal { // deadline/session progress. Replays cannot start it again. return Ok(current); } + if matches!( + transition, + JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(_)) + ) { + let operation = current + .head() + .maintenance() + .ok_or(OperationError::NotFound)?; + let key = FleetAction::maintenance_action_key( + current.head().scope(), + MaintenanceAction::SettleRoles, + operation, + )?; + let (accepted, result) = db + .accepted( + key, + operation.node(), + operation.session(), + )? + .ok_or(OperationError::Invalid( + "ready-to-close lacks committed role settlement", + ))?; + let accepted_operation = match accepted.action().kind() { + FleetActionKind::Maintenance { + action: MaintenanceAction::SettleRoles, + operation, + } => operation, + _ => { + return Err(OperationError::Invalid( + "ready-to-close action is not SettleRoles", + ) + .into()); + } + }; + if accepted_operation.id() != operation.id() + || accepted_operation.request_digest() != operation.request_digest() + || accepted_operation.node() != operation.node() + || accepted_operation.session() != operation.session() + || accepted_operation.intent_revision() != operation.intent_revision() + || accepted_operation.phase() != MaintenancePhase::Evacuating + { + return Err(OperationError::Invalid( + "ready-to-close SettleRoles identity differs", + ) + .into()); + } + let result = result.ok_or(OperationError::Invalid( + "ready-to-close lacks committed role settlement", + ))?; + let (result_registry, result_head_revision) = match result.outcome { + FleetOutcome::RolesSettledAt { + registry, + head_revision, + .. + } => (registry, head_revision), + _ => { + return Err(OperationError::Invalid( + "ready-to-close lacks current successful role settlement", + ) + .into()); + } + }; + if result_registry != current.registry() { + return Err(OperationError::Conflict.into()); + } + if result_head_revision != current.head().revision() { + return Err(OperationError::Conflict.into()); + } + if result.observed_at_ms > now_ms { + return Err(OperationError::Invalid( + "ready-to-close lacks current successful role settlement", + ) + .into()); + } + } let head = current.head().transition(db.profile, current.head().revision(), epoch, now_ms, transition.clone())?; if head == *current.head() { return Ok(current); } if matches!(transition, JournalTransition::Maintenance(MaintenanceEvent::BeginEvacuation)) { diff --git a/crates/cellule-host/minion/journal/mod.rs b/crates/cellule-host/minion/journal/mod.rs index 60e20b91..30bd0f83 100644 --- a/crates/cellule-host/minion/journal/mod.rs +++ b/crates/cellule-host/minion/journal/mod.rs @@ -12,12 +12,21 @@ use cellule_runtime::node::NodeMode; use rusqlite::{Connection, OptionalExtension, Transaction, TransactionBehavior, params}; use tokio::sync::{Notify, Semaphore}; +#[cfg(test)] +mod action_result_faults; mod actions; +#[cfg(test)] +pub(crate) use action_result_faults::ResultWriteBoundary; mod controller; mod enrollment; mod follower_evacuation; mod maintenance_enrollments; mod reader_evacuation; +mod receiver_recovery; +#[cfg(test)] +mod receiver_recovery_faults; +#[cfg(test)] +pub(crate) use receiver_recovery_faults::{RecoveryWrite, RecoveryWriteBoundary}; mod records; mod writer_inventory; use records::{Db, blob, profile_bytes, scope_bytes}; @@ -39,6 +48,8 @@ struct Inner { scope: FleetScope, profile: FleetProfile, #[cfg(test)] + role_result_fault: Mutex>, + #[cfg(test)] lose_commit_reply: std::sync::atomic::AtomicBool, #[cfg(test)] boot_reply: Mutex>, @@ -50,6 +61,8 @@ struct Inner { follower_evacuation_reply: Mutex>, #[cfg(test)] original_writer_reply: Mutex>, + #[cfg(test)] + receiver_recovery_reply: Mutex>, } #[cfg(test)] @@ -136,6 +149,8 @@ impl SqliteJournal { scope, profile, #[cfg(test)] + role_result_fault: Mutex::new(None), + #[cfg(test)] lose_commit_reply: std::sync::atomic::AtomicBool::new(false), #[cfg(test)] boot_reply: Mutex::new(None), @@ -147,6 +162,8 @@ impl SqliteJournal { follower_evacuation_reply: Mutex::new(None), #[cfg(test)] original_writer_reply: Mutex::new(None), + #[cfg(test)] + receiver_recovery_reply: Mutex::new(None), }), }) } diff --git a/crates/cellule-host/minion/journal/receiver_recovery.rs b/crates/cellule-host/minion/journal/receiver_recovery.rs new file mode 100644 index 00000000..8cc87f60 --- /dev/null +++ b/crates/cellule-host/minion/journal/receiver_recovery.rs @@ -0,0 +1,171 @@ +//! Receiver recovery has its own table; existing source basis kinds are fixed. +use super::*; + +impl Db<'_> { + pub(super) fn receiver_basis( + &self, + accepted: &AcceptedFleetAction, + kind: u8, + ) -> JournalResult>> { + // Exact original acceptance is required in the same transaction. + let (original, _) = self + .accepted( + accepted.action().key()?, + accepted.node(), + accepted.session(), + )? + .ok_or(OperationError::NotFound)?; + if &original != accepted { + return Err(OperationError::Conflict.into()); + } + Ok(self.tx.query_row("SELECT body FROM receiver_recoveries WHERE key=?1 AND node=?2 AND session=?3 AND kind=?4", params![accepted.action().key()?.as_bytes().as_slice(), accepted.node().as_bytes().as_slice(), accepted.session().as_bytes().as_slice(), kind], |row| blob(row, 0, MAX_RECORD_BYTES)).optional()?) + } + fn write_receiver_basis( + &self, + accepted: &AcceptedFleetAction, + kind: u8, + bytes: Vec, + ) -> JournalResult<()> { + self.receiver_basis(accepted, kind)?; + self.tx.execute( + "INSERT INTO receiver_recoveries(key,node,session,kind,body) VALUES (?1,?2,?3,?4,?5)", + params![ + accepted.action().key()?.as_bytes().as_slice(), + accepted.node().as_bytes().as_slice(), + accepted.session().as_bytes().as_slice(), + kind, + bytes + ], + )?; + Ok(()) + } +} + +impl SqliteJournal { + pub(super) fn receiver_recovery_basis<'a>( + &'a self, + basis: &'a ReceiverRecoveryBasis, + ) -> FleetAdapterFuture<'a, ReceiverRecoveryBasis> { + let basis = basis.clone(); + Box::pin(async move { + #[cfg(test)] + self.receiver_recovery_write_boundary( + RecoveryWrite::Basis, + RecoveryWriteBoundary::BeforeCommit, + ) + .await?; + let retained = self + .run(move |db| { + let bytes = basis.to_bytes()?; + if db.basis(basis.accepted(), 1)?.is_some() { + return Err(OperationError::Conflict.into()); + } + if let Some(bytes) = db.receiver_basis(basis.accepted(), 1)? { + let original = ReceiverRecoveryBasis::from_bytes(&bytes)?; + if original.accepted() != basis.accepted() + || original.control() != basis.control() + { + return Err(OperationError::Conflict.into()); + } + return Ok(original); + } + db.write_receiver_basis(basis.accepted(), 1, bytes)?; + Ok(basis) + }) + .await?; + #[cfg(test)] + self.receiver_recovery_write_boundary( + RecoveryWrite::Basis, + RecoveryWriteBoundary::AfterCommit, + ) + .await?; + Ok(retained) + }) + } + pub(super) fn receiver_recovery_input<'a>( + &'a self, + accepted: &'a AcceptedFleetAction, + ) -> FleetAdapterFuture<'a, Option> { + let accepted = accepted.clone(); + Box::pin(self.run(move |db| { + db.receiver_basis(&accepted, 1)? + .map(|bytes| { + let basis = ReceiverRecoveryBasis::from_bytes(&bytes)?; + if basis.accepted() != &accepted { + return Err(OperationError::Conflict.into()); + } + Ok(basis) + }) + .transpose() + })) + } + pub(super) fn receiver_recovery_evidence<'a>( + &'a self, + evidence: &'a ReceiverRecoveryEvidence, + ) -> FleetAdapterFuture<'a, ReceiverRecoveryEvidence> { + let evidence = evidence.clone(); + Box::pin(async move { + #[cfg(test)] + self.receiver_recovery_write_boundary( + RecoveryWrite::Evidence, + RecoveryWriteBoundary::BeforeCommit, + ) + .await?; + let retained = self + .run(move |db| { + let accepted = evidence.basis().accepted(); + let basis = ReceiverRecoveryBasis::from_bytes( + &db.receiver_basis(accepted, 1)? + .ok_or(OperationError::NotFound)?, + )?; + if &basis != evidence.basis() { + return Err(OperationError::Conflict.into()); + } + let bytes = evidence.to_bytes()?; + if let Some(bytes) = db.receiver_basis(accepted, 2)? { + let original = ReceiverRecoveryEvidence::from_bytes(&bytes)?; + if original.basis() != evidence.basis() + || original.restored() != evidence.restored() + { + return Err(OperationError::Conflict.into()); + } + return Ok(original); + } + db.write_receiver_basis(accepted, 2, bytes)?; + Ok(evidence) + }) + .await?; + #[cfg(test)] + self.receiver_recovery_write_boundary( + RecoveryWrite::Evidence, + RecoveryWriteBoundary::AfterCommit, + ) + .await?; + Ok(retained) + }) + } + pub(super) fn receiver_recovery_result<'a>( + &'a self, + accepted: &'a AcceptedFleetAction, + ) -> FleetAdapterFuture<'a, Option> { + let accepted = accepted.clone(); + Box::pin(self.run(move |db| { + db.receiver_basis(&accepted, 2)? + .map(|bytes| { + let evidence = ReceiverRecoveryEvidence::from_bytes(&bytes)?; + if evidence.basis().accepted() != &accepted { + return Err(OperationError::Conflict.into()); + } + let basis = ReceiverRecoveryBasis::from_bytes( + &db.receiver_basis(&accepted, 1)? + .ok_or(OperationError::NotFound)?, + )?; + if &basis != evidence.basis() { + return Err(OperationError::Conflict.into()); + } + Ok(evidence) + }) + .transpose() + })) + } +} diff --git a/crates/cellule-host/minion/journal/receiver_recovery_faults.rs b/crates/cellule-host/minion/journal/receiver_recovery_faults.rs new file mode 100644 index 00000000..2c1d1505 --- /dev/null +++ b/crates/cellule-host/minion/journal/receiver_recovery_faults.rs @@ -0,0 +1,76 @@ +//! Faults at the actual receiver record transaction and reply boundaries. +use super::*; + +#[derive(Clone, Copy, PartialEq, Eq)] +pub(crate) enum RecoveryWrite { + Basis, + Evidence, +} + +#[derive(Clone, Copy, PartialEq, Eq)] +pub(crate) enum RecoveryWriteBoundary { + BeforeCommit, + AfterCommit, +} + +pub(super) struct RecoveryWritePause { + write: RecoveryWrite, + boundary: RecoveryWriteBoundary, + fail: bool, + entered: tokio::sync::oneshot::Sender<()>, + resume: tokio::sync::oneshot::Receiver<()>, +} + +impl SqliteJournal { + pub(crate) fn pause_receiver_recovery_write( + &self, + write: RecoveryWrite, + boundary: RecoveryWriteBoundary, + fail: bool, + ) -> ( + tokio::sync::oneshot::Receiver<()>, + tokio::sync::oneshot::Sender<()>, + ) { + let (entered, captured) = tokio::sync::oneshot::channel(); + let (resume, paused) = tokio::sync::oneshot::channel(); + let mut slot = self.inner.receiver_recovery_reply.lock().unwrap(); + assert!(slot.is_none(), "receiver recovery fault already armed"); + *slot = Some(RecoveryWritePause { + write, + boundary, + fail, + entered, + resume: paused, + }); + (captured, resume) + } + + pub(super) async fn receiver_recovery_write_boundary( + &self, + write: RecoveryWrite, + boundary: RecoveryWriteBoundary, + ) -> JournalResult<()> { + let pause = { + let mut slot = self.inner.receiver_recovery_reply.lock().unwrap(); + if slot + .as_ref() + .is_some_and(|pause| pause.write == write && pause.boundary == boundary) + { + slot.take() + } else { + None + } + }; + if let Some(pause) = pause { + let _ = pause.entered.send(()); + let _ = pause.resume.await; + if pause.fail { + return Err(std::io::Error::other( + "injected receiver recovery write reply failure", + ) + .into()); + } + } + Ok(()) + } +} diff --git a/crates/cellule-host/minion/journal/schema.sql b/crates/cellule-host/minion/journal/schema.sql index 18663ff0..9611e5bd 100644 --- a/crates/cellule-host/minion/journal/schema.sql +++ b/crates/cellule-host/minion/journal/schema.sql @@ -29,11 +29,18 @@ CREATE TABLE IF NOT EXISTS actions ( PRIMARY KEY(key,node,session) ); CREATE UNIQUE INDEX IF NOT EXISTS movement_index ON actions(operation,sequence,effect,node,session) WHERE length(sequence)=8; +DROP INDEX IF EXISTS maintenance_action_index; CREATE TABLE IF NOT EXISTS bases ( key BLOB NOT NULL, node BLOB NOT NULL, session BLOB NOT NULL, kind INTEGER NOT NULL CHECK (kind IN (1,2,3)), body BLOB NOT NULL CHECK (length(body)<=65536), PRIMARY KEY(key,node,session,kind), FOREIGN KEY(key,node,session) REFERENCES actions(key,node,session) ); +CREATE TABLE IF NOT EXISTS receiver_recoveries ( + key BLOB NOT NULL CHECK(length(key)=32), node BLOB NOT NULL CHECK(length(node)=16), + session BLOB NOT NULL CHECK(length(session)=16), kind INTEGER NOT NULL CHECK(kind IN (1,2)), + body BLOB NOT NULL CHECK(length(body)<=65536), PRIMARY KEY(key,node,session,kind), + FOREIGN KEY(key,node,session) REFERENCES actions(key,node,session) +); -- Role evidence shares registry revisions and the same accepted backend jobs. CREATE TABLE IF NOT EXISTS reader_evacuations ( diff --git a/crates/cellule-host/minion/journal/tests/maintenance_enrollments.rs b/crates/cellule-host/minion/journal/tests/maintenance_enrollments.rs index f7fac10e..459ca342 100644 --- a/crates/cellule-host/minion/journal/tests/maintenance_enrollments.rs +++ b/crates/cellule-host/minion/journal/tests/maintenance_enrollments.rs @@ -28,6 +28,175 @@ async fn freeze(fixture: &Fixture) -> FleetJournalSnapshot { .await } +#[tokio::test] +async fn ready_to_close_requires_committed_settlement_at_the_current_registry_version() { + let fixture = Fixture::new().await; + let snapshot = freeze(&fixture).await; + let close = + JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(ready_drain_proof(3))); + let missing = fixture + .journal + .compare_exchange(&snapshot, 1, 0, &close) + .await + .unwrap_err(); + assert!(matches!( + missing.downcast_ref::(), + Some(OperationError::Invalid(_)) + )); + + let operation = snapshot.head().maintenance().unwrap(); + let action = snapshot + .head() + .maintenance_action(MaintenanceAction::SettleRoles, 0) + .unwrap(); + let accepted = match fixture + .journal + .accept_action(&action, operation.node(), operation.session(), 0) + .await + .unwrap() + { + FleetActionAcceptance::New(accepted) => accepted, + FleetActionAcceptance::Existing { .. } => panic!("first SettleRoles acceptance expected"), + }; + let unknown = result(&accepted, FleetOutcome::Unknown, 0); + fixture + .journal + .publish_action_result(&accepted, &unknown) + .await + .unwrap(); + let unresolved = fixture + .journal + .compare_exchange(&snapshot, 1, 0, &close) + .await + .unwrap_err(); + assert!(matches!( + unresolved.downcast_ref::(), + Some(OperationError::Invalid(_)) + )); + + let settled = result( + &accepted, + FleetOutcome::RolesSettledAt { + inventory: Digest::from_bytes([85; 32]), + head_revision: snapshot.head().revision(), + registry: snapshot.registry(), + }, + 0, + ); + fixture + .journal + .publish_action_result(&accepted, &settled) + .await + .unwrap(); + let closed = fixture + .journal + .compare_exchange(&snapshot, 1, 0, &close) + .await + .unwrap(); + assert_eq!( + closed.head().maintenance().unwrap().phase(), + MaintenancePhase::Closing + ); +} + +#[tokio::test] +async fn registry_change_after_settlement_fences_ready_to_close() { + let fixture = Fixture::new().await; + let snapshot = freeze(&fixture).await; + roles_settled(&fixture.journal, 0).await; + let settled = fixture.journal.load_snapshot(scope()).await.unwrap(); + let changed = fixture + .journal + .set_scheduling(settled.registry(), false) + .await + .unwrap(); + let current = fixture.journal.load_snapshot(scope()).await.unwrap(); + assert_eq!(current.registry(), changed); + let error = fixture + .journal + .compare_exchange( + ¤t, + 1, + 0, + &JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(ready_drain_proof(3))), + ) + .await + .unwrap_err(); + assert!(matches!( + error.downcast_ref::(), + Some(OperationError::Conflict) + )); + assert_eq!( + fixture.journal.load_snapshot(scope()).await.unwrap().head(), + snapshot.head() + ); +} + +#[tokio::test] +async fn head_change_after_settlement_fences_ready_to_close() { + let fixture = Fixture::new().await; + freeze(&fixture).await; + roles_settled(&fixture.journal, 0).await; + let settled = fixture.journal.load_snapshot(scope()).await.unwrap(); + let current = fixture + .journal + .claim_controller( + scope(), + settled.head().revision(), + SessionId::from_bytes([206; 16]), + 1, + ) + .await + .unwrap(); + assert_eq!(current.registry(), settled.registry()); + assert_ne!(current.head().revision(), settled.head().revision()); + let error = fixture + .journal + .compare_exchange( + ¤t, + current.head().controller().unwrap().epoch, + 1, + &JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(ready_drain_proof(3))), + ) + .await + .unwrap_err(); + assert!(matches!( + error.downcast_ref::(), + Some(OperationError::Conflict) + )); +} + +#[tokio::test] +async fn deadline_revision_can_accept_a_new_settle_roles_action_for_the_same_boot() { + let fixture = Fixture::new().await; + freeze(&fixture).await; + roles_settled(&fixture.journal, 0).await; + let before = fixture.journal.load_snapshot(scope()).await.unwrap(); + let operation = before.head().maintenance().unwrap(); + let first_key = + FleetAction::maintenance_action_key(scope(), MaintenanceAction::SettleRoles, operation) + .unwrap(); + fixture + .journal + .compare_exchange( + &before, + before.head().controller().unwrap().epoch, + 1, + &JournalTransition::Maintenance(MaintenanceEvent::ExtendDeadline( + operation.deadline_ms() + 10_000, + )), + ) + .await + .unwrap(); + roles_settled(&fixture.journal, 1).await; + let after = fixture.journal.load_snapshot(scope()).await.unwrap(); + let operation = after.head().maintenance().unwrap(); + let second_key = + FleetAction::maintenance_action_key(scope(), MaintenanceAction::SettleRoles, operation) + .unwrap(); + assert_ne!(first_key, second_key); +} + #[tokio::test] async fn original_roles_commit_with_evacuation_and_survive_lost_reply_retirement_and_restart() { let fixture = Fixture::new().await; diff --git a/crates/cellule-host/minion/journal/tests/mod.rs b/crates/cellule-host/minion/journal/tests/mod.rs index a9ab4546..31e8fafc 100644 --- a/crates/cellule-host/minion/journal/tests/mod.rs +++ b/crates/cellule-host/minion/journal/tests/mod.rs @@ -128,9 +128,46 @@ fn drain_proof(node: u8) -> DrainEvidence { withdrawn: true, } } +fn ready_drain_proof(node: u8) -> DrainEvidence { + let mut evidence = drain_proof(node); + evidence.facilities_closed = false; + evidence.stopped = false; + evidence.withdrawn = false; + evidence +} + +async fn roles_settled(journal: &SqliteJournal, now_ms: i64) { + let snapshot = journal.load_snapshot(scope()).await.unwrap(); + let operation = snapshot.head().maintenance().unwrap(); + let action = snapshot + .head() + .maintenance_action(MaintenanceAction::SettleRoles, now_ms) + .unwrap(); + let accepted = match journal + .accept_action(&action, operation.node(), operation.session(), now_ms) + .await + .unwrap() + { + FleetActionAcceptance::New(accepted) => accepted, + FleetActionAcceptance::Existing { .. } => panic!("first SettleRoles acceptance expected"), + }; + let settled = result( + &accepted, + FleetOutcome::RolesSettledAt { + inventory: Digest::from_bytes([84; 32]), + head_revision: snapshot.head().revision(), + registry: snapshot.registry(), + }, + now_ms, + ); + journal + .publish_action_result(&accepted, &settled) + .await + .unwrap(); +} #[tokio::test] -async fn closing_source_fences_new_roles_and_preserves_original_completion_after_reconstruction() { +async fn closing_source_fences_new_roles_and_preserves_settled_roles_after_reconstruction() { for role in [ EnrollmentRole::Follower { log_epoch: 7 }, EnrollmentRole::Reader { @@ -162,10 +199,32 @@ async fn closing_source_fences_new_roles_and_preserves_original_completion_after FleetEnrollmentAcceptance::New(record) => record, _ => panic!("first acceptance expected"), }; + let established = fixture + .journal + .publish_enrollment_result( + &accepted, + EnrollmentEvent::Established(Digest::from_bytes([62; 32])), + 2, + ) + .await + .unwrap(); + let retired = fixture + .journal + .publish_enrollment_result( + &established, + EnrollmentEvent::Retired(Digest::from_bytes([63; 32])), + 3, + ) + .await + .unwrap(); + roles_settled(&fixture.journal, 4).await; fixture - .transition(JournalTransition::Maintenance( - MaintenanceEvent::ReadyToClose(drain_proof(3)), - )) + .transition_at( + JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(ready_drain_proof( + 3, + ))), + 4, + ) .await; fixture.journal.close().await.unwrap(); let restarted = fixture.client().await; @@ -185,24 +244,15 @@ async fn closing_source_fences_new_roles_and_preserves_original_completion_after ); assert!(matches!( restarted.accept_enrollment(&original, 2).await.unwrap(), - FleetEnrollmentAcceptance::Existing(record) if record == accepted + FleetEnrollmentAcceptance::Existing(record) if record == retired )); - let established = restarted - .publish_enrollment_result( - &accepted, - EnrollmentEvent::Established(Digest::from_bytes([62; 32])), - 3, - ) - .await - .unwrap(); - let retired = restarted - .publish_enrollment_result( - &established, - EnrollmentEvent::Retired(Digest::from_bytes([63; 32])), - 4, - ) - .await - .unwrap(); + assert_eq!( + restarted + .load_enrollment(scope(), original.key().unwrap()) + .await + .unwrap(), + Some(retired.clone()) + ); assert_eq!(retired.accepted_at_ms(), accepted.accepted_at_ms()); assert!(!retired.unresolved()); restarted.close().await.unwrap(); @@ -212,7 +262,7 @@ async fn closing_source_fences_new_roles_and_preserves_original_completion_after .compare_exchange( ¤t, 1, - 4, + 5, &JournalTransition::Maintenance(MaintenanceEvent::Stopped(drain_proof(3))), ) .await @@ -223,15 +273,15 @@ async fn closing_source_fences_new_roles_and_preserves_original_completion_after .compare_exchange( &completed, 1, - 4, + 5, &JournalTransition::BeginMaintenance(request(2, 4)), ) .await .unwrap(); - assert!(resumed.accept_enrollment(&fresh, 5).await.is_err()); + assert!(resumed.accept_enrollment(&fresh, 6).await.is_err()); assert_eq!(resumed.load_snapshot(scope()).await.unwrap(), later); assert!(matches!( - resumed.accept_enrollment(&original, 5).await.unwrap(), + resumed.accept_enrollment(&original, 6).await.unwrap(), FleetEnrollmentAcceptance::Existing(record) if record == retired )); resumed.close().await.unwrap(); @@ -252,6 +302,7 @@ async fn delayed_source_acceptance_checks_closing_in_its_original_transaction() MaintenanceEvent::BeginEvacuation, )) .await; + roles_settled(&fixture.journal, 0).await; let independent = fixture.client().await; let mut spec = enrollment(64, 1); spec.source.as_mut().unwrap().intent_revision = 2; @@ -261,9 +312,10 @@ async fn delayed_source_acceptance_checks_closing_in_its_original_transaction() let task = tokio::spawn(async move { client.accept_enrollment(&original, 1).await }); paused.await.unwrap(); let closed = fixture - .transition(JournalTransition::Maintenance( - MaintenanceEvent::ReadyToClose(drain_proof(3)), - )) + .transition_at( + JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(ready_drain_proof(3))), + 1, + ) .await; resume.send(()).unwrap(); assert!(task.await.unwrap().is_err()); @@ -292,11 +344,13 @@ async fn source_acceptance_and_closing_have_one_committed_registry_barrier() { fixture .transition(JournalTransition::Maintenance(MaintenanceEvent::Cordoned)) .await; - let expected = fixture + fixture .transition(JournalTransition::Maintenance( MaintenanceEvent::BeginEvacuation, )) .await; + roles_settled(&fixture.journal, 0).await; + let expected = fixture.journal.load_snapshot(scope()).await.unwrap(); let independent = fixture.client().await; let mut enrollment_spec = enrollment(65, 1); enrollment_spec.source.as_mut().unwrap().intent_revision = 2; @@ -307,7 +361,7 @@ async fn source_acceptance_and_closing_have_one_committed_registry_barrier() { }; } let transition = - JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(drain_proof(3))); + JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(ready_drain_proof(3))); let (accepted, closing) = tokio::join!( independent.accept_enrollment(&enrollment_spec, 1), fixture @@ -692,12 +746,19 @@ impl Fixture { } } async fn transition(&self, transition: JournalTransition) -> FleetJournalSnapshot { + self.transition_at(transition, 0).await + } + async fn transition_at( + &self, + transition: JournalTransition, + now_ms: i64, + ) -> FleetJournalSnapshot { let current = self.journal.load_snapshot(scope()).await.unwrap(); self.journal .compare_exchange( ¤t, current.head().controller().unwrap().epoch, - 0, + now_ms, &transition, ) .await @@ -751,6 +812,44 @@ impl Fixture { } } +#[tokio::test] +async fn routed_receiver_acceptance_requires_typed_closed_boot_proof() { + let fixture = Fixture::new().await; + fixture.activating().await; + let snapshot = fixture.journal.load_snapshot(scope()).await.unwrap(); + let route = ReceiverRoute::begin( + scope(), + &spec(1), + endpoint(3).node, + endpoint(3).session, + Digest::from_bytes([207; 32]), + snapshot.registry(), + ) + .unwrap(); + let action = snapshot + .head() + .movement_action_with_receiver_route(spec(1).id, MovementAction::Activate, route, 1) + .unwrap(); + assert!( + fixture + .journal + .accept_action(&action, endpoint(3).node, endpoint(3).session, 1) + .await + .is_err() + ); + let records = fixture + .journal + .load_movement_actions( + scope(), + &snapshot.head().attempts()[0], + MovementAction::Activate, + ) + .await + .unwrap(); + assert_eq!(records.len(), 1); + fixture.journal.close().await.unwrap(); +} + #[tokio::test] async fn reconstruction_preserves_scope_profile_header_and_stopped_policy() { let fixture = Fixture::new().await; @@ -1565,18 +1664,8 @@ async fn later_operations_retain_old_cordons_and_return_to_service_does_not_rewr MaintenanceEvent::BeginEvacuation, )) .await; - let proof = DrainEvidence { - node: endpoint(1).node, - session: endpoint(1).session, - remaining_cells: 0, - unresolved_attempts: 0, - relocated: true, - readers_settled: true, - followers_settled: true, - facilities_closed: true, - stopped: true, - withdrawn: true, - }; + roles_settled(&fixture.journal, 0).await; + let proof = ready_drain_proof(1); fixture .transition(JournalTransition::Maintenance( MaintenanceEvent::ReadyToClose(proof), @@ -1584,7 +1673,7 @@ async fn later_operations_retain_old_cordons_and_return_to_service_does_not_rewr .await; fixture .transition(JournalTransition::Maintenance(MaintenanceEvent::Stopped( - proof, + drain_proof(1), ))) .await; let completed = fixture diff --git a/crates/cellule-host/minion/main.rs b/crates/cellule-host/minion/main.rs index 7746777a..d2036d3a 100644 --- a/crates/cellule-host/minion/main.rs +++ b/crates/cellule-host/minion/main.rs @@ -14,18 +14,30 @@ use journal::{JournalResult, SqliteJournal}; async fn main() -> JournalResult<()> { let mut args = std::env::args_os().skip(1); let command = args.next(); - if matches!(command.as_deref(), Some(value) if value == "overload" || value == "controller-restart" || value == "balance") + if matches!(command.as_deref(), Some(value) if value == "overload" || value == "controller-restart" || value == "balance" || value == "maintenance" || value == "maintenance-busy" || value == "maintenance-reader" || value == "maintenance-follower" || value == "maintenance-roles" || value == "receiver-loss") && args.next().is_none() { let summary = if command.as_deref() == Some(std::ffi::OsStr::new("controller-restart")) { scenario::controller_restart().await? } else if command.as_deref() == Some(std::ffi::OsStr::new("balance")) { scenario::count_balance().await? + } else if command.as_deref() == Some(std::ffi::OsStr::new("maintenance")) { + scenario::maintenance().await? + } else if command.as_deref() == Some(std::ffi::OsStr::new("maintenance-busy")) { + scenario::maintenance_busy().await? + } else if command.as_deref() == Some(std::ffi::OsStr::new("maintenance-reader")) { + scenario::maintenance_reader().await? + } else if command.as_deref() == Some(std::ffi::OsStr::new("maintenance-follower")) { + scenario::maintenance_follower().await? + } else if command.as_deref() == Some(std::ffi::OsStr::new("maintenance-roles")) { + scenario::maintenance_roles().await? + } else if command.as_deref() == Some(std::ffi::OsStr::new("receiver-loss")) { + scenario::receiver_loss().await? } else { scenario::overload().await? }; println!( - "released={} activated={} retired={} receipt_checks={} max_inflight={} max_restore_bytes={} joined_nodes={} boot_retirements={} receiver_nodes={} lost_release_replies={} controller_epoch={} expired_receiver_cleanups={} blocker_count={} final_counts={:?}", + "released={} activated={} retired={} receipt_checks={} max_inflight={} max_restore_bytes={} joined_nodes={} boot_retirements={} receiver_nodes={} lost_release_replies={} controller_epoch={} expired_receiver_cleanups={} blocker_count={} final_counts={:?} maintenance_completed={} maintenance_boot_withdrawn={} receiver_process_closures={} lost_activation_replies={} routed_activation_replays={}", summary.released, summary.activated, summary.retired, @@ -39,7 +51,12 @@ async fn main() -> JournalResult<()> { summary.controller_epoch, summary.expired_receiver_cleanups, summary.blockers.len(), - summary.final_counts + summary.final_counts, + summary.maintenance_completed, + summary.maintenance_boot_withdrawn, + summary.receiver_process_closures, + summary.lost_activation_replies, + summary.routed_activation_replays ); println!("blockers={:?}", summary.blockers); return Ok(()); @@ -50,7 +67,7 @@ async fn main() -> JournalResult<()> { || args.next().is_some() { return Err(std::io::Error::other( - "usage: fleet_operations overload | controller-restart | balance | inspect-journal ", + "usage: fleet_operations overload | controller-restart | balance | maintenance | maintenance-busy | maintenance-reader | maintenance-follower | maintenance-roles | receiver-loss | inspect-journal ", ) .into()); } diff --git a/crates/cellule-host/minion/reconciler_tests.rs b/crates/cellule-host/minion/reconciler_tests.rs index 45f9a2d0..ad94ffee 100644 --- a/crates/cellule-host/minion/reconciler_tests.rs +++ b/crates/cellule-host/minion/reconciler_tests.rs @@ -166,6 +166,10 @@ impl FleetObserver for Observer { target: target(n), generation: u64::from(n), incarnation: position(1).incarnation, + owner_fence: cellule_runtime::control::OwnerFence { + incarnation: position(1).incarnation, + epoch: 1, + }, code: Digest::from_bytes([9; 32]), schema: 1, role: CatalogRole::Sql, @@ -860,12 +864,7 @@ async fn draining_advertisement_or_pressure_alone_cannot_plan_busy_work() { #[tokio::test] async fn maintenance_planning_keeps_blob_role_and_missing_evidence_blocked() { - for state in [ - CellState::Blob, - CellState::RoleBlocked, - CellState::NoCost, - CellState::NoPosition, - ] { + for state in [CellState::Blob, CellState::RoleBlocked, CellState::NoCost] { let fixture = Fixture::new(false, false, false).await; *fixture.observer.cell_state.lock().unwrap() = state; fixture.request_maintenance(NOW + 20_000).await; @@ -1038,16 +1037,21 @@ async fn public_reconciler_drives_durable_phases_fresh_activation_cleanup_and_re assert_eq!(first.dispatched, 0); assert!(first.next_wake_at_ms < NOW + 100); assert_eq!(first.snapshot.head().reserved_restore_bytes(), 8192); + assert_eq!(fixture.observer.calls.load(Ordering::SeqCst), 1); fixture.stop().await; assert_eq!(fixture.step(&driver, 1).await.dispatched, 2); // prepare let released = fixture.step(&driver, 2).await; assert_eq!(released.dispatched, 2); assert_eq!(released.released, 2); assert_eq!(released.activated, 0); + assert_eq!(fixture.observer.calls.load(Ordering::SeqCst), 1); let activated = fixture.step(&driver, 3).await; assert_eq!(activated.inspected, 2); assert_eq!(activated.activated, 2); assert_eq!(activated.released, 0); + // Stop-new-moves suppresses planning, but each receiver effect still needs + // a fresh complete observation to decide whether its boot has closed. + assert_eq!(fixture.observer.calls.load(Ordering::SeqCst), 3); assert!( activated .snapshot @@ -1057,6 +1061,7 @@ async fn public_reconciler_drives_durable_phases_fresh_activation_cleanup_and_re .all(|a| a.phase() == AttemptPhase::Activated) ); assert_eq!(fixture.step(&driver, 4).await.dispatched, 2); // independent resource proof + assert_eq!(fixture.observer.calls.load(Ordering::SeqCst), 5); let retired = fixture.step(&fixture.driver(206), 5).await; assert_eq!(retired.retired, 2); assert_eq!(retired.inspected, 2); @@ -1089,7 +1094,7 @@ async fn public_reconciler_drives_durable_phases_fresh_activation_cleanup_and_re entry.completed_at_ms() ); } - assert_eq!(fixture.observer.calls.load(Ordering::SeqCst), 1); + assert_eq!(fixture.observer.calls.load(Ordering::SeqCst), 5); fixture.journal.close().await.unwrap(); } @@ -1774,3 +1779,20 @@ async fn unknown_enrollment_disables_counts_and_keeps_pressure_relief_available( fixture.journal.close().await.unwrap(); } } + +#[tokio::test] +async fn explicit_maintenance_can_prepare_from_retained_fence_without_borrowed_publisher_position() +{ + let fixture = Fixture::new(false, false, false).await; + *fixture.observer.cell_state.lock().unwrap() = CellState::NoPosition; + fixture.request_maintenance(NOW + 20_000).await; + let driver = fixture.driver(206); + fixture.step(&driver, 1).await; + let planned = fixture.step(&driver, 2).await; + assert_eq!(planned.allocated, 2); + assert_eq!(fixture.transport.released.load(Ordering::SeqCst), 0); + assert!(planned.snapshot.head().attempts().iter().all(|attempt| { + attempt.spec().source_epoch == 1 && attempt.spec().cost.disk_bytes == 8192 + })); + fixture.journal.close().await.unwrap(); +} diff --git a/crates/cellule-host/minion/scenario/adapters.rs b/crates/cellule-host/minion/scenario/adapters.rs index c5e806b9..f26526f7 100644 --- a/crates/cellule-host/minion/scenario/adapters.rs +++ b/crates/cellule-host/minion/scenario/adapters.rs @@ -1,4 +1,5 @@ use super::*; +use cellule_host::fleet::FleetReaderEvacuationVerifier; use cellule_host::fleet::*; use cellule_runtime::fleet::operations::*; @@ -6,6 +7,7 @@ pub(super) struct Cells { pub records: Arc>, pub local: usize, pub root: PathBuf, + pub receiver_directory: Option, } impl FleetCellProvider for Cells { fn cell_inputs<'a>( @@ -43,23 +45,102 @@ impl FleetCellProvider for Cells { )) }) } + fn receiver_recovery_inputs<'a>( + &'a self, + accepted: &'a AcceptedFleetAction, + control: &'a cellule_runtime::control::Control, + ) -> FleetAdapterFuture<'a, Option> { + Box::pin(async move { + let Some(directory) = &self.receiver_directory else { + return Ok(None); + }; + let FleetActionKind::Movement { + action: MovementAction::Activate, + attempt, + } = accepted.action().kind() + else { + return Err(invalid("receiver recovery requires routed activation")); + }; + if accepted.action().receiver_route().is_none() + || accepted.session() != session(self.local) + || control.cell != attempt.spec().target.cell_id() + || control.incarnation != attempt.spec().incarnation + { + return Err(invalid("receiver recovery binding differs")); + } + let Some(owner) = &control.owner else { + return Ok(None); + }; + let Some(takeover) = directory + .takeover_proof(owner.session, session(self.local), clock()?) + .await? + else { + return Ok(None); + }; + let record = self + .records + .get(&control.cell) + .ok_or_else(|| invalid("receiver recovery Cell absent"))?; + Ok(Some(FleetRecoveryInputs { + takeover, + manifests: cellule_runtime::recovery::manifest::RecoveryManifestStore::new( + record.authority.layout().clone(), + record.replica.limits(), + ), + })) + }) + } } -/// Fixed three-boot in-process transport. Trusted composition pins identities; +/// Finite managed-boot in-process transport. Trusted composition pins identities; /// production endpoints must provide equivalent authentication independently. pub(super) struct LocalFleet { pub nodes: Vec>, pub journal: Arc, pub boots: Vec, pub records: Arc>, + pub reader_verifier: Option, pub capture_sequence: std::sync::atomic::AtomicU64, pub lose_release_replies: bool, pub lost_release_replies: std::sync::atomic::AtomicUsize, + pub drop_closed_finalize_replies: std::sync::atomic::AtomicUsize, pub expired_receiver_cleanups: std::sync::atomic::AtomicUsize, } + +pub(super) struct LocalSnapshots { + nodes: Vec>, +} +impl LocalSnapshots { + pub(super) fn new(nodes: Vec>) -> Self { + Self { nodes } + } +} +impl FleetSnapshotTransport for LocalSnapshots { + fn capture<'a>( + &'a self, + request: &'a FleetSnapshotRequest, + deadline: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + Box::pin(async move { + if Instant::now() >= deadline { + return Err(Box::new(cellule_runtime::Error::Deadline) as JournalError); + } + let index = (0..self.nodes.len()) + .find(|index| { + node_id(*index) == request.node() && session(*index) == request.session() + }) + .ok_or_else(|| Box::new(cellule_runtime::Error::Fenced) as JournalError)?; + self.nodes[index] + .fleet_snapshot(request.clone()) + .await + .map_err(|error| Box::new(error) as JournalError) + }) + } +} + impl LocalFleet { fn endpoint(&self, physical: NodeId, boot: SessionId) -> JournalResult<&Arc> { - let index = (0..3) + let index = (0..self.nodes.len()) .find(|n| node_id(*n) == physical && session(*n) == boot) .ok_or_else(|| invalid("unrecognized example boot endpoint"))?; self.nodes @@ -74,18 +155,24 @@ impl FleetTransport for LocalFleet { _: Instant, ) -> FleetAdapterFuture<'a, Arc> { Box::pin(async move { - let FleetActionKind::Movement { - action: effect, - attempt, - } = action.kind() - else { - return Err(invalid("non-movement example action")); - }; - let spec = attempt.spec(); - let (node, boot) = if effect.is_source_release() { - (spec.source_node, spec.source) - } else { - (spec.destination_node, spec.destination) + let (node, boot, source_release) = match action.kind() { + FleetActionKind::Movement { + action: effect, + attempt, + } => { + let spec = attempt.spec(); + let (node, boot) = if effect.is_source_release() { + (spec.source_node, spec.source) + } else { + action + .receiver_endpoint() + .ok_or_else(|| invalid("receiver action has no endpoint"))? + }; + (node, boot, effect.is_source_release()) + } + FleetActionKind::Maintenance { operation, .. } => { + (operation.node(), operation.session(), false) + } }; let completion = self .endpoint(node, boot)? @@ -93,18 +180,26 @@ impl FleetTransport for LocalFleet { .await .map_err(|error| Box::new(error) as super::JournalError)?; let now = clock()?; - if *effect == MovementAction::Cancel - && attempt.phase() == AttemptPhase::CleaningReceiver - && attempt - .reservation() - .is_some_and(|r| r.expires_at_ms <= now) + let expired_receiver_cleanup = match action.kind() { + FleetActionKind::Movement { + action: MovementAction::Cancel, + attempt, + } => { + attempt.phase() == AttemptPhase::CleaningReceiver + && attempt + .reservation() + .is_some_and(|reservation| reservation.expires_at_ms <= now) + } + _ => false, + }; + if expired_receiver_cleanup && completion.committed && matches!(completion.outcome.outcome, FleetOutcome::ReceiverCleaned) { self.expired_receiver_cleanups .fetch_add(1, std::sync::atomic::Ordering::SeqCst); } - if self.lose_release_replies && effect.is_source_release() { + if self.lose_release_replies && source_release { if !completion.committed || !matches!(completion.outcome.outcome, FleetOutcome::Released(_)) { @@ -120,6 +215,216 @@ impl FleetTransport for LocalFleet { Ok(completion) }) } + fn settle_roles<'a>( + &'a self, + action: &'a FleetAction, + settlement: &'a FleetRoleSettlement, + _: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + Box::pin(async move { + let FleetActionKind::Maintenance { operation, .. } = action.kind() else { + return Err(invalid("role settlement requires a maintenance action")); + }; + self.endpoint(operation.node(), operation.session())? + .apply_fleet_role_settlement(action.clone(), settlement.clone(), clock()?) + .await + .map_err(|error| Box::new(error) as super::JournalError) + }) + } + fn settle_roles_after_process_closure<'a>( + &'a self, + action: &'a FleetAction, + settlement: &'a FleetRoleSettlement, + closure: &'a FleetFailedBootClosure, + _: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + Box::pin(async move { + settlement.validate_failed_boot_closure(action, closure)?; + let (node, session) = match action.kind() { + FleetActionKind::Maintenance { operation, .. } => { + (operation.node(), operation.session()) + } + FleetActionKind::Movement { .. } => { + return Err(invalid("closed-boot settlement requires maintenance")); + } + }; + let now = clock()?; + if now < action.issued_at_ms() { + return Err(invalid("closed-boot settlement clock regressed")); + } + let (accepted, previous) = match self + .journal + .accept_action(action, node, session, now) + .await? + { + FleetActionAcceptance::New(accepted) => (accepted, None), + FleetActionAcceptance::Existing { accepted, result } => (accepted, result), + }; + accepted.validate_replay(action, node, session)?; + let outcome = FleetActionOutcome { + scope: action.scope(), + action_key: action.key()?, + node, + session, + observed_at_ms: now, + outcome: FleetOutcome::RolesSettledAt { + inventory: settlement.inventory(), + head_revision: settlement.head_revision(), + registry: settlement.registry(), + }, + }; + accepted.validate_result(&outcome)?; + if let Some(previous) = previous { + let previous = *previous; + accepted.validate_result(&previous)?; + match &previous.outcome { + FleetOutcome::Unknown => {} + FleetOutcome::RolesSettledAt { + inventory, + head_revision, + registry, + } if *inventory == settlement.inventory() + && *head_revision == settlement.head_revision() + && *registry == settlement.registry() => + { + return Ok(Arc::new(FleetActionCompletion { + accepted, + outcome: previous, + committed: true, + execution_error: None, + journal_error: None, + })); + } + _ => return Err(invalid("closed-boot role-settlement result differs")), + } + } + self.journal + .publish_action_result(&accepted, &outcome) + .await?; + Ok(Arc::new(FleetActionCompletion { + accepted, + outcome, + committed: true, + execution_error: None, + journal_error: None, + })) + }) + } + fn finalize_after_process_closure<'a>( + &'a self, + action: &'a FleetAction, + closure: &'a FleetFailedBootClosure, + _: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + Box::pin(async move { + let operation = match action.kind() { + FleetActionKind::Maintenance { + action: MaintenanceAction::Finalize, + operation, + } => operation, + _ => return Err(invalid("closed-boot finalization requires Finalize")), + }; + let target = closure.boot().spec().target; + let mut evidence = operation + .drain_evidence() + .ok_or_else(|| invalid("closed-boot finalization lacks drain evidence"))?; + if operation.phase() != MaintenancePhase::Closing + || evidence.remaining_cells != 0 + || evidence.unresolved_attempts != 0 + || !evidence.relocated + || !evidence.readers_settled + || !evidence.followers_settled + || evidence.facilities_closed + || evidence.stopped + || evidence.withdrawn + || closure.snapshot().head().scope() != action.scope() + || closure.snapshot().head().maintenance() != Some(operation.as_ref()) + || closure.snapshot().head().revision() != action.journal_revision() + || closure.boot().status() != EnrollmentStatus::Retired + || !matches!(closure.boot().spec().role, EnrollmentRole::Node { .. }) + || target.node != operation.node() + || target.session != operation.session() + || closure.canonical().node() != operation.node() + || closure.canonical().session() != operation.session() + || closure.interval().1 > action.issued_at_ms() + { + return Err(invalid("closed-boot finalization evidence differs")); + } + let now = clock()?; + action.authorize_against(closure.snapshot().head(), now)?; + evidence.facilities_closed = true; + evidence.stopped = true; + evidence.withdrawn = true; + let (node, session) = (operation.node(), operation.session()); + let (accepted, previous) = match self + .journal + .accept_action(action, node, session, now) + .await? + { + FleetActionAcceptance::New(accepted) => (accepted, None), + FleetActionAcceptance::Existing { accepted, result } => (accepted, result), + }; + accepted.validate_replay(action, node, session)?; + let outcome = FleetActionOutcome { + scope: action.scope(), + action_key: action.key()?, + node, + session, + observed_at_ms: now, + outcome: FleetOutcome::Stopped(evidence), + }; + accepted.validate_result(&outcome)?; + if let Some(previous) = previous { + let previous = *previous; + accepted.validate_result(&previous)?; + match &previous.outcome { + FleetOutcome::Unknown => {} + FleetOutcome::Stopped(previous_evidence) if previous_evidence == &evidence => { + return Ok(Arc::new(FleetActionCompletion { + accepted, + outcome: previous, + committed: true, + execution_error: None, + journal_error: None, + })); + } + _ => return Err(invalid("closed-boot finalization result differs")), + } + } + self.journal + .publish_action_result(&accepted, &outcome) + .await?; + let mut remaining = self + .drop_closed_finalize_replies + .load(std::sync::atomic::Ordering::SeqCst); + let drop_reply = loop { + if remaining == 0 { + break false; + } + match self.drop_closed_finalize_replies.compare_exchange_weak( + remaining, + remaining - 1, + std::sync::atomic::Ordering::SeqCst, + std::sync::atomic::Ordering::SeqCst, + ) { + Ok(_) => break true, + Err(current) => remaining = current, + } + }; + if drop_reply { + return Err(invalid( + "injected loss after committed closed-boot finalization", + )); + } + Ok(Arc::new(FleetActionCompletion { + accepted, + outcome, + committed: true, + execution_error: None, + journal_error: None, + })) + }) + } fn inspect<'a>( &'a self, request: &'a FleetInspectionRequest, diff --git a/crates/cellule-host/minion/scenario/balance/mod.rs b/crates/cellule-host/minion/scenario/balance/mod.rs index 0f5842ec..3b85a13b 100644 --- a/crates/cellule-host/minion/scenario/balance/mod.rs +++ b/crates/cellule-host/minion/scenario/balance/mod.rs @@ -18,10 +18,12 @@ pub(super) async fn run( nodes: nodes.clone(), boots: boots.clone(), records: records.clone(), + reader_verifier: None, journal: journal.clone(), capture_sequence: std::sync::atomic::AtomicU64::new(0), lose_release_replies: false, lost_release_replies: std::sync::atomic::AtomicUsize::new(0), + drop_closed_finalize_replies: std::sync::atomic::AtomicUsize::new(0), expired_receiver_cleanups: std::sync::atomic::AtomicUsize::new(0), }); let driver = FleetReconciler::new( @@ -46,7 +48,12 @@ pub(super) async fn run( controller_epoch: 0, expired_receiver_cleanups: 0, blockers: Vec::new(), - final_counts: [0; 3], + final_counts: vec![0; 3], + maintenance_completed: false, + maintenance_boot_withdrawn: false, + receiver_process_closures: 0, + lost_activation_replies: 0, + routed_activation_replays: 0, }; let mut specs = HashMap::new(); let deadline = Instant::now() + Duration::from_secs(120); diff --git a/crates/cellule-host/minion/scenario/follower_maintenance/mod.rs b/crates/cellule-host/minion/scenario/follower_maintenance/mod.rs new file mode 100644 index 00000000..b57f3fd4 --- /dev/null +++ b/crates/cellule-host/minion/scenario/follower_maintenance/mod.rs @@ -0,0 +1,413 @@ +//! Live-follower maintenance through the existing supervisor and public driver. +use super::*; +use cellule_host::{ + NodeLogRotationPhase, + fleet::{ + FleetFollowerEvacuationJournal, FleetFollowerEvacuationPublication, + FleetFollowerEvacuationVerifier, FleetRoster, + }, +}; +use cellule_runtime::{ + fleet::operations::{ + EnrollmentRole, EnrollmentStatus, FollowerReplacementPolicy, MaintenancePhase, + }, + follower::FollowerStore, + node::NodeDirectory, +}; +use std::sync::atomic::{AtomicU64, AtomicUsize, Ordering}; + +mod provider; +mod readers; +mod service; +mod setup; +#[cfg(test)] +mod tests; + +struct Inputs { + records: Arc>, + acknowledged: Acknowledged, + readers: Option, +} + +pub(super) enum Roles { + Followers, + ReadersAndFollowers, +} +fn deadline() -> Instant { + Instant::now() + Duration::from_secs(10) +} + +pub(super) async fn run( + root: &tempfile::TempDir, + journal: Arc, + nodes: &mut Vec>, + boots: &mut Vec, + profile: FleetProfile, + roles: Roles, +) -> JournalResult { + // These role-enabled futures contain native peer activation and complete + // fleet capture. Keep each composition boundary off the caller stack. + let inputs = Box::pin(setup::initialize(root, &journal, nodes, boots, roles)).await?; + let version = journal.load_snapshot(scope()).await?.registry(); + let page = journal.enrollments_page(version, None, 128).await?; + let original = page + .entries() + .iter() + .find(|row| { + row.spec().target.node == node_id(1) + && matches!(row.spec().role, EnrollmentRole::Follower { log_epoch: 1 }) + }) + .ok_or_else(|| invalid("original managed follower is missing"))? + .clone(); + if original.status() != EnrollmentStatus::Established { + return Err(invalid("original follower is not established")); + } + let fleet = Arc::new(adapters::LocalFleet { + nodes: nodes.clone(), + boots: boots.clone(), + records: inputs.records.clone(), + journal: journal.clone(), + reader_verifier: inputs + .readers + .as_ref() + .map(|readers| readers.verifier.clone()), + capture_sequence: AtomicU64::new(0), + lose_release_replies: false, + lost_release_replies: AtomicUsize::new(0), + drop_closed_finalize_replies: AtomicUsize::new(0), + expired_receiver_cleanups: AtomicUsize::new(0), + }); + let claimant = SessionId::from_bytes([206; 16]); + let now = clock()?; + let snapshot = journal.load_snapshot(scope()).await?; + let claimed = journal + .claim_controller(scope(), snapshot.head().revision(), claimant, now) + .await?; + let epoch = claimed + .head() + .controller() + .ok_or_else(|| invalid("follower maintenance controller is absent"))? + .epoch; + let operation = MaintenanceOperation::new( + OperationId::from_bytes([85; 16])?, + Digest::from_bytes([86; 32]), + node_id(1), + session(1), + 2, + now, + now.checked_add(60_000) + .ok_or_else(|| invalid("follower maintenance deadline overflow"))?, + )?; + journal + .compare_exchange( + &claimed, + epoch, + now, + &JournalTransition::BeginMaintenance(operation.clone()), + ) + .await?; + let driver = FleetReconciler::new( + scope(), + claimant, + profile, + journal.clone(), + fleet.clone(), + fleet.clone(), + )?; + let wall_deadline = Instant::now() + Duration::from_secs(60); + let mut blockers = Vec::new(); + // The driver owns cordon and evacuation transitions. No fixture writes a + // ready-to-close row or an invented native role proof. + loop { + let report = Box::pin(driver.reconcile_once(clock, deadline())).await?; + checked_report(&report)?; + for blocker in report.blockers { + if !blockers.contains(&blocker) { + blockers.push(blocker); + } + } + let phase = report + .snapshot + .head() + .maintenance() + .ok_or_else(|| invalid("follower maintenance operation disappeared"))? + .phase(); + if phase == MaintenancePhase::Evacuating { + break; + } + if Instant::now() >= wall_deadline { + return Err(invalid("follower maintenance did not enter evacuation")); + } + } + boots[1] + .refresh_capacity(1, journal.as_ref(), deadline()) + .await?; + if nodes[1].state() == NodeState::Stopped || !nodes[1].is_management_ready() { + return Err(invalid("follower source stopped before replacement")); + } + // The missing replacement remains an explicit policy blocker and cannot + // retire an uncovered original follower or authorize native Finalize. + let blocked = Box::pin(driver.reconcile_once(clock, deadline())).await?; + checked_report(&blocked)?; + if blocked + .snapshot + .head() + .maintenance() + .map(MaintenanceOperation::phase) + != Some(MaintenancePhase::Evacuating) + || journal + .load_enrollment(scope(), original.spec().key()?) + .await? + .as_ref() + != Some(&original) + { + return Err(invalid("follower maintenance bypassed replacement policy")); + } + for blocker in blocked.blockers { + if !blockers.contains(&blocker) { + blockers.push(blocker); + } + } + if let Some(readers) = &inputs.readers { + readers.check_original(&journal, true).await?; + } + let store = nodes[1] + .try_owned_component::(cellule_host::FOLLOWER_STORE_COMPONENT)? + .ok_or_else(|| invalid("original follower store is missing"))?; + if store.retained_bytes() == 0 || nodes[1].stats().active_cells() != 0 { + return Err(invalid( + "follower maintenance fixture lacks a foreign retained tail", + )); + } + let donor = boots[1] + .directory + .load_if_live(session(1), clock()?) + .await? + .ok_or_else(|| invalid("original follower advertisement is missing"))?; + if donor.advertisement().capacity().follower_retained_bytes != store.retained_bytes() { + return Err(invalid( + "follower advertisement does not report its original retained bytes", + )); + } + // Enable the spare only after observing that the complete foreign obligation + // blocks finalization despite the donor having zero local writers. + boots[3] + .refresh_capacity(3, journal.as_ref(), deadline()) + .await?; + let selected = boots[0] + .directory + .select_log_members(session(0), 1, clock()?, 4) + .await?; + if selected != [node_id(2), node_id(3)] { + return Err(invalid( + "follower replacement selection includes the donor or lacks redundancy", + )); + } + // This canonical actor barrier covers every acknowledged frame before the + // existing supervisor fences each old member and closes the original epoch. + inputs.acknowledged.source.drain().await?; + let request = nodes[0].request_node_log_rotation(1)?; + let completion = tokio::time::timeout_at(deadline(), async { + loop { + let observed = request.observe()?; + if observed.phase() == NodeLogRotationPhase::Completed { + return observed + .completion() + .cloned() + .ok_or_else(|| invalid("follower rotation completion is missing")); + } + if observed.phase() == NodeLogRotationPhase::Interrupted { + return match observed.latest_failure() { + Some(source) => Err(Box::new(RotationFailure(source.clone())) as JournalError), + None => Err(invalid("follower rotation was interrupted")), + }; + } + tokio::time::sleep(Duration::from_millis(10)).await; + } + }) + .await??; + if completion.retirement().barrier().covered_through() == 0 + || completion.replacement_epoch() != 2 + { + return Err(invalid( + "follower rotation lacks original tail coverage or the newer epoch", + )); + } + let renewed = service::resume(root, nodes, boots, &inputs).await?; + let current = journal.load_snapshot(scope()).await?; + let policy = FollowerReplacementPolicy::new(scope(), 1, 2)?; + journal + .set_follower_replacement_policy(¤t, policy, clock()?) + .await?; + let current = journal.load_snapshot(scope()).await?; + let operation = current + .head() + .maintenance() + .ok_or_else(|| invalid("follower maintenance operation disappeared"))?; + let capture = nodes[0] + .follower_evacuation(&original, operation, 2, deadline()) + .await?; + if !Arc::ptr_eq(capture.rotation(), &completion) + || capture.retired_members().len() != 2 + || capture + .retired_members() + .iter() + .any(|row| row.status() != EnrollmentStatus::Retired) + || capture.replacements().len() != 2 + || capture + .authority() + .log() + .is_none_or(|log| log.epoch() != 2 || log.members() != [node_id(2), node_id(3)]) + { + return Err(invalid( + "follower evacuation lacks the exact original and replacement ensembles", + )); + } + let verifier = FleetFollowerEvacuationVerifier::new( + boots[0].directory.clone(), + Arc::new(adapters::LocalSnapshots::new(nodes.clone())), + ); + let publication = FleetFollowerEvacuationPublication::publish( + &capture, + policy, + journal.as_ref(), + &verifier, + deadline(), + clock, + ) + .await?; + let record = publication.record()?; + let reader_capture = if let Some(readers) = &inputs.readers { + // A follower proof alone cannot close a donor that also owns a reader. + let blocked = Box::pin(driver.reconcile_once(clock, deadline())).await?; + checked_report(&blocked)?; + readers + .require_blocked(&blocked, &journal, &nodes[1]) + .await?; + for blocker in blocked.blockers { + if !blockers.contains(&blocker) { + blockers.push(blocker); + } + } + Box::pin(readers.prepare_replacement(&journal, boots)).await?; + let blocked = Box::pin(driver.reconcile_once(clock, deadline())).await?; + checked_report(&blocked)?; + readers + .require_blocked(&blocked, &journal, &nodes[1]) + .await?; + Some(Box::pin(readers.complete_replacement(&journal, boots)).await?) + } else { + None + }; + let completed = loop { + let report = Box::pin(driver.reconcile_once(clock, deadline())).await?; + checked_report(&report)?; + for blocker in report.blockers { + if !blockers.contains(&blocker) { + blockers.push(blocker); + } + } + if report.snapshot.head().maintenance().is_some_and(|current| { + current.id() == operation.id() && current.phase() == MaintenancePhase::Completed + }) { + break report.snapshot; + } + if Instant::now() >= wall_deadline { + return Err(invalid( + "follower maintenance did not complete before its deadline", + )); + } + }; + let evidence = completed + .head() + .maintenance() + .and_then(MaintenanceOperation::drain_evidence) + .ok_or_else(|| invalid("follower maintenance completion evidence is missing"))?; + if !evidence.relocated + || !evidence.readers_settled + || !evidence.followers_settled + || evidence.remaining_cells != 0 + || evidence.unresolved_attempts != 0 + || !completed.head().attempts().is_empty() + || nodes[1].state() != NodeState::Stopped + || !boots[1].directory.is_withdrawn(session(1)).await? + { + return Err(invalid( + "follower maintenance completed without native shutdown and withdrawal", + )); + } + for retired in capture.retired_members() { + if journal + .load_enrollment(scope(), retired.spec().key()?) + .await? + .as_ref() + != Some(retired) + { + return Err(invalid("original follower retirement history changed")); + } + } + if record.retired() != capture.retired_members() { + return Err(invalid( + "published follower policy lost original retirement history", + )); + } + service::readback(root, &inputs, &renewed).await?; + if let Some((readers, capture)) = inputs.readers.as_ref().zip(reader_capture.as_ref()) { + readers.readback(&journal, capture).await?; + } + let roster = FleetRoster::collect(journal.as_ref(), &completed, deadline()).await?; + let final_counts = observation::complete_counts(&fleet, &roster, deadline()) + .await? + .ok_or_else(|| invalid("follower maintenance final observation is incomplete"))?; + if final_counts != [1, 0, 0, 0] { + return Err(invalid("follower maintenance changed writer ownership")); + } + Ok(ScenarioSummary { + released: 0, + activated: 0, + retired: 0, + receipt_checks: if inputs.readers.is_some() { 8 } else { 5 }, + max_inflight: 0, + max_restore_bytes: 0, + joined_nodes: 0, + boot_retirements: 0, + receiver_nodes: 2, + lost_release_replies: 0, + controller_epoch: completed + .head() + .controller() + .ok_or_else(|| invalid("controller is absent"))? + .epoch, + expired_receiver_cleanups: 0, + blockers, + final_counts, + maintenance_completed: true, + maintenance_boot_withdrawn: true, + receiver_process_closures: 0, + lost_activation_replies: 0, + routed_activation_replays: 0, + }) +} +fn checked_report(report: &cellule_host::fleet::FleetReconcileReport) -> JournalResult<()> { + if let Some(failure) = report.failures.first() { + return Err(Box::new(failure.error.clone())); + } + if let Some(failure) = &report.maintenance_failure { + return Err(Box::new(failure.clone())); + } + Ok(()) +} + +// Preserve the supervisor's shared source error when this finite command exits. +#[derive(Debug)] +struct RotationFailure(Arc); +impl std::fmt::Display for RotationFailure { + fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + self.0.fmt(formatter) + } +} +impl std::error::Error for RotationFailure { + fn source(&self) -> Option<&(dyn std::error::Error + 'static)> { + Some(self.0.as_ref()) + } +} diff --git a/crates/cellule-host/minion/scenario/follower_maintenance/provider.rs b/crates/cellule-host/minion/scenario/follower_maintenance/provider.rs new file mode 100644 index 00000000..dda2812f --- /dev/null +++ b/crates/cellule-host/minion/scenario/follower_maintenance/provider.rs @@ -0,0 +1,276 @@ +//! Finite live-owner embedding adapter; all effects use the canonical directory. +use super::*; +use bytes::Bytes; +use cellule_host::{FacilityResult, FleetNodeDurabilityProvider, FleetNodeLogRecruitment}; +use cellule_runtime::{ + Error, + follower::FollowerReceipt, + node::{ + VersionedNodeAdvertisement, + durability::NodeLogAuthority, + log::{NodeLogRetirementObservation, NodeLogRotationBarrier}, + log_transport::{ + AppendRequest, LocalFollowerTransport, NodeLogTransport, RetireRequest, SealRequest, + TailRequest, + }, + }, +}; +use futures_util::future::BoxFuture; + +pub(super) struct LiveFollowers { + directory: NodeDirectory, + locals: Vec<(NodeId, LocalFollowerTransport)>, + pub appends: AtomicUsize, +} +impl LiveFollowers { + pub fn new(directory: NodeDirectory, locals: Vec<(NodeId, LocalFollowerTransport)>) -> Self { + Self { + directory, + locals, + appends: AtomicUsize::new(0), + } + } + fn local(&self, member: NodeId) -> cellule_runtime::Result<&LocalFollowerTransport> { + self.locals + .iter() + .find(|(node, _)| *node == member) + .map(|(_, local)| local) + .ok_or(Error::Fenced) + } +} +impl NodeLogTransport for LiveFollowers { + fn append<'a>( + &'a self, + member: NodeId, + request: AppendRequest, + ) -> BoxFuture<'a, cellule_runtime::Result> { + Box::pin(async move { + self.directory + .authorize_log_append( + request.leader_session, + member, + request.log_epoch, + request.covered_through, + clock()?, + ) + .await?; + let receipt = self.local(member)?.append(member, request).await?; + self.appends.fetch_add(1, Ordering::AcqRel); + Ok(receipt) + }) + } + fn retire<'a>( + &'a self, + member: NodeId, + request: RetireRequest, + ) -> BoxFuture<'a, cellule_runtime::Result> { + Box::pin(async move { + self.directory + .authorize_log_retire( + request.leader_session, + member, + request.log_epoch, + request.covered_through, + clock()?, + ) + .await?; + self.local(member)?.retire(member, request).await + }) + } + fn seal<'a>( + &'a self, + _member: NodeId, + _request: SealRequest, + ) -> BoxFuture<'a, cellule_runtime::Result> { + Box::pin(async { + Err(Error::PeerAuthorization( + "live maintenance adapter has no recovery claimant", + )) + }) + } + fn tail<'a>( + &'a self, + _member: NodeId, + _request: TailRequest, + ) -> BoxFuture<'a, cellule_runtime::Result>> { + Box::pin(async { + Err(Error::PeerAuthorization( + "live maintenance adapter has no recovery claimant", + )) + }) + } +} + +#[derive(Default)] +struct History { + members: HashMap>, + closed: HashMap, +} +pub(super) struct Authority { + directory: NodeDirectory, + history: tokio::sync::Mutex, +} +impl Authority { + pub fn new(directory: NodeDirectory) -> Self { + Self { + directory, + history: tokio::sync::Mutex::new(History::default()), + } + } + async fn current( + &self, + epoch: u64, + history: &History, + ) -> cellule_runtime::Result { + let observed = self + .directory + .load_if_live(session(0), clock()?) + .await? + .ok_or(Error::Fenced)?; + if observed.advertisement().node() != node_id(0) + || observed.advertisement().log().is_none_or(|log| { + log.epoch() != epoch + || history + .members + .get(&epoch) + .is_none_or(|members| log.members() != members) + }) + { + return Err(Error::Fenced); + } + Ok(observed) + } +} +impl NodeLogAuthority for Authority { + fn activate<'a>(&'a self, epoch: u64) -> BoxFuture<'a, cellule_runtime::Result<()>> { + Box::pin(async move { + let history = self.history.lock().await; + let observed = self.current(epoch, &history).await?; + self.directory.activate_log(&observed, clock()?).await?; + Ok(()) + }) + } + fn advance_coverage<'a>( + &'a self, + epoch: u64, + through: u64, + ) -> BoxFuture<'a, cellule_runtime::Result<()>> { + Box::pin(async move { + let history = self.history.lock().await; + let observed = self.current(epoch, &history).await?; + self.directory + .advance_log_coverage(&observed, through, clock()?) + .await?; + Ok(()) + }) + } + fn close<'a>( + &'a self, + retirement: &'a NodeLogRetirementObservation, + ) -> BoxFuture<'a, cellule_runtime::Result<()>> { + Box::pin(async move { + retirement.confirmed()?; + let mut history = self.history.lock().await; + let barrier = retirement.barrier(); + if let Some(original) = history.closed.get(&barrier.log_epoch()) { + return if original == barrier { + Ok(()) + } else { + Err(Error::Fenced) + }; + } + let observed = self.current(barrier.log_epoch(), &history).await?; + self.directory + .close_log(&observed, barrier, clock()?) + .await?; + // Retain the exact checked barrier before returning the close reply. + // Mere absence of the log cannot prove a particular original close. + history.closed.insert(barrier.log_epoch(), barrier.clone()); + Ok(()) + }) + } +} + +pub(super) struct Provider { + directory: NodeDirectory, + transport: Arc, + authority: Arc, + lease: NodeLeaseGuard, + prepared: AtomicU64, +} +impl Provider { + pub fn new( + directory: NodeDirectory, + transport: Arc, + lease: NodeLeaseGuard, + ) -> Self { + Self { + authority: Arc::new(Authority::new(directory.clone())), + directory, + transport, + lease, + prepared: AtomicU64::new(0), + } + } +} +impl FleetNodeDurabilityProvider for Provider { + fn rotation_required( + self: Arc, + _live_node_limit: usize, + ) -> std::pin::Pin> + Send>> { + Box::pin(async { Ok(false) }) + } + fn prepare( + self: Arc, + _limits: Limits, + bytes: u64, + live: usize, + ) -> std::pin::Pin< + Box< + dyn std::future::Future>> + + Send, + >, + > { + Box::pin(async move { + // This finite command owns exactly two epochs. Failed preparation + // still consumes its epoch; native attempts are never restamped. + let index = match self + .prepared + .try_update(Ordering::AcqRel, Ordering::Acquire, |old| { + (old < 2).then_some(old + 1) + }) { + Ok(index) => index, + Err(_) => return Ok(None), + }; + let now = clock()?; + let source = self + .directory + .load_if_live(session(0), now) + .await? + .ok_or(Error::Fenced)?; + let prepared = self + .directory + .prepare_log_enrollment(&source, index + 1, bytes, live, now) + .await? + .ok_or(Error::Fenced)?; + let attempt = self + .directory + .prepare_log_enrollment_attempt(&prepared, now) + .await?; + self.authority + .history + .lock() + .await + .members + .insert(prepared.log().epoch(), prepared.log().members().to_vec()); + Ok(Some(FleetNodeLogRecruitment::new( + self.directory.clone(), + attempt, + self.transport.clone(), + self.authority.clone(), + self.lease.clone(), + Default::default(), + )?)) + }) + } +} diff --git a/crates/cellule-host/minion/scenario/follower_maintenance/readers.rs b/crates/cellule-host/minion/scenario/follower_maintenance/readers.rs new file mode 100644 index 00000000..fb24d1a5 --- /dev/null +++ b/crates/cellule-host/minion/scenario/follower_maintenance/readers.rs @@ -0,0 +1,255 @@ +//! Reader obligations on the same physical donor as the foreign follower tail. +use super::*; +use cellule_host::{ + fleet::{ + FleetReaderEvacuationPublication, FleetReaderEvacuationVerifier, FleetReconcileReport, + }, + read_replicas::{ReadReplicaManager, ReaderEvacuation}, +}; +use cellule_runtime::{ + client::{CellDescription, CellReadReplica}, + fleet::operations::EnrollmentRecord, + peer::{PeerReplicaResolver, ReplicaPeerClient}, +}; + +pub(super) struct Readers { + managers: Vec, + target: CellTarget, + description: CellDescription, + peer: ReplicaPeerClient, + pub(super) verifier: FleetReaderEvacuationVerifier, + original: EnrollmentRecord, + reader: CellReadReplica, +} + +impl Readers { + pub(super) async fn initialize( + journal: &SqliteJournal, + managers: Vec, + target: CellTarget, + description: CellDescription, + peer: ReplicaPeerClient, + verifier: FleetReaderEvacuationVerifier, + boots: &[startup::BootOwner], + ) -> JournalResult { + // Both original follower receivers advertise reader capacity. Requiring + // two readers makes the eventual placement exclude the cordoned donor. + managers[1] + .set_target(&target, 0, 2) + .await? + .ok_or_else(|| invalid("combined maintenance reader policy is absent"))?; + let ad = boots[1].refresh_capacity(1, journal, deadline()).await?; + peer.activate(&target, &boots[0].directory, ad, description) + .await?; + let reader = managers[1].resolve(target.clone()).await?; + let completion = managers[1] + .enrollment_completion(target.cell_id()) + .await? + .ok_or_else(|| invalid("combined maintenance reader acceptance is absent"))?; + let original = journal + .load_enrollment(scope(), completion.spec.key()?) + .await? + .ok_or_else(|| invalid("combined maintenance reader enrollment is absent"))?; + let inputs = Self { + managers, + target, + description, + peer, + verifier, + original, + reader, + }; + inputs.check_original(journal, true).await?; + Ok(inputs) + } + + pub(super) async fn check_original( + &self, + journal: &SqliteJournal, + read: bool, + ) -> JournalResult<()> { + if self.original.status() != EnrollmentStatus::Established + || journal + .load_enrollment(scope(), self.original.spec().key()?) + .await? + .as_ref() + != Some(&self.original) + || self.reader.lifecycle_observation().await.admission_closed() + { + return Err(invalid( + "combined maintenance prematurely retired its reader", + )); + } + if read + && self + .reader + .query::(None, 0) + .await? + .output + != 29 + { + return Err(invalid( + "combined maintenance original reader lost acknowledged state", + )); + } + Ok(()) + } + + pub(super) async fn require_blocked( + &self, + report: &FleetReconcileReport, + journal: &SqliteJournal, + donor: &CellNode, + ) -> JournalResult<()> { + if report + .snapshot + .head() + .maintenance() + .map(MaintenanceOperation::phase) + != Some(MaintenancePhase::Evacuating) + || donor.state() == NodeState::Stopped + || !donor.is_management_ready() + || report.blockers.is_empty() + { + return Err(invalid( + "combined maintenance finalized with an unsettled reader", + )); + } + self.check_original(journal, false).await + } + + pub(super) async fn prepare_replacement( + &self, + journal: &SqliteJournal, + boots: &[startup::BootOwner], + ) -> JournalResult<()> { + tokio::time::timeout_at(deadline(), async { + loop { + let selected = boots[0] + .directory + .select_readers( + self.target.cell_id(), + session(0), + self.description.code, + 2, + clock()?, + 128, + ) + .await?; + if selected.len() == 2 + && selected + .iter() + .all(|ad| ad.session() == session(2) || ad.session() == session(3)) + { + return Ok::<_, JournalError>(()); + } + tokio::time::sleep(Duration::from_millis(10)).await; + } + }) + .await??; + let replacement = boots[2].refresh_capacity(2, journal, deadline()).await?; + self.peer + .activate( + &self.target, + &boots[0].directory, + replacement, + self.description, + ) + .await?; + let current = journal.load_snapshot(scope()).await?; + let operation = current + .head() + .maintenance() + .ok_or_else(|| invalid("combined maintenance operation disappeared"))?; + // Node 3 is live and holds the replacement follower, but has no ready + // reader. Native evacuation must refuse before closing the old view. + match self.managers[1] + .evacuate(&self.original, operation, &self.peer, deadline()) + .await + { + Err(cellule_runtime::Error::Control( + "reader replacement lacks Established enrollment", + )) => {} + Err(error) => return Err(Box::new(error)), + Ok(_) => { + return Err(invalid( + "combined maintenance accepted an absent replacement reader", + )); + } + } + self.check_original(journal, false).await + } + + pub(super) async fn complete_replacement( + &self, + journal: &SqliteJournal, + boots: &[startup::BootOwner], + ) -> JournalResult { + let replacement = boots[3].refresh_capacity(3, journal, deadline()).await?; + self.peer + .activate( + &self.target, + &boots[0].directory, + replacement, + self.description, + ) + .await?; + let current = journal.load_snapshot(scope()).await?; + let operation = current + .head() + .maintenance() + .ok_or_else(|| invalid("combined maintenance operation disappeared"))?; + let capture = self.managers[1] + .evacuate(&self.original, operation, &self.peer, deadline()) + .await?; + let publication = FleetReaderEvacuationPublication::publish( + &capture, + journal, + &self.verifier, + deadline(), + clock, + ) + .await?; + let record = publication.record()?; + if record.retired() != capture.retired() || capture.replacements().len() != 2 { + return Err(invalid( + "combined maintenance reader publication lost native history", + )); + } + Ok(capture) + } + + pub(super) async fn readback( + &self, + journal: &SqliteJournal, + capture: &ReaderEvacuation, + ) -> JournalResult<()> { + let retired = journal + .load_enrollment(scope(), self.original.spec().key()?) + .await? + .ok_or_else(|| invalid("combined maintenance reader history disappeared"))?; + let lifecycle = self.reader.lifecycle_observation().await; + if retired != *capture.retired() + || retired.status() != EnrollmentStatus::Retired + || !lifecycle.locally_joined() + { + return Err(invalid( + "combined maintenance original reader was not joined and retired", + )); + } + for index in [2, 3] { + let reader = self.managers[index].resolve(self.target.clone()).await?; + if reader + .query::(Some(capture.minimum()), 0) + .await? + .output + != 29 + { + return Err(invalid( + "combined maintenance replacement reader lost acknowledged state", + )); + } + } + Ok(()) + } +} diff --git a/crates/cellule-host/minion/scenario/follower_maintenance/service.rs b/crates/cellule-host/minion/scenario/follower_maintenance/service.rs new file mode 100644 index 00000000..cd041dff --- /dev/null +++ b/crates/cellule-host/minion/scenario/follower_maintenance/service.rs @@ -0,0 +1,178 @@ +//! Resume through canonical acquisition and check exact durable command evidence. +use super::*; + +pub(super) async fn resume( + root: &tempfile::TempDir, + nodes: &[Arc], + boots: &[startup::BootOwner], + inputs: &Inputs, +) -> JournalResult { + let cell = inputs + .records + .values() + .next() + .ok_or_else(|| invalid("follower maintenance Cell is absent"))?; + let idle = cell + .authority + .load(cell.target.cell_id()) + .await? + .ok_or_else(|| invalid("drained writer authority is absent"))?; + let resumed = nodes[0] + .runtime() + .acquire_idle_restored( + cell.catalog.clone(), + cell.replica.clone(), + cell.authority.clone(), + idle, + root.path().join("resumed-writer.sqlite"), + owner(0), + ) + .await?; + if resumed + .resolve( + inputs.acknowledged.identity, + inputs.acknowledged.digest, + clock()?, + 64, + ) + .await? + != Resolution::Committed(inputs.acknowledged.outcome.clone()) + || resumed + .query(64, 64, |tx| { + let value: i64 = tx.query_row("SELECT value FROM counter", [], |row| row.get(0))?; + Ok(value.to_be_bytes().to_vec()) + }) + .await? + != inputs.acknowledged.value.to_be_bytes() + { + return Err(invalid( + "resumed writer lost the original acknowledged command", + )); + } + let now = clock()?; + let renewed_identity = MutationIdentity { + request_id: RequestId::from_bytes([2; 16]), + issued_at_ms: now, + expires_at_ms: now + .checked_add(120_000) + .ok_or_else(|| invalid("renewed receipt deadline overflow"))?, + }; + let renewed_digest = Digest::from_bytes([2; 32]); + let renewed_outcome = resumed + .execute(renewed_identity, renewed_digest, now, 64, 64, |tx| { + tx.execute_batch("UPDATE counter SET value = 29")?; + Ok(HandlerOutcome::Success(29i64.to_be_bytes().to_vec())) + }) + .await?; + if resumed + .resolve(renewed_identity, renewed_digest, clock()?, 64) + .await? + != Resolution::Committed(renewed_outcome.clone()) + { + return Err(invalid( + "replacement ensemble lost its acknowledged command", + )); + } + // Object proof may win before authoritative follower activation. A real + // installed ensemble is checked independently by follower evacuation below. + let leader = boots[0] + .directory + .load_if_live(session(0), clock()?) + .await? + .ok_or_else(|| invalid("replacement leader is absent"))?; + if leader + .advertisement() + .log() + .is_none_or(|log| log.epoch() != 2 || log.members() != [node_id(2), node_id(3)]) + { + return Err(invalid( + "replacement ensemble changed during resumed service", + )); + } + Ok(renewed_outcome) +} + +pub(super) async fn readback( + root: &tempfile::TempDir, + inputs: &Inputs, + renewed: &StoredOutcome, +) -> JournalResult<()> { + let record = inputs + .records + .values() + .next() + .ok_or_else(|| invalid("follower maintenance Cell is absent"))?; + // Restore only a canonical root covering both acknowledgements. Follower + // proof can precede asynchronous object publication, so wait for real CAS. + let (control, published) = tokio::time::timeout_at(deadline(), async { + loop { + let control = record + .authority + .load(record.target.cell_id()) + .await? + .ok_or_else(|| invalid("follower maintenance authority is absent"))?; + let published = control + .value() + .ltx_root() + .ok_or_else(|| invalid("follower maintenance canonical root is absent"))?; + if published.commit_sequence >= renewed.commit_sequence() { + return Ok::<_, JournalError>((control, published)); + } + tokio::time::sleep(Duration::from_millis(10)).await; + } + }) + .await??; + if control.value().owner.as_ref() != Some(&owner(0)) + || control.value().incarnation != record.incarnation + || published.commit_sequence < inputs.acknowledged.outcome.commit_sequence() + { + return Err(invalid( + "follower maintenance changed the canonical writer or lost its receipt prefix", + )); + } + let destination = root.path().join("follower-maintenance-readback.sqlite"); + if record + .replica + .open_root(&published) + .await? + .restore(&destination) + .await? + != published.position + { + return Err(invalid("follower maintenance restore position differs")); + } + let acknowledged = &inputs.acknowledged; + let request = acknowledged.identity.request_id; + let expected = acknowledged.value; + let (value, result, digest, expires, stored) = tokio::task::spawn_blocking(move || { + let database = rusqlite::Connection::open_with_flags( + destination, + rusqlite::OpenFlags::SQLITE_OPEN_READ_ONLY, + )?; + let value = + database.query_row("SELECT value FROM counter", [], |row| row.get::<_, i64>(0))?; + let (result, digest, expires, outcome, sequence) = database.query_row( + "SELECT result, operation_digest, expires_at_ms, outcome, commit_sequence FROM sys_requests WHERE request_id=?1", + [request.as_bytes().as_slice()], + |row| Ok((row.get::<_, Vec>(0)?, row.get::<_, Vec>(1)?, row.get::<_, i64>(2)?, row.get::<_, i64>(3)?, row.get::<_, u64>(4)?)), + )?; + let stored = match outcome { + 1 => StoredOutcome::Success { result: result.clone(), commit_sequence: sequence }, + 2 => StoredOutcome::Rejected { result: result.clone(), commit_sequence: sequence }, + _ => return Err(rusqlite::Error::InvalidQuery), + }; + Ok::<_, rusqlite::Error>((value, result, digest, expires, stored)) + }) + .await??; + if value != expected + || result != expected.to_be_bytes() + || digest != acknowledged.digest.as_bytes() + || expires != acknowledged.identity.expires_at_ms + || stored != acknowledged.outcome + { + return Err(invalid( + "follower maintenance lost acknowledged value or original stored outcome", + )); + } + Ok(()) +} diff --git a/crates/cellule-host/minion/scenario/follower_maintenance/setup.rs b/crates/cellule-host/minion/scenario/follower_maintenance/setup.rs new file mode 100644 index 00000000..81ba2fed --- /dev/null +++ b/crates/cellule-host/minion/scenario/follower_maintenance/setup.rs @@ -0,0 +1,293 @@ +//! Register every partially constructed owner before fallible startup work. +use super::*; +use cellule_host::NodeDurabilitySupervisorConfig; +use cellule_runtime::node::log_transport::LocalFollowerTransport; +use cellule_runtime::{ + client::CellDescription, + peer::{PeerPrincipal, PeerSigner, ReplicaPeerClient}, +}; +use ed25519_dalek::SigningKey; + +pub(super) async fn initialize( + root: &tempfile::TempDir, + journal: &Arc, + nodes: &mut Vec>, + boots: &mut Vec, + roles: Roles, +) -> JournalResult { + let app = application::compile()?; + let code = *app + .registry() + .module_digests() + .first() + .ok_or_else(|| invalid("follower maintenance module is absent"))?; + let layout = CellStorageLayout::new( + Store::new(Arc::new(InMemory::new())), + ObjectPath::from("follower-maintenance-cells"), + [3; 16], + ); + let directory = NodeDirectory::new( + layout.clone(), + scope().fleet, + Digest::from_bytes([31; 32]), + app.registry().release_digest(), + ); + let limits = Limits { + max_database_bytes: 64 << 20, + max_capture_bytes: 16 << 20, + ..Limits::default() + }; + let target = CellTarget::new( + TenantId::from_bytes([1; 16]), + scope().application, + application::NAMESPACE, + &[1], + )?; + let incarnation = IncarnationId::from_bytes([1; 16]); + let catalog = CellCatalog::new(layout.clone(), target.tenant()) + .provision(CatalogEntry::new(&target, CatalogRole::Sql, code, 1)?) + .await?; + let authority = CellAuthority::new(layout.clone()); + let initial = authority + .create_initial(&catalog, incarnation, owner(0)) + .await?; + let replica = CellReplica::new( + layout.clone(), + *target.cell_id().as_bytes(), + *incarnation.as_bytes(), + limits, + )?; + let records = Arc::new(HashMap::from([( + target.cell_id(), + Record { + target: target.clone(), + incarnation, + catalog: catalog.clone(), + replica: replica.clone(), + authority: authority.clone(), + }, + )])); + let mut managers = Vec::new(); + for index in 0..4 { + let intent = journal + .register_initial_intent(&NodeIntent::initial( + scope(), + node_id(index), + session(index), + )?) + .await?; + let mut builder = CellNodeBuilder::new(app.clone()) + .with_runtime( + SqlWorkerPool::new(2, 8)?.with_native_memory_limit(128 << 20)?, + 16 << 20, + ) + .with_replica_host(Host::default().with_local_disk_budget(DiskBudget::new(8 << 30))) + .with_session(session(index)) + .with_fleet_startup_intent(intent.clone()); + if index != 0 { + builder = builder.with_follower_store( + root.path().join(format!("follower-{index}")), + limits, + DiskBudget::new(1 << 30), + ); + } + let node = Arc::new(builder.build()?); + nodes.push(node.clone()); + node.install_task_group(CancellationToken::new(), CancellationToken::new())?; + if matches!(roles, Roles::ReadersAndFollowers) { + managers.push(node.install_read_replicas( + layout.clone(), + directory.clone(), + root.path().join(format!("readers-{index}")), + limits, + )?); + node.install_fleet_reader_enrollment(scope(), node_id(index), journal.clone())?; + } + node.install_fleet_actions( + scope(), + node_id(index), + journal.clone(), + Arc::new(adapters::Cells { + records: records.clone(), + local: index, + root: root.path().into(), + receiver_directory: None, + }), + )?; + let ad = startup::advertisement(index, &node, &intent).await?; + let spec = startup::spec(&intent)?; + boots.push(startup::BootOwner { + node: node.clone(), + directory: directory.clone(), + spec: spec.clone(), + advertisement: ad.clone(), + guard: None, + }); + let original = startup::enroll(journal, &directory, &spec, ad.clone(), clock()?).await?; + let guard = NodeLeaseGuard::new(clock()?, ad.expires_at_ms())?; + node.install_node_lease_for_startup(guard.clone())?; + boots[index].guard = Some(guard); + node.confirm_fleet_startup(journal.as_ref(), spec.key()?) + .await?; + let observed = directory + .load(session(index), clock()?) + .await? + .ok_or_else(|| invalid("original follower maintenance boot is absent"))?; + node.install_fleet_boot_withdrawal(directory.clone(), observed, original, journal.clone())?; + } + let mut locals = Vec::new(); + for (index, node) in nodes.iter().enumerate().skip(1) { + let store = node + .try_owned_component::(cellule_host::FOLLOWER_STORE_COMPONENT)? + .ok_or_else(|| invalid("follower maintenance store is absent"))?; + locals.push(( + node_id(index), + LocalFollowerTransport::new(node_id(index), (*store).clone()), + )); + } + let transport = Arc::new(provider::LiveFollowers::new(directory.clone(), locals)); + let lease = boots[0] + .guard + .as_ref() + .ok_or_else(|| invalid("leader lease is absent"))? + .clone(); + let provider = Arc::new(provider::Provider::new( + directory.clone(), + transport.clone(), + lease, + )); + nodes[0].install_fleet_node_durability_provider( + scope(), + node_id(0), + journal.clone(), + provider, + NodeDurabilitySupervisorConfig::new( + scope().application, + limits, + 1, + 4, + Duration::from_millis(10), + Duration::from_secs(60), + u64::MAX, + )?, + )?; + // Register all physical boots before permitting the existing producer. + let version = journal.load_snapshot(scope()).await?.registry(); + journal.bootstrap_registry(version).await?; + for index in 1..4 { + nodes[index].start()?; + if index < 3 { + boots[index] + .refresh_capacity(index, journal.as_ref(), deadline()) + .await?; + } + } + nodes[0].start()?; + tokio::time::timeout_at(deadline(), async { + while nodes[0].runtime().node_durability().is_none() { + tokio::time::sleep(Duration::from_millis(10)).await; + } + }) + .await?; + boots[0] + .refresh_capacity(0, journal.as_ref(), deadline()) + .await?; + let handle = nodes[0] + .runtime() + .bootstrap( + catalog, + replica, + authority, + initial, + root.path().join("source.sqlite"), + |tx| { + tx.execute_batch( + "CREATE TABLE counter(value INTEGER); INSERT INTO counter VALUES (17)", + )?; + Ok(()) + }, + ) + .await?; + let now = clock()?; + let identity = MutationIdentity { + request_id: RequestId::from_bytes([1; 16]), + issued_at_ms: now, + expires_at_ms: now + .checked_add(120_000) + .ok_or_else(|| invalid("follower maintenance receipt deadline overflow"))?, + }; + let digest = Digest::from_bytes([1; 32]); + let outcome = handle + .execute(identity, digest, now, 64, 64, |tx| { + tx.execute_batch("UPDATE counter SET value = 29")?; + Ok(HandlerOutcome::Success(29i64.to_be_bytes().to_vec())) + }) + .await?; + let acknowledged = Acknowledged { + identity, + digest, + outcome, + value: 29, + source: handle, + }; + tokio::time::timeout_at(deadline(), async { + while transport.appends.load(Ordering::Acquire) < 2 { + tokio::time::sleep(Duration::from_millis(10)).await; + } + }) + .await?; + let readers = if matches!(roles, Roles::ReadersAndFollowers) { + let peer = ReplicaPeerClient::new( + app.registry(), + Arc::new(PeerSigner::new( + session(0), + app.registry().release_digest(), + SigningKey::from_bytes(&[1; 32]), + )), + PeerPrincipal { + issuer: "managed-owner".into(), + subject: "live-owner".into(), + actions: vec!["replica-maintenance".into()], + }, + Arc::new(native_peers::NativePeers::new( + nodes, + &managers, + &layout, + directory.clone(), + )), + ); + let description = CellDescription { + cell: target.cell_id(), + incarnation, + code, + schema: 1, + }; + let verifier = cellule_host::fleet::FleetReaderEvacuationVerifier::new( + directory, + CellAuthority::new(layout.clone()), + cellule_runtime::read_policy::ReadPolicyStore::new(layout), + peer.clone(), + ); + Some( + Box::pin(readers::Readers::initialize( + journal, + managers, + target, + description, + peer, + verifier, + boots, + )) + .await?, + ) + } else { + None + }; + let version = journal.load_snapshot(scope()).await?.registry(); + journal.set_scheduling(version, true).await?; + Ok(Inputs { + records, + acknowledged, + readers, + }) +} diff --git a/crates/cellule-host/minion/scenario/follower_maintenance/tests.rs b/crates/cellule-host/minion/scenario/follower_maintenance/tests.rs new file mode 100644 index 00000000..9cf6932d --- /dev/null +++ b/crates/cellule-host/minion/scenario/follower_maintenance/tests.rs @@ -0,0 +1,38 @@ +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn executable_follower_maintenance_covers_old_tail_replaces_both_members_and_joins_every_owner() + { + let summary = super::super::maintenance_follower().await.unwrap(); + assert!(summary.maintenance_completed); + assert!(summary.maintenance_boot_withdrawn); + assert_eq!(summary.final_counts, [1, 0, 0, 0]); + assert_eq!(summary.receipt_checks, 5); + assert_eq!(summary.joined_nodes, 4); + assert_eq!(summary.boot_retirements, 4); + assert_eq!(summary.receiver_nodes, 2); + assert_eq!( + (summary.released, summary.activated, summary.retired), + (0, 0, 0) + ); + assert_eq!(summary.max_inflight, 0); + assert_eq!(summary.max_restore_bytes, 0); + assert!(!summary.blockers.is_empty()); +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn executable_combined_maintenance_requires_both_role_policies_and_every_ready_replacement() { + let summary = super::super::maintenance_roles().await.unwrap(); + assert!(summary.maintenance_completed); + assert!(summary.maintenance_boot_withdrawn); + assert_eq!(summary.final_counts, [1, 0, 0, 0]); + assert_eq!(summary.receipt_checks, 8); + assert_eq!(summary.joined_nodes, 4); + assert_eq!(summary.boot_retirements, 4); + assert_eq!(summary.receiver_nodes, 2); + assert_eq!( + (summary.released, summary.activated, summary.retired), + (0, 0, 0) + ); + assert_eq!(summary.max_inflight, 0); + assert_eq!(summary.max_restore_bytes, 0); + assert!(!summary.blockers.is_empty()); +} diff --git a/crates/cellule-host/minion/scenario/follower_tests/coverage.rs b/crates/cellule-host/minion/scenario/follower_tests/coverage.rs index b5a735e4..79341540 100644 --- a/crates/cellule-host/minion/scenario/follower_tests/coverage.rs +++ b/crates/cellule-host/minion/scenario/follower_tests/coverage.rs @@ -10,11 +10,20 @@ pub(super) async fn captures( ) -> (Vec, Vec, u64) { let mut sequence = 0; let mut native = Vec::new(); - let mut foreign = Vec::new(); for index in 0..fixture.nodes.len() { native.push(aggregate::collect(fixture, roster, index, &mut sequence).await); - foreign.push(aggregate::references(fixture, roster, index).await); } + let members = (0..fixture.nodes.len()).map(node_id).collect::>(); + let foreign = FleetFollowerReferences::collect_all( + &fixture.native.directory, + roster, + &members, + 1, + Instant::now() + Duration::from_secs(5), + clock, + ) + .await + .unwrap(); (native, foreign, sequence) } @@ -41,18 +50,16 @@ pub(super) async fn foreign_rechecks( roster: &FleetRoster, foreign: &mut [FleetFollowerReferences], ) { - for references in foreign { - references - .recheck( - &fixture.native.directory, - roster, - 1, - Instant::now() + Duration::from_secs(5), - clock, - ) - .await - .unwrap(); - } + FleetFollowerReferences::recheck_all( + foreign, + &fixture.native.directory, + roster, + 1, + Instant::now() + Duration::from_secs(5), + clock, + ) + .await + .unwrap(); } pub(super) fn check( diff --git a/crates/cellule-host/minion/scenario/follower_tests/fixture/managed.rs b/crates/cellule-host/minion/scenario/follower_tests/fixture/managed.rs index fef0dd74..4319d8d1 100644 --- a/crates/cellule-host/minion/scenario/follower_tests/fixture/managed.rs +++ b/crates/cellule-host/minion/scenario/follower_tests/fixture/managed.rs @@ -8,6 +8,7 @@ pub(crate) struct ManagedFixture { pub nodes: Vec>, pub boots: Vec, pub handle: CellHandle, + pub records: Arc>, } impl ManagedFixture { @@ -118,6 +119,7 @@ impl ManagedFixture { node_id(index), journal.clone(), Arc::new(adapters::Cells { + receiver_directory: None, records: records.clone(), local: index, root: root.path().into(), @@ -323,6 +325,7 @@ impl ManagedFixture { nodes, boots, handle, + records, } } diff --git a/crates/cellule-host/minion/scenario/follower_tests/mod.rs b/crates/cellule-host/minion/scenario/follower_tests/mod.rs index ea87cd97..84aad55c 100644 --- a/crates/cellule-host/minion/scenario/follower_tests/mod.rs +++ b/crates/cellule-host/minion/scenario/follower_tests/mod.rs @@ -35,6 +35,7 @@ mod inventory; mod nonexecution; mod observation; mod persisted; +mod reference_batches; mod tests; async fn captured(reply: tokio::sync::oneshot::Receiver<()>) { diff --git a/crates/cellule-host/minion/scenario/follower_tests/persisted/cancelled_settlement.rs b/crates/cellule-host/minion/scenario/follower_tests/persisted/cancelled_settlement.rs new file mode 100644 index 00000000..bb02e578 --- /dev/null +++ b/crates/cellule-host/minion/scenario/follower_tests/persisted/cancelled_settlement.rs @@ -0,0 +1,244 @@ +//! The original native owner survives a cancelled controller publication waiter. +use super::*; +use crate::journal::ResultWriteBoundary; +use cellule_host::fleet::{ + FleetActionAcceptance, FleetActionCompletion, FleetActionJournal, FleetActionWorkState, + FleetAdapterFuture, FleetReconciler, FleetRoleSettlement, FleetTransport, +}; +use cellule_runtime::fleet::operations::{ + FleetAction, FleetInspectionObservation, FleetInspectionRequest, FleetOutcome, MaintenancePhase, +}; +use std::future::Future; +use std::task::Poll; + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn cancelled_controller_waiter_before_role_result_write_keeps_original_owner() { + run(ResultWriteBoundary::BeforeCommit).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn cancelled_controller_waiter_after_role_result_commit_keeps_original_owner() { + run(ResultWriteBoundary::AfterCommit).await; +} + +async fn run(boundary: ResultWriteBoundary) { + let (fixture, capture, policy) = setup().await; + let publication = publish(&fixture, &capture, policy).await; + let fleet = Arc::new(crate::scenario::adapters::LocalFleet { + nodes: fixture.nodes.clone(), + journal: fixture.native.journal.clone(), + boots: fixture.boots.clone(), + records: fixture.records.clone(), + reader_verifier: None, + capture_sequence: std::sync::atomic::AtomicU64::new(0), + lose_release_replies: false, + lost_release_replies: AtomicUsize::new(0), + drop_closed_finalize_replies: AtomicUsize::new(0), + expired_receiver_cleanups: AtomicUsize::new(0), + }); + let transport = Arc::new(IssuedSettlement { + inner: fleet.clone(), + issued: Mutex::new(None), + }); + let first = Arc::new( + FleetReconciler::new( + scope(), + session(9), + FleetProfile::default(), + fixture.native.journal.clone(), + fleet.clone(), + transport.clone(), + ) + .unwrap(), + ); + let (entered, resume) = fixture.native.journal.pause_next_role_result(boundary); + let task = tokio::spawn({ + let first = first.clone(); + async move { first.reconcile_once(clock, pass_deadline()).await } + }); + super::super::captured(entered).await; + let (action, proof) = transport.issued.lock().unwrap().take().unwrap(); + task.abort(); + assert!(task.await.unwrap_err().is_cancelled()); + { + let work = fixture.nodes[1].fleet_action_work().unwrap().unwrap(); + assert_eq!(work.entries().len(), 1); + assert_eq!(work.entries()[0].state(), FleetActionWorkState::Running); + assert!(!work.entries()[0].response_received()); + assert_eq!(work.entries()[0].committed(), None); + } + let old = fixture.native.journal.load_snapshot(scope()).await.unwrap(); + assert_eq!( + old.head().maintenance().unwrap().phase(), + MaintenancePhase::Evacuating + ); + let original = fixture + .native + .journal + .accept_action(&action, node_id(1), session(1), clock().unwrap()) + .await + .unwrap(); + let FleetActionAcceptance::Existing { accepted, result } = original else { + panic!("original acceptance missing"); + }; + assert_eq!( + result.is_some(), + boundary == ResultWriteBoundary::AfterCommit + ); + // An exact duplicate joins the retained native task, rather than executing + // the observation again. Poll it while the journal gate is still closed. + let mut replay = Box::pin(fixture.nodes[1].apply_fleet_role_settlement( + action.clone(), + proof, + clock().unwrap(), + )); + std::future::poll_fn(|cx| { + assert!(matches!(replay.as_mut().poll(cx), Poll::Pending)); + Poll::Ready(()) + }) + .await; + let client = Arc::new(client(&fixture).await); + let driver = FleetReconciler::new( + scope(), + session(9), + FleetProfile::default(), + client.clone(), + fleet.clone(), + fleet.clone(), + ) + .unwrap(); + // The new request carries fresh evidence at a renewed head. It cannot + // replace a running original capture or use its historical receipt to close. + let blocked = driver.reconcile_once(clock, pass_deadline()).await.unwrap(); + assert_eq!(blocked.allocated, 0); + assert_eq!( + blocked.snapshot.head().maintenance().unwrap().phase(), + MaintenancePhase::Evacuating + ); + assert!( + blocked + .blockers + .contains(&cellule_runtime::fleet::operations::DrainBlocker::OutcomeUnknown) + ); + let refusal = blocked.maintenance_failure.as_ref().unwrap(); + assert!( + super::restart::has_conflict(refusal.as_ref()), + "wrong running-owner refusal: {refusal:?}" + ); + let current = client.load_snapshot(scope()).await.unwrap(); + assert!(current.head().revision() > old.head().revision()); + assert_eq!( + current.head().maintenance().unwrap().phase(), + MaintenancePhase::Evacuating + ); + { + let work = fixture.nodes[1].fleet_action_work().unwrap().unwrap(); + assert_eq!(work.entries().len(), 1); + assert_eq!(work.entries()[0].state(), FleetActionWorkState::Running); + } + resume.send(()).unwrap(); + let completion = replay.await.unwrap(); + assert_eq!(completion.accepted, accepted); + assert!(!completion.committed); + assert!( + matches!(completion.outcome.outcome, FleetOutcome::RolesSettledAt { head_revision, .. } + if head_revision == old.head().revision()) + ); + assert!( + matches!(completion.journal_error.as_deref(), Some(Error::Facility { name: "fleet-action-journal", source }) + if source.downcast_ref::().is_some()) + ); + assert_eq!(client.load_snapshot(scope()).await.unwrap(), current); + // Join/removal is its own failed publication boundary. Return the original + // Conflict and require another full capture before publishing a new receipt. + let refusal = driver + .reconcile_once(clock, pass_deadline()) + .await + .unwrap_err(); + assert!(super::restart::has_conflict(&refusal)); + let next = driver.reconcile_once(clock, pass_deadline()).await.unwrap(); + assert!(next.maintenance_failure.is_none()); + assert_eq!( + next.snapshot.head().maintenance().unwrap().phase(), + MaintenancePhase::Closing + ); + let published = client + .accept_action(&action, node_id(1), session(1), clock().unwrap()) + .await + .unwrap(); + let FleetActionAcceptance::Existing { + accepted: unchanged, + result: Some(published), + } = published + else { + panic!("fresh settlement missing"); + }; + assert_eq!(unchanged, accepted); + assert!( + matches!(published.outcome, FleetOutcome::RolesSettledAt { head_revision, .. } + if head_revision > old.head().revision()) + ); + assert_eq!( + super::latest(&fixture, publication.record().unwrap()).await, + *publication.record().unwrap() + ); + let completed = driver.reconcile_once(clock, pass_deadline()).await.unwrap(); + assert_eq!( + completed.snapshot.head().maintenance().unwrap().phase(), + MaintenancePhase::Completed + ); + assert_eq!(fixture.nodes[1].state(), NodeState::Stopped); + assert!( + fixture + .native + .directory + .is_withdrawn(session(1)) + .await + .unwrap() + ); + super::restart::readback(&fixture).await; + drop((first, driver, transport, fleet, capture, completion)); + client.close().await.unwrap(); + fixture.finish().await; +} + +fn pass_deadline() -> Instant { + Instant::now() + Duration::from_secs(10) +} + +struct IssuedSettlement { + inner: Arc, + issued: Mutex>, +} + +impl FleetTransport for IssuedSettlement { + fn dispatch<'a>( + &'a self, + action: &'a FleetAction, + end: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + self.inner.dispatch(action, end) + } + fn inspect<'a>( + &'a self, + request: &'a FleetInspectionRequest, + end: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + self.inner.inspect(request, end) + } + fn settle_roles<'a>( + &'a self, + action: &'a FleetAction, + proof: &'a FleetRoleSettlement, + end: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + assert!( + self.issued + .lock() + .unwrap() + .replace((action.clone(), proof.clone())) + .is_none() + ); + self.inner.settle_roles(action, proof, end) + } +} diff --git a/crates/cellule-host/minion/scenario/follower_tests/persisted/mod.rs b/crates/cellule-host/minion/scenario/follower_tests/persisted/mod.rs index a0359be5..922cc05f 100644 --- a/crates/cellule-host/minion/scenario/follower_tests/persisted/mod.rs +++ b/crates/cellule-host/minion/scenario/follower_tests/persisted/mod.rs @@ -9,9 +9,11 @@ use cellule_host::{ }, }; use cellule_runtime::fleet::operations::{FollowerEvacuationRecord, FollowerReplacementPolicy}; +mod cancelled_settlement; mod maintenance; mod observation; mod races; +mod restart; mod tests; fn deadline() -> Instant { diff --git a/crates/cellule-host/minion/scenario/follower_tests/persisted/observation.rs b/crates/cellule-host/minion/scenario/follower_tests/persisted/observation.rs index 1772ecde..3b749a78 100644 --- a/crates/cellule-host/minion/scenario/follower_tests/persisted/observation.rs +++ b/crates/cellule-host/minion/scenario/follower_tests/persisted/observation.rs @@ -3,8 +3,8 @@ use super::super::{aggregate, coverage}; use super::*; use cellule_host::fleet::{ FleetActionCompletion, FleetAdapterFuture, FleetFollowerEvacuationCheck, - FleetMaintenanceEnrollments, FleetObservation, FleetObserver, FleetRoleCoverage, FleetRoster, - FleetTransport, + FleetMaintenanceEnrollments, FleetObservation, FleetObserver, FleetReconciler, + FleetRoleCoverage, FleetRoster, FleetTransport, }; use cellule_runtime::fleet::operations::{ FleetAction, FleetInspectionObservation, FleetInspectionRequest, @@ -251,6 +251,119 @@ async fn follower_evacuation_observation_retains_native_policy_and_role_graph_in Arc::try_unwrap(fixture).ok().unwrap().finish().await; } +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn reference_observer_reconciles_follower_only_maintenance_to_completion() { + let (fixture, capture, policy) = setup().await; + let publication = publish(&fixture, &capture, policy).await; + let record = publication.record().unwrap().clone(); + let fleet = Arc::new(crate::scenario::adapters::LocalFleet { + nodes: fixture.nodes.clone(), + journal: fixture.native.journal.clone(), + boots: fixture.boots.clone(), + records: fixture.records.clone(), + reader_verifier: None, + capture_sequence: std::sync::atomic::AtomicU64::new(0), + lose_release_replies: false, + lost_release_replies: AtomicUsize::new(0), + drop_closed_finalize_replies: std::sync::atomic::AtomicUsize::new(0), + expired_receiver_cleanups: AtomicUsize::new(0), + }); + + let snapshot = fixture.native.journal.load_snapshot(scope()).await.unwrap(); + let roster = FleetRoster::collect(fixture.native.journal.as_ref(), &snapshot, deadline()) + .await + .unwrap(); + let observation = fleet + .observe(&roster, clock().unwrap(), deadline()) + .await + .unwrap(); + let graph = observation.role_coverage().unwrap(); + assert_eq!(graph.native_boots(), 4); + assert_eq!(graph.physical_nodes(), 4); + let policies = observation.maintenance_policy_coverage().unwrap(); + assert!(policies.is_complete()); + assert_eq!(policies.progress().required, 1); + assert_eq!(policies.progress().checked, 1); + assert_eq!( + policies.obligations()[0].status(), + cellule_host::fleet::FleetMaintenancePolicyStatus::Follower(record.digest().unwrap()) + ); + + fixture + .native + .journal + .set_scheduling(roster.snapshot().registry(), true) + .await + .unwrap(); + let driver = FleetReconciler::new( + scope(), + session(9), + FleetProfile::default(), + fixture.native.journal.clone(), + fleet.clone(), + fleet.clone(), + ) + .unwrap(); + let mut report = None; + let mut policy_progress = None; + for _ in 0..3 { + let next = driver.reconcile_once(clock, deadline()).await.unwrap(); + if let Some(progress) = next.maintenance_policy { + policy_progress = Some(progress); + } + let completed = next.snapshot.head().maintenance().is_some_and(|operation| { + operation.phase() == cellule_runtime::fleet::operations::MaintenancePhase::Completed + }); + report = Some(next); + if completed { + break; + } + } + let report = report.unwrap(); + assert_eq!( + report.snapshot.head().maintenance().unwrap().phase(), + cellule_runtime::fleet::operations::MaintenancePhase::Completed, + "follower-only maintenance failed to settle: {report:?}" + ); + let progress = policy_progress.unwrap(); + assert_eq!(progress.required, 1); + assert_eq!(progress.checked, 1); + assert_eq!(fixture.nodes[1].state(), NodeState::Stopped); + assert!( + fixture + .native + .directory + .is_withdrawn(session(1)) + .await + .unwrap() + ); + let rows = fixture.native.rows().await; + let donor = rows + .iter() + .find(|row| { + row.spec().target.session == session(1) + && matches!( + row.spec().role, + cellule_runtime::fleet::operations::EnrollmentRole::Follower { log_epoch: 1 } + ) + }) + .unwrap(); + assert_eq!(donor.status(), EnrollmentStatus::Retired); + let leader = fixture + .native + .directory + .load(session(0), clock().unwrap()) + .await + .unwrap() + .unwrap(); + let current_log = leader.advertisement().log().unwrap(); + assert_eq!(current_log.epoch(), 2); + assert_eq!(current_log.members(), [node_id(2), node_id(3)]); + + drop((driver, fleet, capture)); + fixture.finish().await; +} + #[tokio::test(flavor = "multi_thread", worker_threads = 2)] async fn follower_evacuation_observation_refuses_duplicate_ensembles_and_stale_full_head_in_both_orders() { diff --git a/crates/cellule-host/minion/scenario/follower_tests/persisted/restart.rs b/crates/cellule-host/minion/scenario/follower_tests/persisted/restart.rs new file mode 100644 index 00000000..ab4bb6b3 --- /dev/null +++ b/crates/cellule-host/minion/scenario/follower_tests/persisted/restart.rs @@ -0,0 +1,418 @@ +//! A lost settlement reply must be refreshed at the reconstructed driver's head. +use super::*; +use cellule_host::fleet::{ + FleetActionCompletion, FleetActionJournal, FleetAdapterFuture, FleetReconciler, + FleetRoleSettlement, FleetTransport, +}; +use cellule_runtime::fleet::operations::{ + FleetAction, FleetInspectionObservation, FleetInspectionRequest, FleetOutcome, MaintenancePhase, +}; + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn reconstructed_controller_refreshes_native_settlement_after_a_lost_reply() { + run(Restart::Renew, Publication::LostTransportReply).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn new_claimant_after_real_expiry_refreshes_settlement_and_fences_old_controller() { + run(Restart::ReplaceAfterExpiry, Publication::LostTransportReply).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn failed_role_result_write_reconstructs_current_settlement() { + run( + Restart::Renew, + Publication::Failure(ResultWriteBoundary::BeforeCommit), + ) + .await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn unconfirmed_role_result_commit_reconstructs_current_settlement() { + run( + Restart::Renew, + Publication::Failure(ResultWriteBoundary::AfterCommit), + ) + .await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn new_claimant_after_expiry_resumes_failed_role_result_write() { + run( + Restart::ReplaceAfterExpiry, + Publication::Failure(ResultWriteBoundary::BeforeCommit), + ) + .await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn new_claimant_after_expiry_resumes_unconfirmed_role_result_commit() { + run( + Restart::ReplaceAfterExpiry, + Publication::Failure(ResultWriteBoundary::AfterCommit), + ) + .await; +} + +use crate::journal::ResultWriteBoundary; + +#[derive(Clone, Copy)] +enum Publication { + LostTransportReply, + Failure(ResultWriteBoundary), +} + +#[derive(Clone, Copy)] +enum Restart { + Renew, + ReplaceAfterExpiry, +} + +async fn run(restart: Restart, publication_fault: Publication) { + let (fixture, capture, policy) = setup().await; + let publication = publish(&fixture, &capture, policy).await; + let fleet = Arc::new(crate::scenario::adapters::LocalFleet { + nodes: fixture.nodes.clone(), + journal: fixture.native.journal.clone(), + boots: fixture.boots.clone(), + records: fixture.records.clone(), + reader_verifier: None, + capture_sequence: std::sync::atomic::AtomicU64::new(0), + lose_release_replies: false, + lost_release_replies: AtomicUsize::new(0), + drop_closed_finalize_replies: AtomicUsize::new(0), + expired_receiver_cleanups: AtomicUsize::new(0), + }); + if let Publication::Failure(boundary) = publication_fault { + fixture.native.journal.fail_next_role_result(boundary); + } + let transport = Arc::new(LostSettlementReply { + inner: fleet.clone(), + publication: publication_fault, + original: Mutex::new(None), + }); + let first = FleetReconciler::new( + scope(), + session(9), + FleetProfile::default(), + fixture.native.journal.clone(), + fleet.clone(), + transport.clone(), + ) + .unwrap(); + let before = first + .reconcile_once(clock, restart_deadline()) + .await + .unwrap(); + assert!(before.maintenance_failure.is_some()); + assert_eq!( + before.snapshot.head().maintenance().unwrap().phase(), + MaintenancePhase::Evacuating + ); + let original = transport.original.lock().unwrap().take().unwrap(); + let action = original.accepted.action().clone(); + assert_eq!( + original.committed, + matches!(publication_fault, Publication::LostTransportReply) + ); + if let Publication::Failure(boundary) = publication_fault { + let expected = match boundary { + ResultWriteBoundary::BeforeCommit => "injected role result before commit", + ResultWriteBoundary::AfterCommit => "injected role result after commit", + }; + assert!( + matches!(original.journal_error.as_deref(), Some(Error::Facility { name: "fleet-action-journal", source }) + if source.downcast_ref::().is_some_and(|error| error.to_string() == expected)) + ); + let accepted = fixture + .native + .journal + .accept_action(&action, node_id(1), session(1), clock().unwrap()) + .await + .unwrap(); + let cellule_host::fleet::FleetActionAcceptance::Existing { result, .. } = accepted else { + panic!("original acceptance missing"); + }; + assert_eq!( + result.is_some(), + boundary == ResultWriteBoundary::AfterCommit + ); + if let Some(result) = result { + assert_eq!(*result, original.outcome); + } + let work = fixture.nodes[1].fleet_action_work().unwrap().unwrap(); + assert_eq!(work.entries().len(), 1); + assert_eq!(work.entries()[0].committed(), Some(false)); + assert_eq!(work.entries()[0].outcome_known(), Some(true)); + assert!(work.entries()[0].journal_error().is_some()); + } + assert!( + matches!(original.outcome.outcome, FleetOutcome::RolesSettledAt { head_revision, .. } if head_revision == before.snapshot.head().revision()) + ); + // The transport dropped the successful reply before ReadyToClose. A fresh + // driver over an independent client must renew and recheck the same intent. + let claimant = match restart { + Restart::Renew => session(9), + Restart::ReplaceAfterExpiry => { + let expires = before.snapshot.head().controller().unwrap().expires_at_ms; + // Advance real wall time while the application's original lease + // owners publish actual heartbeats. No logical clock jump or + // fabricated advertisement keeps native participants admitted. + while clock().unwrap() <= expires { + for (index, boot) in fixture.boots.iter().enumerate() { + boot.refresh_capacity(index, fixture.native.journal.as_ref(), deadline()) + .await + .unwrap(); + } + tokio::time::sleep(Duration::from_millis(100)).await; + } + assert!(clock().unwrap() > expires); + session(10) + } + }; + let client = Arc::new(client(&fixture).await); + let driver = FleetReconciler::new( + scope(), + claimant, + FleetProfile::default(), + client.clone(), + fleet.clone(), + fleet.clone(), + ) + .unwrap(); + if matches!(publication_fault, Publication::Failure(_)) { + // The old read-only proof cannot be published at this renewed head. + // Preserve that refusal, join its original owner, and require a wholly + // fresh observation on the next pass rather than restamping its rows. + let refused = driver + .reconcile_once(clock, restart_deadline()) + .await + .unwrap_err(); + assert!( + has_conflict(&refused), + "wrong publication refusal: {refused:?}" + ); + let snapshot = client.load_snapshot(scope()).await.unwrap(); + assert_eq!( + snapshot.head().maintenance().unwrap().phase(), + MaintenancePhase::Evacuating + ); + let work = fixture.nodes[1].fleet_action_work().unwrap().unwrap(); + assert!( + work.entries().is_empty(), + "joined stale role proof blocked fresh observation: {:?}", + work.entries() + ); + assert!(work.failure().is_none()); + assert_eq!(fixture.nodes[1].state(), NodeState::Ready); + } + let next = driver + .reconcile_once(clock, restart_deadline()) + .await + .unwrap(); + assert!( + next.maintenance_failure.is_none(), + "fresh settlement failed: {next:?}" + ); + assert_eq!( + next.snapshot.head().maintenance().unwrap().phase(), + MaintenancePhase::Closing + ); + assert!(next.snapshot.head().revision() > before.snapshot.head().revision()); + let controller = next.snapshot.head().controller().unwrap(); + assert_eq!(controller.claimant, claimant); + assert_eq!( + controller.epoch, + match restart { + Restart::Renew => 1, + Restart::ReplaceAfterExpiry => 2, + } + ); + if matches!(restart, Restart::ReplaceAfterExpiry) { + let current = client.load_snapshot(scope()).await.unwrap(); + let rejected = first + .reconcile_once(clock, restart_deadline()) + .await + .unwrap_err(); + assert!( + matches!(&rejected, Error::Facility { name: "fleet-journal", source } + if matches!(source.downcast_ref::(), + Some(cellule_runtime::fleet::operations::OperationError::Fenced))), + "old controller returned the wrong refusal: {rejected:?}" + ); + assert_eq!(client.load_snapshot(scope()).await.unwrap(), current); + } + drop(first); + let latest = fixture + .native + .journal + .accept_action(&action, node_id(1), session(1), clock().unwrap()) + .await + .unwrap(); + let cellule_host::fleet::FleetActionAcceptance::Existing { + accepted, + result: Some(result), + } = latest + else { + panic!("settlement result missing") + }; + assert_eq!(accepted.action(), &action); + assert_eq!(accepted, original.accepted); + assert!( + matches!(result.outcome, FleetOutcome::RolesSettledAt { head_revision, .. } if head_revision > before.snapshot.head().revision()) + ); + assert_eq!( + super::latest(&fixture, publication.record().unwrap()).await, + *publication.record().unwrap() + ); + let completed = driver + .reconcile_once(clock, restart_deadline()) + .await + .unwrap(); + assert_eq!( + completed.snapshot.head().maintenance().unwrap().phase(), + MaintenancePhase::Completed + ); + assert_eq!(fixture.nodes[1].state(), NodeState::Stopped); + assert!( + fixture + .native + .directory + .is_withdrawn(session(1)) + .await + .unwrap() + ); + readback(&fixture).await; + drop((driver, fleet, capture, original)); + client.close().await.unwrap(); + fixture.finish().await; +} + +pub(super) async fn readback(fixture: &ManagedFixture) { + // Rotation drained the original actor. Restore the exact current canonical + // root instead of reading a stale handle or bypassing its admission gate. + assert_eq!(fixture.records.len(), 1); + let record = fixture.records.values().next().unwrap(); + let control = record + .authority + .load(record.target.cell_id()) + .await + .unwrap() + .unwrap(); + let root = control.value().ltx_root().unwrap(); + let destination = fixture.native.root.path().join("restart-readback.sqlite"); + assert_eq!( + record + .replica + .open_root(&root) + .await + .unwrap() + .restore(&destination) + .await + .unwrap(), + root.position + ); + let database = rusqlite::Connection::open_with_flags( + destination, + rusqlite::OpenFlags::SQLITE_OPEN_READ_ONLY, + ) + .unwrap(); + assert_eq!( + database + .query_row("SELECT value FROM counter", [], |row| row.get::<_, i64>(0)) + .unwrap(), + 29 + ); + assert_eq!( + database + .query_row( + "SELECT result FROM sys_requests WHERE request_id=?1", + [RequestId::from_bytes([1; 16]).as_bytes().as_slice()], + |row| row.get::<_, Vec>(0) + ) + .unwrap(), + 29i64.to_be_bytes() + ); + drop(database); +} + +fn restart_deadline() -> Instant { + Instant::now() + Duration::from_secs(10) +} + +struct LostSettlementReply { + inner: Arc, + publication: Publication, + original: Mutex>>, +} +impl FleetTransport for LostSettlementReply { + fn dispatch<'a>( + &'a self, + action: &'a FleetAction, + end: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + self.inner.dispatch(action, end) + } + fn inspect<'a>( + &'a self, + request: &'a FleetInspectionRequest, + end: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + self.inner.inspect(request, end) + } + fn settle_roles<'a>( + &'a self, + action: &'a FleetAction, + proof: &'a FleetRoleSettlement, + end: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + Box::pin(async move { + let completion = self.inner.settle_roles(action, proof, end).await?; + assert_eq!( + completion.committed, + matches!(self.publication, Publication::LostTransportReply) + ); + assert!(matches!( + completion.outcome.outcome, + FleetOutcome::RolesSettledAt { .. } + )); + assert!( + self.original + .lock() + .unwrap() + .replace(completion.clone()) + .is_none() + ); + match self.publication { + Publication::LostTransportReply => { + Err(invalid("injected loss after committed role settlement")) + } + Publication::Failure(_) => Ok(completion), + } + }) + } +} + +pub(super) fn has_conflict(mut source: &(dyn std::error::Error + 'static)) -> bool { + loop { + if let Some(error) = source.downcast_ref::>() { + return has_conflict(error.as_ref()); + } + if let Some(Error::FleetOperation(error)) = source.downcast_ref::() { + return matches!( + error.as_ref(), + cellule_runtime::fleet::operations::OperationError::Conflict + ); + } + if matches!( + source.downcast_ref::(), + Some(cellule_runtime::fleet::operations::OperationError::Conflict) + ) { + return true; + } + match source.source() { + Some(next) => source = next, + None => return false, + } + } +} diff --git a/crates/cellule-host/minion/scenario/follower_tests/persisted/tests.rs b/crates/cellule-host/minion/scenario/follower_tests/persisted/tests.rs index a8f79f13..4b43b3b9 100644 --- a/crates/cellule-host/minion/scenario/follower_tests/persisted/tests.rs +++ b/crates/cellule-host/minion/scenario/follower_tests/persisted/tests.rs @@ -215,6 +215,9 @@ async fn follower_policy_publication_closing_and_deadline_refresh_preserve_origi let record = result.record().unwrap(); result.confirmed().unwrap(); let snapshot = fixture.native.journal.load_snapshot(scope()).await.unwrap(); + crate::scenario::commit_test_role_settlement(&fixture.native.journal, clock().unwrap()) + .await + .unwrap(); let operation = snapshot.head().maintenance().unwrap(); let after = fixture .native diff --git a/crates/cellule-host/minion/scenario/follower_tests/reference_batches.rs b/crates/cellule-host/minion/scenario/follower_tests/reference_batches.rs new file mode 100644 index 00000000..35960da7 --- /dev/null +++ b/crates/cellule-host/minion/scenario/follower_tests/reference_batches.rs @@ -0,0 +1,160 @@ +//! Shared foreign traversal keeps the complete roster and global recheck barrier. +use super::*; +use cellule_host::fleet::FleetFollowerReferences; + +fn deadline() -> Instant { + Instant::now() + Duration::from_secs(5) +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn shared_reference_windows_match_individual_complete_captures() { + let fixture = ManagedFixture::new().await; + let roster = aggregate::roster(&fixture).await; + let members = [node_id(2), node_id(0), node_id(1)]; + let mut batch = FleetFollowerReferences::collect_all( + &fixture.native.directory, + &roster, + &members, + 1, + deadline(), + clock, + ) + .await + .unwrap(); + assert_eq!(batch.len(), members.len()); + for (references, member) in batch.iter().zip(members) { + assert_eq!(references.member(), member); + let single = FleetFollowerReferences::collect( + &fixture.native.directory, + &roster, + member, + 1, + deadline(), + clock, + ) + .await + .unwrap(); + assert_eq!(references.entries(), single.entries()); + references.validate_enrollments(&roster).unwrap(); + } + assert!(batch[1].entries().is_empty()); + assert_eq!(batch[0].entries().len(), 1); + assert_eq!(batch[2].entries().len(), 1); + FleetFollowerReferences::recheck_all( + &mut batch, + &fixture.native.directory, + &roster, + 1, + deadline(), + clock, + ) + .await + .unwrap(); + for members in [ + vec![], + vec![node_id(1), node_id(1)], + vec![NodeId::from_bytes([250; 16])], + ] { + assert!( + FleetFollowerReferences::collect_all( + &fixture.native.directory, + &roster, + &members, + 1, + deadline(), + clock + ) + .await + .is_err() + ); + } + drop(batch); + fixture.finish().await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn failed_shared_recheck_invalidates_every_original_without_restamping() { + let fixture = ManagedFixture::new().await; + let roster = aggregate::roster(&fixture).await; + let (mut native, mut foreign, mut sequence) = coverage::captures(&fixture, &roster).await; + coverage::native_rechecks(&fixture, &roster, &mut native, &mut sequence).await; + coverage::foreign_rechecks(&fixture, &roster, &mut foreign).await; + assert!(coverage::check(&roster, &native, &foreign).is_ok()); + let original = foreign + .iter() + .map(|references| (references.interval(), references.entries().to_vec())) + .collect::>(); + assert!( + FleetFollowerReferences::recheck_all( + &mut foreign, + &fixture.native.directory, + &roster, + 1, + Instant::now(), + clock + ) + .await + .is_err() + ); + for (references, (interval, entries)) in foreign.iter().zip(&original) { + assert_eq!(&references.interval(), interval); + assert_eq!(references.entries(), entries); + } + assert!(coverage::check(&roster, &native, &foreign).is_err()); + for index in 0..foreign.len() { + foreign[index] + .recheck(&fixture.native.directory, &roster, 1, deadline(), clock) + .await + .unwrap(); + assert_eq!( + coverage::check(&roster, &native, &foreign).is_ok(), + index + 1 == foreign.len(), + "every member needs a new confirmation after the failed batch" + ); + } + assert!(coverage::check(&roster, &native, &foreign).is_ok()); + // Exact rows include liveness even when topology/cursor bytes are unchanged. + // This negative logical-clock capture grants no lease or takeover authority. + let owner = fixture + .native + .directory + .load(session(0), clock().unwrap()) + .await + .unwrap() + .unwrap(); + let expired_at = owner.advertisement().expires_at_ms() + 1; + let retried = foreign + .iter() + .map(FleetFollowerReferences::interval) + .collect::>(); + let result = FleetFollowerReferences::recheck_all( + &mut foreign, + &fixture.native.directory, + &roster, + 1, + deadline(), + || Ok(expired_at), + ) + .await; + assert!(matches!( + result, + Err(Error::Node("authoritative follower inventory changed")) + )); + for ((references, (interval, entries)), retried) in foreign.iter().zip(&original).zip(retried) { + // The successful retry may extend the interval; this failed recheck + // must retain that retry's end instead of stamping the future clock. + assert!(references.interval().1 < expired_at); + assert_eq!(references.interval(), retried); + assert_eq!(references.interval().0, interval.0); + assert_eq!(references.entries(), entries); + } + assert!(coverage::check(&roster, &native, &foreign).is_err()); + coverage::foreign_rechecks(&fixture, &roster, &mut foreign).await; + roster + .confirm(fixture.native.journal.as_ref(), deadline()) + .await + .unwrap(); + assert!(coverage::check(&roster, &native, &foreign).is_ok()); + drop((native, foreign)); + fixture.finish().await; +} diff --git a/crates/cellule-host/minion/scenario/maintenance/mod.rs b/crates/cellule-host/minion/scenario/maintenance/mod.rs new file mode 100644 index 00000000..e261814d --- /dev/null +++ b/crates/cellule-host/minion/scenario/maintenance/mod.rs @@ -0,0 +1,279 @@ +//! Planned maintenance through the canonical driver, with optional busy SQL admission. +use super::*; + +#[cfg(test)] +mod tests; +mod traffic; + +pub(super) async fn run( + root: &tempfile::TempDir, + journal: Arc, + nodes: &mut Vec>, + boots: &mut Vec, + profile: FleetProfile, + busy: bool, +) -> JournalResult { + let (records, acknowledged) = initialize(root, &journal, nodes, boots, 300_000).await?; + let mut traffic = if busy { + Some(traffic::Traffic::start(&acknowledged)?) + } else { + None + }; + let result = async { + if let Some(traffic) = &mut traffic { + traffic.started().await?; + } + Box::pin(run_nodes( + journal, + nodes, + boots, + profile, + records.clone(), + &acknowledged, + )) + .await + } + .await; + // Stop new offering on every exit and join all accepted client calls before + // the shared node cleanup can close the journal or temporary databases. + let joined = match traffic { + Some(traffic) => Some(traffic.finish().await), + None => None, + }; + let mut summary = match result { + Ok(summary) => summary, + Err(error) => { + if let Some(Err(client_error)) = joined { + eprintln!("additional joined maintenance client failure: {client_error:?}"); + } + return Err(error); + } + }; + if let Some(joined) = joined { + summary.receipt_checks += joined?.verify(nodes, &records).await?; + } + Ok(summary) +} + +async fn run_nodes( + journal: Arc, + nodes: &[Arc], + boots: &[startup::BootOwner], + profile: FleetProfile, + records: Arc>, + acknowledged: &HashMap, +) -> JournalResult { + use cellule_host::fleet::FleetRoster; + use cellule_runtime::fleet::operations::{EnrollmentStatus, MaintenancePhase}; + let fleet = Arc::new(adapters::LocalFleet { + nodes: nodes.to_vec(), + journal: journal.clone(), + boots: boots.to_vec(), + records: records.clone(), + reader_verifier: None, + capture_sequence: std::sync::atomic::AtomicU64::new(0), + lose_release_replies: false, + lost_release_replies: std::sync::atomic::AtomicUsize::new(0), + drop_closed_finalize_replies: std::sync::atomic::AtomicUsize::new(0), + expired_receiver_cleanups: std::sync::atomic::AtomicUsize::new(0), + }); + let claimant = SessionId::from_bytes([206; 16]); + let now_ms = clock()?; + let initial = journal.load_snapshot(scope()).await?; + let claimed = journal + .claim_controller(scope(), initial.head().revision(), claimant, now_ms) + .await?; + let operation = MaintenanceOperation::new( + OperationId::from_bytes([222; 16])?, + Digest::from_bytes([223; 32]), + node_id(0), + session(0), + 2, + now_ms, + now_ms + .checked_add(180_000) + .ok_or_else(|| invalid("maintenance deadline overflow"))?, + )?; + let epoch = claimed + .head() + .controller() + .ok_or_else(|| invalid("maintenance controller lease absent"))? + .epoch; + journal + .compare_exchange( + &claimed, + epoch, + now_ms, + &JournalTransition::BeginMaintenance(operation.clone()), + ) + .await?; + let driver = FleetReconciler::new( + scope(), + claimant, + profile, + journal.clone(), + fleet.clone(), + fleet.clone(), + )?; + let mut summary = ScenarioSummary { + released: 0, + activated: 0, + retired: 0, + receipt_checks: 0, + max_inflight: 0, + max_restore_bytes: 0, + joined_nodes: 0, + boot_retirements: 0, + receiver_nodes: 0, + lost_release_replies: 0, + controller_epoch: 0, + expired_receiver_cleanups: 0, + blockers: Vec::new(), + final_counts: vec![0; 3], + maintenance_completed: false, + maintenance_boot_withdrawn: false, + receiver_process_closures: 0, + lost_activation_replies: 0, + routed_activation_replays: 0, + }; + let wall_deadline = Instant::now() + Duration::from_secs(180); + let mut specs = HashMap::new(); + let mut completed_snapshot = None; + while Instant::now() < wall_deadline { + let report = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(8)) + .await?; + if let Some(failure) = report.failures.first() { + return Err(Box::new(Arc::clone(&failure.error)) as JournalError); + } + if let Some(failure) = &report.maintenance_failure { + return Err(Box::new(Arc::clone(failure)) as JournalError); + } + for attempt in report.snapshot.head().attempts() { + let spec = attempt.spec(); + if let Some(original) = specs.insert(spec.target.cell_id(), spec.clone()) + && original != *spec + { + return Err(invalid("maintenance moved a Cell more than once")); + } + } + summary.released += report.released; + summary.activated += report.activated; + summary.retired += report.retired; + summary.max_inflight = summary + .max_inflight + .max(report.snapshot.head().attempts().len()); + summary.max_restore_bytes = summary + .max_restore_bytes + .max(report.snapshot.head().reserved_restore_bytes()); + summary.controller_epoch = report + .snapshot + .head() + .controller() + .ok_or_else(|| invalid("maintenance controller lease absent"))? + .epoch; + for blocker in report.blockers { + if !summary.blockers.contains(&blocker) { + summary.blockers.push(blocker); + } + } + if report.snapshot.head().maintenance().is_some_and(|current| { + current.id() == operation.id() && current.phase() == MaintenancePhase::Completed + }) { + completed_snapshot = Some(report.snapshot); + break; + } + tokio::time::sleep(Duration::from_millis(10)).await; + } + let snapshot = match completed_snapshot { + Some(snapshot) => snapshot, + None => { + let retained = journal.load_snapshot(scope()).await?; + return Err(std::io::Error::other(format!( + "maintenance did not complete before its deadline: operation={:?} phase={:?} attempts={:?} blockers={:?}", + operation.id(), + retained.head().maintenance().map(MaintenanceOperation::phase), + retained.head().attempts(), + summary.blockers, + )).into()); + } + }; + let current = snapshot + .head() + .maintenance() + .ok_or_else(|| invalid("completed maintenance operation is absent"))?; + let evidence = current + .drain_evidence() + .ok_or_else(|| invalid("completed maintenance evidence is absent"))?; + if current.id() != operation.id() + || current.phase() != MaintenancePhase::Completed + || !snapshot.head().attempts().is_empty() + || evidence.remaining_cells != 0 + || evidence.unresolved_attempts != 0 + || !evidence.relocated + || !evidence.readers_settled + || !evidence.followers_settled + || nodes[0].state() != NodeState::Stopped + || nodes[0].stats().active_cells() != 0 + { + return Err(invalid( + "maintenance completed without the full drain barrier", + )); + } + let boot = journal + .load_enrollment(scope(), boots[0].spec.key()?) + .await? + .ok_or_else(|| invalid("maintenance boot enrollment is absent"))?; + if boot.status() != EnrollmentStatus::Retired + || !boots[0].directory.is_withdrawn(session(0)).await? + { + return Err(invalid("maintenance stopped without exact boot withdrawal")); + } + summary.maintenance_completed = true; + summary.maintenance_boot_withdrawn = true; + summary.lost_release_replies = fleet + .lost_release_replies + .load(std::sync::atomic::Ordering::SeqCst); + summary.expired_receiver_cleanups = fleet + .expired_receiver_cleanups + .load(std::sync::atomic::Ordering::SeqCst); + if specs.len() != CELL_COUNT + || summary.released != CELL_COUNT + || summary.activated != CELL_COUNT + || summary.retired != CELL_COUNT + { + return Err(std::io::Error::other(format!( + "maintenance did not relocate every Cell: summary={summary:?} specs={}", + specs.len() + )) + .into()); + } + let mut destinations = std::collections::HashSet::new(); + for spec in specs.values() { + destinations.insert(spec.destination); + verify_movement(&fleet, &records, acknowledged, spec).await?; + summary.receipt_checks += 1; + } + summary.receiver_nodes = destinations.len(); + if summary.receiver_nodes != 2 { + return Err(invalid("maintenance did not use both eligible receivers")); + } + let roster = FleetRoster::collect( + journal.as_ref(), + &snapshot, + Instant::now() + Duration::from_secs(8), + ) + .await?; + summary.final_counts = + observation::complete_counts(&fleet, &roster, Instant::now() + Duration::from_secs(8)) + .await? + .ok_or_else(|| invalid("post-maintenance observation is incomplete"))?; + if summary.final_counts[0] != 0 || summary.final_counts.iter().sum::() != CELL_COUNT { + return Err(std::io::Error::other(format!( + "maintenance left Cells on the stopped node or lost inventory: counts={:?}", + summary.final_counts + )) + .into()); + } + Ok(summary) +} diff --git a/crates/cellule-host/minion/scenario/maintenance/tests.rs b/crates/cellule-host/minion/scenario/maintenance/tests.rs new file mode 100644 index 00000000..60af623a --- /dev/null +++ b/crates/cellule-host/minion/scenario/maintenance/tests.rs @@ -0,0 +1,69 @@ +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn executable_busy_maintenance_closes_sustained_admission_preserves_every_outcome_and_joins() +{ + let summary = super::super::maintenance_busy().await.unwrap(); + assert_eq!( + (summary.released, summary.activated, summary.retired), + (12, 12, 12) + ); + assert!(summary.receipt_checks >= 14); + assert_eq!(summary.final_counts[0], 0); + assert_eq!(summary.final_counts.iter().sum::(), 12); + assert_eq!(summary.receiver_nodes, 2); + assert!(summary.maintenance_completed); + assert!(summary.maintenance_boot_withdrawn); + assert_eq!(summary.joined_nodes, 3); + assert_eq!(summary.boot_retirements, 3); +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn failed_startup_retains_native_error_and_joins_every_client_before_node_cleanup() { + use super::*; + let root = tempfile::tempdir().unwrap(); + let journal = Arc::new( + SqliteJournal::open( + root.path().join("busy-startup.sqlite"), + scope(), + FleetProfile::default(), + clock().unwrap(), + ) + .await + .unwrap(), + ); + let mut nodes = Vec::new(); + let mut boots = Vec::new(); + let (_, acknowledged) = initialize(&root, &journal, &mut nodes, &mut boots, 300_000) + .await + .unwrap(); + nodes[0].shutdown().await.unwrap(); + let mut traffic = traffic::Traffic::start(&acknowledged).unwrap(); + let error = traffic.started().await.unwrap_err(); + fn native(error: &(dyn std::error::Error + 'static)) -> bool { + matches!( + error.downcast_ref::(), + Some( + cellule_runtime::Error::RuntimeClosed + | cellule_runtime::Error::CellDraining + | cellule_runtime::Error::CellNotActive + | cellule_runtime::Error::Fenced + ) + ) || error.source().is_some_and(native) + } + assert!(native(error.as_ref()), "lost native source: {error:?}"); + let joined = match traffic.finish().await { + Ok(_) => panic!("closed original writer accepted sustained traffic"), + Err(error) => error, + }; + assert!(native(joined.as_ref()), "lost joined source: {joined:?}"); + for (node, boot) in nodes.iter().zip(&boots) { + node.shutdown().await.unwrap(); + boot.withdraw(&journal).await.unwrap(); + let stats = node.stats(); + assert_eq!(stats.active_cells(), 0); + assert_eq!(stats.worker_jobs(), 0); + assert_eq!(stats.retained_bytes(), 0); + assert_eq!(stats.file_descriptors(), 0); + assert_eq!(stats.local_disk_reserved_bytes(), 0); + } + journal.close().await.unwrap(); +} diff --git a/crates/cellule-host/minion/scenario/maintenance/traffic.rs b/crates/cellule-host/minion/scenario/maintenance/traffic.rs new file mode 100644 index 00000000..1912c440 --- /dev/null +++ b/crates/cellule-host/minion/scenario/maintenance/traffic.rs @@ -0,0 +1,305 @@ +//! Bounded client lanes retain every accepted outcome across native quiescence. +use super::*; +use cellule_runtime::Error; +use tokio::{sync::oneshot, task::JoinHandle}; + +const LANES: usize = 2; +const MAX_COMMANDS_PER_LANE: u64 = 512; +type ClientError = Arc; + +struct Receipt { + identity: MutationIdentity, + digest: Digest, + outcome: StoredOutcome, +} +struct Lane { + receipts: Vec, + refused: bool, +} +pub(super) struct Traffic { + cell: CellId, + source: CellHandle, + stop: CancellationToken, + started: Vec>>, + tasks: Vec>>, +} +pub(super) struct JoinedTraffic { + cell: CellId, + source: CellHandle, + lanes: Vec, +} + +impl Traffic { + pub(super) fn start(acknowledged: &HashMap) -> JournalResult { + let cell = acknowledged + .keys() + .min_by_key(|cell| *cell.as_bytes()) + .copied() + .ok_or_else(|| invalid("busy maintenance lacks an original writer"))?; + let source = acknowledged + .get(&cell) + .ok_or_else(|| invalid("busy maintenance original writer disappeared"))? + .source + .clone(); + let stop = CancellationToken::new(); + let mut started = Vec::with_capacity(LANES); + let mut tasks = Vec::with_capacity(LANES); + for lane in 0..LANES { + let (signal, ready) = oneshot::channel(); + started.push(ready); + let source = source.clone(); + let stop = stop.clone(); + tasks.push(tokio::spawn(async move { + offer(source, stop, lane, signal).await + })); + } + Ok(Self { + cell, + source, + stop, + started, + tasks, + }) + } + + pub(super) async fn started(&mut self) -> JournalResult<()> { + for ready in self.started.drain(..) { + let ready = tokio::time::timeout(Duration::from_secs(8), ready).await??; + ready.map_err(|error| Box::new(ClientFailure(error)) as JournalError)?; + } + Ok(()) + } + + pub(super) async fn finish(mut self) -> JournalResult { + self.stop.cancel(); + let mut lanes = Vec::with_capacity(LANES); + let mut errors = Vec::new(); + // Join every sibling even after one failed. A dropped execute waiter + // cannot establish whether its native SQL or durable response finished. + for task in std::mem::take(&mut self.tasks) { + match task.await { + Ok(Ok(lane)) => lanes.push(lane), + Ok(Err(error)) => errors.push(error), + Err(error) => errors.push(Box::new(error) as JournalError), + } + } + if !errors.is_empty() { + return Err(Box::new(TrafficFailures(errors))); + } + Ok(JoinedTraffic { + cell: self.cell, + source: self.source.clone(), + lanes, + }) + } +} + +impl Drop for Traffic { + fn drop(&mut self) { + // Cancellation stops new offering. Accepted execute calls stay with + // their native actors and the finite client tasks until completion. + self.stop.cancel(); + } +} + +async fn offer( + source: CellHandle, + stop: CancellationToken, + lane: usize, + signal: oneshot::Sender>, +) -> JournalResult { + let mut signal = Some(signal); + match offer_commands(source, stop, lane, &mut signal).await { + Err(error) if signal.is_some() => { + let shared = ClientError::from(error); + if let Some(signal) = signal.take() { + let _ = signal.send(Err(shared.clone())); + } + Err(Box::new(ClientFailure(shared))) + } + other => other, + } +} + +async fn offer_commands( + source: CellHandle, + stop: CancellationToken, + lane: usize, + signal: &mut Option>>, +) -> JournalResult { + let mut receipts = Vec::new(); + for sequence in 1..=MAX_COMMANDS_PER_LANE { + if stop.is_cancelled() { + return Ok(Lane { + receipts, + refused: false, + }); + } + let now = clock()?; + let mut id = [240; 16]; + id[1] = lane as u8; + id[2..10].copy_from_slice(&sequence.to_be_bytes()); + let identity = MutationIdentity { + request_id: RequestId::from_bytes(id), + issued_at_ms: now, + expires_at_ms: now + .checked_add(300_000) + .ok_or_else(|| invalid("busy maintenance request deadline overflow"))?, + }; + let digest = Digest::from_bytes(*blake3::hash(&id).as_bytes()); + match source + .execute(identity, digest, now, 64, 64, move |transaction| { + transaction.execute_batch( + "CREATE TABLE IF NOT EXISTS maintenance_commands(request BLOB PRIMARY KEY)", + )?; + transaction.execute( + "INSERT INTO maintenance_commands(request) VALUES (?1)", + [id.as_slice()], + )?; + Ok(HandlerOutcome::Success(id.to_vec())) + }) + .await + { + Ok(outcome) => { + receipts.push(Receipt { + identity, + digest, + outcome, + }); + if let Some(signal) = signal.take() { + let _ = signal.send(Ok(())); + } + } + Err(Error::CellDraining) if !receipts.is_empty() => { + return Ok(Lane { + receipts, + refused: true, + }); + } + Err(error) => return Err(Box::new(error)), + } + } + Err(invalid( + "busy maintenance exhausted offered load before closing admission", + )) +} + +impl JoinedTraffic { + pub(super) async fn verify( + self, + nodes: &[Arc], + records: &HashMap, + ) -> JournalResult { + if self.lanes.len() != LANES + || self + .lanes + .iter() + .any(|lane| !lane.refused || lane.receipts.is_empty()) + { + return Err(invalid( + "busy maintenance did not fence every sustained command lane", + )); + } + let record = records + .get(&self.cell) + .ok_or_else(|| invalid("busy maintenance readback record is missing"))?; + let current = record + .authority + .load(self.cell) + .await? + .ok_or_else(|| invalid("busy maintenance successor authority is missing"))?; + let destination = current + .value() + .owner + .as_ref() + .ok_or_else(|| invalid("busy maintenance successor owner is missing"))? + .session; + let node = nodes + .iter() + .enumerate() + .find(|(index, _)| session(*index) == destination) + .map(|(_, node)| node) + .ok_or_else(|| invalid("busy maintenance successor endpoint is missing"))?; + if destination == session(0) || nodes[0].state() != NodeState::Stopped { + return Err(invalid("busy maintenance retained its source writer")); + } + let handle = node + .runtime() + .local_handle(record.catalog.clone(), ¤t) + .await? + .ok_or_else(|| invalid("busy maintenance has no canonical successor actor"))?; + let mut count = 0; + for lane in self.lanes { + for receipt in lane.receipts { + if handle + .resolve(receipt.identity, receipt.digest, clock()?, 64) + .await? + != Resolution::Committed(receipt.outcome.clone()) + { + return Err(invalid("busy maintenance lost an accepted command outcome")); + } + let StoredOutcome::Success { result, .. } = receipt.outcome else { + return Err(invalid( + "busy maintenance accepted a non-success SQL outcome", + )); + }; + if result != receipt.identity.request_id.as_bytes() { + return Err(invalid( + "busy maintenance changed an accepted command result", + )); + } + count += 1; + } + } + let rows = handle + .query(64, 64, |database| { + Ok(database + .query_row("SELECT COUNT(*) FROM maintenance_commands", [], |row| { + row.get::<_, u64>(0) + })? + .to_be_bytes() + .to_vec()) + }) + .await?; + if rows != (count as u64).to_be_bytes() { + return Err(invalid("busy maintenance duplicated or lost SQL commands")); + } + match self.source.query(64, 64, |_| Ok(Vec::new())).await { + Err( + Error::Fenced | Error::CellDraining | Error::CellNotActive | Error::RuntimeClosed, + ) => {} + Err(error) => return Err(Box::new(error)), + Ok(_) => return Err(invalid("busy maintenance retained source service")), + } + println!( + "maintenance_traffic lanes={LANES} accepted_commands={count} admission_refusals={LANES} exact_audit_rows={count}" + ); + Ok(count) + } +} + +#[derive(Debug)] +struct ClientFailure(ClientError); +impl std::fmt::Display for ClientFailure { + fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + self.0.fmt(formatter) + } +} +impl std::error::Error for ClientFailure { + fn source(&self) -> Option<&(dyn std::error::Error + 'static)> { + Some(self.0.as_ref()) + } +} + +#[derive(Debug)] +struct TrafficFailures(Vec); +impl std::fmt::Display for TrafficFailures { + fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + write!(formatter, "busy maintenance client failures: {:?}", self.0) + } +} +impl std::error::Error for TrafficFailures { + fn source(&self) -> Option<&(dyn std::error::Error + 'static)> { + self.0.first().map(|error| error.as_ref() as _) + } +} diff --git a/crates/cellule-host/minion/scenario/mod.rs b/crates/cellule-host/minion/scenario/mod.rs index 200bf36b..0195cff2 100644 --- a/crates/cellule-host/minion/scenario/mod.rs +++ b/crates/cellule-host/minion/scenario/mod.rs @@ -6,11 +6,16 @@ mod adapters; mod application; mod balance; mod failure; +mod follower_maintenance; #[cfg(test)] mod follower_tests; +mod maintenance; +mod native_peers; mod observation; +mod reader_maintenance; #[cfg(test)] mod reader_tests; +mod receiver_loss; #[cfg(test)] mod recovered_followers; mod startup; @@ -20,7 +25,9 @@ mod successor_tests; mod tests; use super::journal::{JournalError, JournalResult, SqliteJournal}; -use cellule_host::fleet::{FleetEnrollmentJournal, FleetJournal, FleetReconciler}; +use cellule_host::fleet::{ + FleetAdapterFuture, FleetEnrollmentJournal, FleetJournal, FleetReconciler, +}; use cellule_host::{CellNode, CellNodeBuilder, NodeState}; use cellule_runtime::{ cell::{ @@ -30,7 +37,9 @@ use cellule_runtime::{ worker::SqlWorkerPool, }, control::{Owner, authority::CellAuthority}, - fleet::operations::{FleetProfile, FleetScope, NodeIntent}, + fleet::operations::{ + FleetProfile, FleetScope, JournalTransition, MaintenanceOperation, NodeIntent, OperationId, + }, identity::{ ApplicationId, CellId, CellTarget, Digest, IncarnationId, NodeId, RequestId, SessionId, TenantId, @@ -121,26 +130,70 @@ pub(super) struct ScenarioSummary { pub controller_epoch: u64, pub expired_receiver_cleanups: usize, pub blockers: Vec, - pub final_counts: [usize; 3], + pub final_counts: Vec, + pub maintenance_completed: bool, + pub maintenance_boot_withdrawn: bool, + pub receiver_process_closures: usize, + pub lost_activation_replies: usize, + pub routed_activation_replays: usize, } /// Owns the private directory until all runtime and journal jobs are joined. pub(super) async fn overload() -> JournalResult { - execute(false, false).await + execute(Scenario::Overload).await } pub(super) async fn controller_restart() -> JournalResult { - execute(true, false).await + execute(Scenario::ControllerRestart).await } pub(super) async fn count_balance() -> JournalResult { - execute(false, true).await + execute(Scenario::CountBalance).await } -async fn execute(restart: bool, count_balance: bool) -> JournalResult { +pub(super) async fn maintenance() -> JournalResult { + execute(Scenario::Maintenance).await +} + +pub(super) async fn maintenance_busy() -> JournalResult { + execute(Scenario::MaintenanceBusy).await +} + +pub(super) async fn maintenance_reader() -> JournalResult { + execute(Scenario::MaintenanceReader).await +} + +pub(super) async fn maintenance_follower() -> JournalResult { + execute(Scenario::MaintenanceFollower).await +} + +pub(super) async fn maintenance_roles() -> JournalResult { + execute(Scenario::MaintenanceRoles).await +} + +pub(super) async fn receiver_loss() -> JournalResult { + execute(Scenario::ReceiverLoss).await +} + +enum Scenario { + Overload, + ControllerRestart, + CountBalance, + Maintenance, + MaintenanceBusy, + MaintenanceReader, + MaintenanceFollower, + MaintenanceRoles, + ReceiverLoss, +} + +async fn execute(scenario: Scenario) -> JournalResult { let root = tempfile::tempdir()?; let path = root.path().join("fleet-journal.sqlite"); - let profile = if restart { + let profile = if matches!( + scenario, + Scenario::ControllerRestart | Scenario::ReceiverLoss + ) { FleetProfile { controller_lease_ms: 3_000, reconcile_interval_ms: 500, @@ -152,20 +205,21 @@ async fn execute(restart: bool, count_balance: bool) -> JournalResult 2, + Scenario::MaintenanceFollower => 4, + Scenario::MaintenanceRoles => 7, + _ => 0, }; + let result = run_scenario( + scenario, + &root, + journal.clone(), + &mut nodes, + &mut boots, + profile, + ) + .await; let mut cleanup_error = None; for node in &nodes { if let Err(error) = node.shutdown().await @@ -187,7 +241,7 @@ async fn execute(restart: bool, count_balance: bool) -> JournalResult JournalResult( + scenario: Scenario, + root: &'a tempfile::TempDir, + journal: Arc, + nodes: &'a mut Vec>, + boots: &'a mut Vec, + profile: FleetProfile, +) -> FleetAdapterFuture<'a, ScenarioSummary> { + match scenario { + Scenario::CountBalance => Box::pin(balance::run(root, journal, nodes, boots, profile)), + Scenario::Maintenance => Box::pin(maintenance::run( + root, journal, nodes, boots, profile, false, + )), + Scenario::MaintenanceBusy => { + Box::pin(maintenance::run(root, journal, nodes, boots, profile, true)) + } + Scenario::MaintenanceReader => Box::pin(reader_maintenance::run( + root, journal, nodes, boots, profile, + )), + Scenario::MaintenanceFollower => Box::pin(follower_maintenance::run( + root, + journal, + nodes, + boots, + profile, + follower_maintenance::Roles::Followers, + )), + Scenario::MaintenanceRoles => Box::pin(follower_maintenance::run( + root, + journal, + nodes, + boots, + profile, + follower_maintenance::Roles::ReadersAndFollowers, + )), + Scenario::ReceiverLoss => receiver_loss::run(root, journal, nodes, boots, profile), + other => Box::pin(run( + root, + root.path().join("fleet-journal.sqlite"), + journal, + nodes, + boots, + profile, + matches!(other, Scenario::ControllerRestart), + )), + } +} + +#[cfg(test)] +pub(super) async fn commit_test_role_settlement( + journal: &SqliteJournal, + now_ms: i64, +) -> JournalResult<()> { + use cellule_host::fleet::{FleetActionAcceptance, FleetActionJournal}; + use cellule_runtime::fleet::operations::{FleetActionOutcome, FleetOutcome, MaintenanceAction}; + + let snapshot = journal.load_snapshot(scope()).await?; + let operation = snapshot + .head() + .maintenance() + .ok_or_else(|| invalid("role settlement fixture has no maintenance operation"))?; + let action = snapshot + .head() + .maintenance_action(MaintenanceAction::SettleRoles, now_ms)?; + let accepted = match journal + .accept_action(&action, operation.node(), operation.session(), now_ms) + .await? + { + FleetActionAcceptance::New(accepted) | FleetActionAcceptance::Existing { accepted, .. } => { + accepted + } + }; + let result = FleetActionOutcome { + scope: action.scope(), + action_key: action.key()?, + node: operation.node(), + session: operation.session(), + observed_at_ms: now_ms, + // This fixture tests journal ordering only; the real observer supplies + // the complete role inventory digest before SettleRoles is dispatched. + outcome: FleetOutcome::RolesSettledAt { + inventory: Digest::from_bytes([253; 32]), + head_revision: snapshot.head().revision(), + registry: snapshot.registry(), + }, + }; + journal.publish_action_result(&accepted, &result).await?; + Ok(()) +} + /// The private reference profile provisions only catalog-backed SQL writers. /// Register every boot before readiness; retain partial owners for exit cleanup. async fn initialize( @@ -252,6 +399,49 @@ async fn initialize( nodes: &mut Vec>, boots: &mut Vec, receipt_lifetime_ms: i64, +) -> JournalResult<(Arc>, HashMap)> { + initialize_inner( + root, + journal, + nodes, + boots, + receipt_lifetime_ms, + None, + false, + ) + .await +} + +#[cfg(test)] +async fn initialize_without_boot_withdrawal( + root: &tempfile::TempDir, + journal: &Arc, + nodes: &mut Vec>, + boots: &mut Vec, + receipt_lifetime_ms: i64, + unbound_index: usize, + recover_receiver: bool, +) -> JournalResult<(Arc>, HashMap)> { + initialize_inner( + root, + journal, + nodes, + boots, + receipt_lifetime_ms, + Some(unbound_index), + recover_receiver, + ) + .await +} + +async fn initialize_inner( + root: &tempfile::TempDir, + journal: &Arc, + nodes: &mut Vec>, + boots: &mut Vec, + receipt_lifetime_ms: i64, + unbound_withdrawal_index: Option, + recover_receiver: bool, ) -> JournalResult<(Arc>, HashMap)> { let application = application::compile()?; let code = *application @@ -337,6 +527,7 @@ async fn initialize( node_id(index), journal.clone(), Arc::new(adapters::Cells { + receiver_directory: recover_receiver.then(|| directory.clone()), records: records.clone(), local: index, root: root.path().into(), @@ -363,7 +554,9 @@ async fn initialize( .load(session(index), clock()?) .await? .ok_or_else(|| invalid("example original boot is absent"))?; - node.install_fleet_boot_withdrawal(directory.clone(), observed, boot, journal.clone())?; + if unbound_withdrawal_index != Some(index) { + node.install_fleet_boot_withdrawal(directory.clone(), observed, boot, journal.clone())?; + } node.start()?; } let mut acknowledged = HashMap::new(); @@ -449,9 +642,11 @@ async fn run( journal: journal.clone(), boots: boots.clone(), records: records.clone(), + reader_verifier: None, capture_sequence: std::sync::atomic::AtomicU64::new(0), lose_release_replies: restart, lost_release_replies: std::sync::atomic::AtomicUsize::new(0), + drop_closed_finalize_replies: std::sync::atomic::AtomicUsize::new(0), expired_receiver_cleanups: std::sync::atomic::AtomicUsize::new(0), }); let driver = FleetReconciler::new( @@ -620,7 +815,12 @@ async fn settle( controller_epoch: 0, expired_receiver_cleanups: 0, blockers, - final_counts: [0; 3], + final_counts: vec![0; 3], + maintenance_completed: false, + maintenance_boot_withdrawn: false, + receiver_process_closures: 0, + lost_activation_replies: 0, + routed_activation_replays: 0, }; let mut passes = Vec::new(); for pass in 0..12 { @@ -773,7 +973,8 @@ async fn verify_movement( receipt.source.query(64, 64, |_| Ok(Vec::new())).await, Err(cellule_runtime::Error::Fenced | cellule_runtime::Error::CellDraining - | cellule_runtime::Error::CellNotActive) + | cellule_runtime::Error::CellNotActive + | cellule_runtime::Error::RuntimeClosed) ) { return Err(invalid("old source handle still served after movement")); } diff --git a/crates/cellule-host/minion/scenario/reader_tests/evacuation/transport.rs b/crates/cellule-host/minion/scenario/native_peers.rs similarity index 82% rename from crates/cellule-host/minion/scenario/reader_tests/evacuation/transport.rs rename to crates/cellule-host/minion/scenario/native_peers.rs index 98099152..df061e18 100644 --- a/crates/cellule-host/minion/scenario/reader_tests/evacuation/transport.rs +++ b/crates/cellule-host/minion/scenario/native_peers.rs @@ -1,16 +1,18 @@ +//! Trusted local routing through the canonical signed peer verifier/dispatcher. use super::*; +use cellule_host::read_replicas::ReadReplicaManager; use cellule_runtime::peer::{ PeerAuthorizer, PeerDispatcher, PeerRoundTrip, ResidentPeerCellResolver, VerifiedPeerRequest, - wire, }; -use std::{ - future::Future, - pin::Pin, - sync::{ - Mutex, - atomic::{AtomicUsize, Ordering}, - }, +use cellule_runtime::{Error, node::NodeDirectory}; +use ed25519_dalek::SigningKey; +#[cfg(test)] +use std::sync::{ + Mutex, + atomic::{AtomicUsize, Ordering}, }; +use std::{future::Future, pin::Pin}; +#[cfg(test)] use tokio::sync::oneshot; #[derive(Clone)] @@ -18,9 +20,12 @@ pub(super) struct NativePeers { directory: NodeDirectory, origin: usize, dispatchers: Vec>, + #[cfg(test)] pub(super) probes: Arc, + #[cfg(test)] pause: Arc>>, } +#[cfg(test)] struct Pause { number: usize, captured: oneshot::Sender<()>, @@ -41,7 +46,7 @@ impl PeerAuthorizer for Authorizer { .any(|action| action == "replica-maintenance") { return Err(Error::PeerAuthorization( - "maintenance fixture principal differs", + "maintenance example principal differs", )); } Ok(()) @@ -89,10 +94,13 @@ impl NativePeers { directory, origin, dispatchers, + #[cfg(test)] probes: Arc::new(AtomicUsize::new(0)), + #[cfg(test)] pause: Arc::new(Mutex::new(None)), } } + #[cfg(test)] pub(super) fn pause_probe( &self, number: usize, @@ -115,7 +123,7 @@ impl PeerRoundTrip for NativePeers { _: Vec, _: u32, ) -> Pin>> + Send + 'static>> { - Box::pin(async { Err(Error::Peer("fixture requires explicit replica routing")) }) + Box::pin(async { Err(Error::Peer("example requires explicit replica routing")) }) } fn send_to_node( &self, @@ -126,7 +134,7 @@ impl PeerRoundTrip for NativePeers { ) -> Pin>> + Send + 'static>> { let peers = self.clone(); Box::pin(async move { - let index = (0..3) + let index = (0..peers.dispatchers.len()) .find(|index| node.node() == node_id(*index) && node.session() == session(*index)) .ok_or(Error::Fenced)?; let now = clock()?; @@ -151,16 +159,24 @@ impl PeerRoundTrip for NativePeers { if verified.target() != &target { return Err(Error::Fenced); } + #[cfg(test)] let status = matches!( verified.operation(), - Some(wire::peer_request::Operation::Read(wire::ReadRequest { - operation: Some(wire::read_request::Operation::ReplicaStatus(true)), - .. - })) + Some(cellule_runtime::peer::wire::peer_request::Operation::Read( + cellule_runtime::peer::wire::ReadRequest { + operation: Some( + cellule_runtime::peer::wire::read_request::Operation::ReplicaStatus( + true + ) + ), + .. + } + )) ); let reply = peers.dispatchers[index] .dispatch_bytes(&verified, clock()?) .await?; + #[cfg(test)] let pause = if status { let number = peers.probes.fetch_add(1, Ordering::SeqCst) + 1; let mut pending = peers.pause.lock().unwrap(); @@ -172,6 +188,7 @@ impl PeerRoundTrip for NativePeers { } else { None }; + #[cfg(test)] if let Some(pause) = pause { let _ = pause.captured.send(()); let _ = pause.resume.await; diff --git a/crates/cellule-host/minion/scenario/observation/mod.rs b/crates/cellule-host/minion/scenario/observation/mod.rs index e0236e36..86163b35 100644 --- a/crates/cellule-host/minion/scenario/observation/mod.rs +++ b/crates/cellule-host/minion/scenario/observation/mod.rs @@ -1,14 +1,18 @@ -//! Full observation of this example's closed, writer-only construction profile. -//! Role-enabled applications need their producer/native/policy collectors. +//! Full bounded observation of the reference fleet's native role graph. +//! Policy checks remain separate evidence and missing checks stay blocking. use super::*; use cellule_host::fleet::{ - FleetFollowerReferences, FleetMaintenanceEnrollments, FleetNodeInventory, - FleetNodeInventoryScan, FleetNodeSnapshot, FleetObservation, FleetOwnedCell, FleetRoleCoverage, - FleetRoster, FleetSnapshotNativePage, FleetSnapshotRequest, FleetSnapshotSubject, + FleetFailedBootProcessRequest, FleetFailedBootProcesses, FleetFailedBootRetirement, + FleetFollowerEvacuationVerifier, FleetFollowerReferences, FleetMaintenanceEnrollments, + FleetNodeInventory, FleetNodeInventoryScan, FleetNodeSnapshot, FleetObservation, + FleetOwnedCell, FleetRoleCoverage, FleetRoster, FleetSnapshotNativePage, FleetSnapshotRequest, + FleetSnapshotSubject, }; use cellule_runtime::control::{Control, ControlState}; -use cellule_runtime::fleet::operations::{EnrollmentRole, EnrollmentStatus, PublishedPosition}; +use cellule_runtime::fleet::operations::{ + DrainBlocker, EnrollmentRole, EnrollmentStatus, PublishedPosition, +}; use cellule_runtime::node::NodeAdvertisement; use std::collections::HashSet; use std::sync::atomic::Ordering; @@ -27,14 +31,14 @@ pub(super) async fn complete_counts( fleet: &adapters::LocalFleet, roster: &FleetRoster, deadline: Instant, -) -> JournalResult> { +) -> JournalResult>> { let capture = collect(fleet, roster, deadline).await?; if !capture.complete { return Ok(None); } - let mut counts = [0; 3]; + let mut counts = vec![0; fleet.nodes.len()]; for owned in capture.cells { - let index = (0..3) + let index = (0..fleet.nodes.len()) .find(|n| node_id(*n) == owned.node && session(*n) == owned.session) .ok_or_else(|| invalid("example count endpoint differs"))?; counts[index] += 1; @@ -46,14 +50,88 @@ pub(super) async fn observe( fleet: &adapters::LocalFleet, roster: &FleetRoster, deadline: Instant, +) -> JournalResult { + observe_inner(fleet, roster, deadline, None).await +} + +pub(super) async fn observe_with_failed_boot_closure( + fleet: &adapters::LocalFleet, + roster: &FleetRoster, + deadline: Instant, + request: &FleetFailedBootProcessRequest, + processes: &dyn FleetFailedBootProcesses, + claimant: SessionId, +) -> JournalResult { + observe_inner( + fleet, + roster, + deadline, + Some((request, processes, claimant)), + ) + .await +} + +async fn observe_inner( + fleet: &adapters::LocalFleet, + roster: &FleetRoster, + deadline: Instant, + failed_boot: Option<( + &FleetFailedBootProcessRequest, + &dyn FleetFailedBootProcesses, + SessionId, + )>, ) -> JournalResult { let capture = collect(fleet, roster, deadline).await?; + let mut reader_checks = Vec::new(); + let mut follower_checks = Vec::new(); + let mut finished = capture.finished; + if let Some(original) = capture.maintenance_enrollments.as_ref() { + if let Some(verifier) = &fleet.reader_verifier { + reader_checks = verifier + .collect_maintenance(fleet.journal.as_ref(), original, roster, deadline, clock) + .await?; + } + let follower_verifier = FleetFollowerEvacuationVerifier::new( + fleet.boots[0].directory.clone(), + Arc::new(adapters::LocalSnapshots::new(fleet.nodes.clone())), + ); + follower_checks = follower_verifier + .collect_maintenance(fleet.journal.as_ref(), original, roster, deadline, clock) + .await?; + roster.confirm(fleet.journal.as_ref(), deadline).await?; + finished = clock()?; + } + let failed_boot_closures = if let Some((request, processes, claimant)) = failed_boot { + let closure = FleetFailedBootRetirement::capture_retained( + fleet.journal.as_ref(), + &fleet.boots[0].directory, + roster, + request, + claimant, + deadline, + clock, + ) + .await? + .confirm( + fleet.journal.as_ref(), + &fleet.boots[0].directory, + processes, + claimant, + deadline, + clock, + ) + .await?; + finished = clock()?; + Some(vec![closure]) + } else { + None + }; let observation = FleetObservation::new( scope(), roster.snapshot().registry(), roster.snapshot().registry().revision(), capture.started, - capture.finished, + finished, capture.complete, capture.nodes, capture.cells, @@ -62,12 +140,18 @@ pub(super) async fn observe( Some(coverage) => observation.with_role_coverage(coverage)?, None => observation, }; - Ok(match capture.maintenance_enrollments { + let observation = match failed_boot_closures { + Some(closures) => observation.with_failed_boot_closures(closures)?, + None => observation, + }; + let observation = match capture.maintenance_enrollments { Some(original) => observation + .with_role_evacuations(reader_checks, follower_checks)? .with_maintenance_enrollments(original)? - .check_maintenance_policies(roster, capture.finished)?, + .check_maintenance_policies(roster, finished)?, None => observation, - }) + }; + Ok(observation) } async fn page( @@ -150,24 +234,49 @@ async fn collect( } else { None }; - if fleet.nodes.len() != 3 || fleet.boots.len() != 3 || fleet.records.len() != CELL_COUNT { + if fleet.nodes.len() < 3 || fleet.boots.len() != fleet.nodes.len() || fleet.records.is_empty() { return Err(invalid("example construction profile differs")); } let directory = &fleet.boots[0].directory; - let expected_sessions = (0..3).map(session).collect::>(); + let mut expected_sessions = roster + .enrollments() + .iter() + .filter(|record| { + record.unresolved() + && record.status() == EnrollmentStatus::Established + && matches!(record.spec().role, EnrollmentRole::Node { .. }) + }) + .map(|record| record.spec().target.session) + .collect::>(); + expected_sessions.sort_by_key(|boot| *boot.as_bytes()); + let node_enrollments_complete = roster.enrollments().iter().all(|record| { + !matches!(record.spec().role, EnrollmentRole::Node { .. }) + || record.status() != EnrollmentStatus::Pending + }); + let mut active_indices = Vec::with_capacity(expected_sessions.len()); + for record in roster.enrollments().iter().filter(|record| { + record.unresolved() + && record.status() == EnrollmentStatus::Established + && matches!(record.spec().role, EnrollmentRole::Node { .. }) + }) { + let endpoint = record.spec().target; + if let Some(index) = (0..fleet.nodes.len()) + .find(|index| node_id(*index) == endpoint.node && session(*index) == endpoint.session) + { + active_indices.push(index); + } + } + active_indices.sort_unstable(); let mut advertised = directory.advertised_sessions(started, 128).await?; advertised.sort_by_key(|boot| *boot.as_bytes()); let mut complete = advertised == expected_sessions - && roster.enrollments().iter().all(|record| { - matches!(record.spec().role, EnrollmentRole::Node { .. }) - && record.status() != EnrollmentStatus::Pending - }); + && active_indices.len() == expected_sessions.len() + && node_enrollments_complete; let mut nodes = Vec::new(); let mut cells = Vec::new(); let mut seen = HashSet::new(); let mut inventories: Vec> = Vec::new(); - let mut references = Vec::new(); - for index in 0..3 { + for index in active_indices.iter().copied() { let mut scan = FleetNodeInventoryScan::new(roster, node_id(index), session(index))?; let mut stable = true; while let Some(subject) = scan.next_subject()? { @@ -200,41 +309,32 @@ async fn collect( cells.extend(scan.cells().iter().cloned()); let inventory = if stable { let inventory = scan.finish()?; - let bindings = inventory.bindings(); - // Absence requires this bootstrapped closed writer composition, - // native traversal and unexpected directory discovery together. - complete &= bindings.managed_startup - && !bindings.readers - && !bindings.follower_store - && !bindings.follower_producer - && !bindings.durability_supervisor - && inventory.node_log().is_none() - && inventory.transitioning_cells().is_empty() - && inventory.validate_enrollments(roster).is_ok(); + // Global role coverage below applies enrollment validation with + // the exact cross-node proof for Established follower lanes that + // have not received their first append yet. + complete &= + inventory.bindings().managed_startup && inventory.transitioning_cells().is_empty(); Some(inventory) } else { None }; inventories.push(inventory); // Include expired and fenced leader obligations; live discovery alone - // could hide a follower role left by a failed boot. - let logs = FleetFollowerReferences::collect( - directory, - roster, - node_id(index), - 128, - deadline, - clock, - ) - .await?; - complete &= logs.entries().is_empty() && logs.validate_enrollments(roster).is_ok(); - references.push(logs); + // could hide a follower role left by a failed boot. The graph may be + // nonempty; its exact match is checked after every native recheck. nodes.push( fleet.boots[index] .refresh_capacity(index, fleet.journal.as_ref(), deadline) .await?, ); } + let members = (0..fleet.nodes.len()).map(node_id).collect::>(); + let mut references = + FleetFollowerReferences::collect_all(directory, roster, &members, 128, deadline, clock) + .await?; + for logs in &references { + complete &= logs.validate_enrollments(roster).is_ok(); + } let mut authority = HashMap::new(); for (cell, record) in fleet.records.iter() { let current = record @@ -245,11 +345,15 @@ async fn collect( authority.insert(*cell, current.value().clone()); } cells.retain(|owned| { - let matches = authority - .get(&owned.observation.target.cell_id()) - .is_some_and(|current| matches_authority(owned, current)); - complete &= matches; - matches + let Some(current) = authority.get(&owned.observation.target.cell_id()) else { + complete = false; + return false; + }; + let exact = matches_authority(owned, current, fleet.nodes.len()); + complete &= exact; + exact + || (maintenance_candidate(roster, owned, started) + && matches_writer_authority(owned, current, fleet.nodes.len())) }); for (cell, current) in &authority { complete &= match current.state { @@ -260,12 +364,13 @@ async fn collect( ControlState::Recovering => false, }; } - // Recheck exact authority after the full scan. A concurrent publication or - // takeover invalidates that Cell's planning row, not just count completeness. - complete &= recheck_authority(fleet, &authority, &mut cells).await?; + // Root changes invalidate complete counts and ordinary movement demand. + // Explicit maintenance may retain this exact writer's peak envelope; the + // prepared action still joins publication and obtains its final root proof. + complete &= recheck_authority(fleet, roster, &authority, &mut cells).await?; // Recheck every role category after *all* authority and membership reads. // Stable local-only traversals cannot supply this fleet-wide interval. - for (index, inventory) in inventories.iter_mut().enumerate() { + for (index, inventory) in active_indices.iter().copied().zip(inventories.iter_mut()) { let Some(inventory) = inventory else { let response = page( fleet, @@ -278,7 +383,7 @@ async fn collect( let FleetSnapshotNativePage::Cells(actors) = response.page() else { return Err(invalid("example repeated actor page category differs")); }; - retain_unchanged_writers(&mut cells, index, actors.entries()); + retain_unchanged_writers(&mut cells, roster, index, actors.entries(), clock()?); continue; }; let mut recheck = inventory.recheck(); @@ -293,7 +398,13 @@ async fn collect( FleetSnapshotNativePage::Cells(actors) => { // Count planning stops on changed topology. Keep only // independently unchanged writer rows for pressure relief. - retain_unchanged_writers(&mut cells, index, actors.entries()); + retain_unchanged_writers( + &mut cells, + roster, + index, + actors.entries(), + clock()?, + ); } FleetSnapshotNativePage::Host => { cells.retain(|owned| owned.node != node_id(index)); @@ -305,14 +416,21 @@ async fn collect( } complete &= recheck.finish().is_ok(); } - for logs in &mut references { - match logs.recheck(directory, roster, 128, deadline, clock).await { - Ok(()) => {} - Err(cellule_runtime::Error::Node("authoritative follower inventory changed")) => { - complete = false; - } - Err(source) => return Err(source.into()), + match FleetFollowerReferences::recheck_all( + &mut references, + directory, + roster, + 128, + deadline, + clock, + ) + .await + { + Ok(()) => {} + Err(cellule_runtime::Error::Node("authoritative follower inventory changed")) => { + complete = false; } + Err(source) => return Err(source.into()), } let mut after = directory.advertised_sessions(clock()?, 128).await?; after.sort_by_key(|boot| *boot.as_bytes()); @@ -334,7 +452,7 @@ async fn collect( None }; roster.confirm(fleet.journal.as_ref(), deadline).await?; - Ok(Capture { + let capture = Capture { started, finished: clock()?, complete, @@ -342,11 +460,13 @@ async fn collect( cells, role_coverage, maintenance_enrollments, - }) + }; + Ok(capture) } async fn recheck_authority( fleet: &adapters::LocalFleet, + roster: &FleetRoster, authority: &HashMap, cells: &mut Vec, ) -> JournalResult { @@ -361,23 +481,14 @@ async fn recheck_authority( .await? .ok_or_else(|| invalid("example authority disappeared during capture"))?; if current.value() != original { - let value = current.value(); - let protected_same = value.cell == original.cell - && value.incarnation == original.incarnation - && value.epoch == original.epoch - && value.state == original.state - && value.owner == original.owner - && value.root == original.root - && value.recovery == original.recovery - && value.code == original.code - && value.schema == original.schema - && value.next_due_ms == original.next_due_ms; - eprintln!( - "FLEET_CAPTURE authority_changed cell={cell:?} revision={}..{} progress={}..{} protected_fields_unchanged={protected_same}", - original.revision, value.revision, original.progress, value.progress, - ); unchanged = false; - cells.retain(|row| row.observation.target.cell_id() != *cell); + let now = clock()?; + cells.retain(|row| { + row.observation.target.cell_id() != *cell + || (maintenance_candidate(roster, row, now) + && matches_writer_authority(row, original, fleet.nodes.len()) + && matches_writer_authority(row, current.value(), fleet.nodes.len())) + }); } } Ok(unchanged) @@ -385,8 +496,10 @@ async fn recheck_authority( fn retain_unchanged_writers( cells: &mut Vec, + roster: &FleetRoster, index: usize, entries: &[CellInventoryEntry], + now: i64, ) { cells.retain(|owned| { owned.node != node_id(index) @@ -400,25 +513,58 @@ fn retain_unchanged_writers( && current.incarnation == original.incarnation && current.code == original.code && current.schema == original.schema - && current.position == original.position - && current.cost == original.cost - && current.blockers == original.blockers + && current.role == original.role + && current.owner_fence == original.owner_fence + && ((current.position == original.position + && current.cost == original.cost + && current.blockers == original.blockers) + || (maintenance_candidate(roster, owned, now) + && current.maintenance_cost == original.maintenance_cost + && current.blockers.iter().all(|blocker| { + matches!( + blocker, + DrainBlocker::BusyExecution + | DrainBlocker::ExternalLease + | DrainBlocker::PendingPublication + | DrainBlocker::UnknownInventory + ) + }))) }) }); } -fn matches_authority(owned: &FleetOwnedCell, current: &Control) -> bool { +fn maintenance_candidate(roster: &FleetRoster, owned: &FleetOwnedCell, now: i64) -> bool { + roster + .snapshot() + .head() + .maintenance() + .is_some_and(|operation| { + operation.phase() == cellule_runtime::fleet::operations::MaintenancePhase::Evacuating + && operation.node() == owned.node + && operation.session() == owned.session + && now < operation.deadline_ms() + }) +} + +fn matches_writer_authority(owned: &FleetOwnedCell, current: &Control, node_count: usize) -> bool { let row = &owned.observation; current.cell == row.target.cell_id() && current.incarnation == row.incarnation + && current.owner_fence() == row.owner_fence && current.code == row.code && current.schema == row.schema && current.state == ControlState::Serving && current.owner.as_ref().is_some_and(|owner| { owner.session == owned.session - && (0..3).any(|n| node_id(n) == owned.node && owner == &super::owner(n)) + && (0..node_count).any(|n| node_id(n) == owned.node && owner == &super::owner(n)) }) && current.recovery.is_none() + && current.root.is_some() +} + +fn matches_authority(owned: &FleetOwnedCell, current: &Control, node_count: usize) -> bool { + let row = &owned.observation; + matches_writer_authority(owned, current, node_count) && current.root.as_ref().is_some_and(|root| { row.position == Some(PublishedPosition { diff --git a/crates/cellule-host/minion/scenario/observation/tests/inventory.rs b/crates/cellule-host/minion/scenario/observation/tests/inventory.rs index 42ecbe44..304a4c90 100644 --- a/crates/cellule-host/minion/scenario/observation/tests/inventory.rs +++ b/crates/cellule-host/minion/scenario/observation/tests/inventory.rs @@ -275,6 +275,7 @@ async fn canonical_owner_renewal_invalidates_the_exact_authority_interval() { assert_eq!(renewed.value().root, original.value().root); let unchanged = recheck_authority( &fixture.fleet, + &roster, &HashMap::from([(cell, original.value().clone())]), &mut cells, ) @@ -290,3 +291,103 @@ async fn canonical_owner_renewal_invalidates_the_exact_authority_interval() { drop(inventory); fixture.close().await; } + +#[tokio::test] +async fn same_writer_publication_retains_only_explicit_maintenance_demand() { + use cellule_runtime::fleet::operations::MaintenanceEvent; + let fixture = Fixture::new().await; + let roster = fixture.roster().await; + let inventory = capture_node(&fixture, &roster, 0).await; + let mut cells = inventory.cells().to_vec(); + let cell = cells[0].observation.target.cell_id(); + let record = fixture.fleet.records.get(&cell).unwrap(); + let original = record.authority.load(cell).await.unwrap().unwrap(); + let handle = fixture.fleet.nodes[0] + .runtime() + .local_handle(record.catalog.clone(), &original) + .await + .unwrap() + .unwrap(); + let now = clock().unwrap(); + let operation = MaintenanceOperation::new( + OperationId::from_bytes([222; 16]).unwrap(), + Digest::from_bytes([223; 32]), + node_id(0), + session(0), + 2, + now, + now + 60_000, + ) + .unwrap(); + let mut snapshot = roster.snapshot().clone(); + for transition in [ + JournalTransition::BeginMaintenance(operation), + JournalTransition::Maintenance(MaintenanceEvent::Cordoned), + JournalTransition::Maintenance(MaintenanceEvent::BeginEvacuation), + ] { + snapshot = fixture + .fleet + .journal + .compare_exchange( + &snapshot, + snapshot.head().controller().unwrap().epoch, + clock().unwrap(), + &transition, + ) + .await + .unwrap(); + } + let roster = fixture.roster().await; + handle + .execute( + MutationIdentity { + request_id: RequestId::from_bytes([239; 16]), + issued_at_ms: now, + expires_at_ms: now + 60_000, + }, + Digest::from_bytes([239; 32]), + now, + 64, + 64, + |transaction| { + transaction.execute("UPDATE counter SET value = value + 1", [])?; + Ok(HandlerOutcome::Success(vec![239])) + }, + ) + .await + .unwrap(); + let published = record.authority.load(cell).await.unwrap().unwrap(); + assert_eq!( + published.value().owner_fence(), + original.value().owner_fence() + ); + assert_ne!(published.value().root, original.value().root); + let unchanged = recheck_authority( + &fixture.fleet, + &roster, + &HashMap::from([(cell, original.value().clone())]), + &mut cells, + ) + .await + .unwrap(); + assert!(!unchanged); + // Incomplete counts stay incomplete. The original activation is still a + // candidate for the separately authorized busy-maintenance action. + assert_eq!(cells.len(), CELL_COUNT); + let original_row = cells + .iter() + .find(|row| row.observation.target.cell_id() == cell) + .unwrap(); + assert!(maintenance_candidate(&roster, original_row, now)); + assert!(!maintenance_candidate(&roster, original_row, now + 60_000)); + let mut foreign = original_row.clone(); + foreign.session = session(1); + assert!(!maintenance_candidate(&roster, &foreign, now)); + assert!(!matches_writer_authority(&foreign, published.value(), 3)); + foreign = original_row.clone(); + foreign.observation.owner_fence.epoch += 1; + assert!(!matches_writer_authority(&foreign, published.value(), 3)); + drop(inventory); + drop(roster); + fixture.close().await; +} diff --git a/crates/cellule-host/minion/scenario/observation/tests/mod.rs b/crates/cellule-host/minion/scenario/observation/tests/mod.rs index 9b6a5877..ca364118 100644 --- a/crates/cellule-host/minion/scenario/observation/tests/mod.rs +++ b/crates/cellule-host/minion/scenario/observation/tests/mod.rs @@ -42,10 +42,12 @@ impl Fixture { nodes, boots, records, + reader_verifier: None, journal, capture_sequence: std::sync::atomic::AtomicU64::new(0), lose_release_replies: false, lost_release_replies: std::sync::atomic::AtomicUsize::new(0), + drop_closed_finalize_replies: std::sync::atomic::AtomicUsize::new(0), expired_receiver_cleanups: std::sync::atomic::AtomicUsize::new(0), }, } @@ -522,7 +524,13 @@ async fn one_canonical_release_keeps_other_native_writers_available_for_pressure assert_ne!(after_page.topology(), topology); assert_eq!(after_page.owned_cells(), CELL_COUNT - 1); let mut candidates = original.cells; - retain_unchanged_writers(&mut candidates, 0, after_page.entries()); + retain_unchanged_writers( + &mut candidates, + &roster, + 0, + after_page.entries(), + clock().unwrap(), + ); assert_eq!(candidates.len(), CELL_COUNT - 1); assert!( candidates diff --git a/crates/cellule-host/minion/scenario/reader_maintenance/mod.rs b/crates/cellule-host/minion/scenario/reader_maintenance/mod.rs new file mode 100644 index 00000000..08e7aa65 --- /dev/null +++ b/crates/cellule-host/minion/scenario/reader_maintenance/mod.rs @@ -0,0 +1,332 @@ +//! Reader-only maintenance using managed boots and the public reconciler. +use super::*; +use cellule_host::{ + fleet::{FleetReaderEvacuationPublication, FleetReaderEvacuationVerifier, FleetRoster}, + read_replicas::ReadReplicaManager, +}; +use cellule_runtime::{ + client::CellDescription, + fleet::operations::{EnrollmentStatus, MaintenancePhase}, + node::NodeDirectory, + peer::{PeerPrincipal, PeerReplicaResolver, PeerSigner, ReplicaPeerClient}, + read_policy::ReadPolicyStore, +}; +use ed25519_dalek::SigningKey; +use std::sync::atomic::{AtomicU64, AtomicUsize}; + +mod setup; +#[cfg(test)] +mod tests; + +struct Inputs { + records: Arc>, + managers: Vec, + target: CellTarget, + description: CellDescription, + peer: ReplicaPeerClient, + verifier: FleetReaderEvacuationVerifier, + acknowledged: Acknowledged, +} +fn deadline() -> Instant { + Instant::now() + Duration::from_secs(8) +} + +pub(super) async fn run( + root: &tempfile::TempDir, + journal: Arc, + nodes: &mut Vec>, + boots: &mut Vec, + profile: FleetProfile, +) -> JournalResult { + let inputs = setup::initialize(root, &journal, nodes, boots).await?; + let original_reader = inputs.managers[1].resolve(inputs.target.clone()).await?; + let completion = inputs.managers[1] + .enrollment_completion(inputs.target.cell_id()) + .await? + .ok_or_else(|| invalid("reader enrollment completion is missing"))?; + let original = journal + .load_enrollment(scope(), completion.spec.key()?) + .await? + .ok_or_else(|| invalid("original managed reader is missing"))?; + if original.status() != EnrollmentStatus::Established { + return Err(invalid("original reader is not established")); + } + let fleet = Arc::new(adapters::LocalFleet { + nodes: nodes.clone(), + boots: boots.clone(), + records: inputs.records.clone(), + journal: journal.clone(), + reader_verifier: Some(inputs.verifier.clone()), + capture_sequence: AtomicU64::new(0), + lose_release_replies: false, + lost_release_replies: AtomicUsize::new(0), + drop_closed_finalize_replies: AtomicUsize::new(0), + expired_receiver_cleanups: AtomicUsize::new(0), + }); + let claimant = SessionId::from_bytes([206; 16]); + let now = clock()?; + let snapshot = journal.load_snapshot(scope()).await?; + let claimed = journal + .claim_controller(scope(), snapshot.head().revision(), claimant, now) + .await?; + let epoch = claimed + .head() + .controller() + .ok_or_else(|| invalid("reader maintenance controller is absent"))? + .epoch; + let operation = MaintenanceOperation::new( + OperationId::from_bytes([83; 16])?, + Digest::from_bytes([84; 32]), + node_id(1), + session(1), + 2, + now, + now.checked_add(60_000) + .ok_or_else(|| invalid("reader maintenance deadline overflow"))?, + )?; + journal + .compare_exchange( + &claimed, + epoch, + now, + &JournalTransition::BeginMaintenance(operation.clone()), + ) + .await?; + let driver = FleetReconciler::new( + scope(), + claimant, + profile, + journal.clone(), + fleet.clone(), + fleet.clone(), + )?; + let wall_deadline = Instant::now() + Duration::from_secs(60); + let mut blockers = Vec::new(); + // The driver owns cordon and evacuation transitions. No fixture writes a + // ready-to-close row or an invented native role proof. + loop { + let report = driver.reconcile_once(clock, deadline()).await?; + checked_report(&report)?; + for blocker in report.blockers { + if !blockers.contains(&blocker) { + blockers.push(blocker); + } + } + let phase = report + .snapshot + .head() + .maintenance() + .ok_or_else(|| invalid("reader maintenance operation disappeared"))? + .phase(); + if phase == MaintenancePhase::Evacuating { + break; + } + if Instant::now() >= wall_deadline { + return Err(invalid("reader maintenance did not enter evacuation")); + } + } + boots[1] + .refresh_capacity(1, journal.as_ref(), deadline()) + .await?; + if nodes[1].state() == NodeState::Stopped || !nodes[1].is_management_ready() { + return Err(invalid("reader source stopped before replacement")); + } + // The missing replacement remains an explicit policy blocker and cannot + // erase the original ready reader or authorize native Finalize. + let blocked = driver.reconcile_once(clock, deadline()).await?; + checked_report(&blocked)?; + if blocked + .snapshot + .head() + .maintenance() + .map(MaintenanceOperation::phase) + != Some(MaintenancePhase::Evacuating) + || journal + .load_enrollment(scope(), original.spec().key()?) + .await? + .as_ref() + != Some(&original) + { + return Err(invalid("reader maintenance bypassed replacement policy")); + } + for blocker in blocked.blockers { + if !blockers.contains(&blocker) { + blockers.push(blocker); + } + } + if original_reader + .query::(None, 0) + .await? + .output + != 29 + { + return Err(invalid( + "original reader lost acknowledged state before evacuation", + )); + } + let spare = boots[2] + .refresh_capacity(2, journal.as_ref(), deadline()) + .await?; + // Reader discovery deliberately uses its bounded membership cache. Wait + // for canonical selection to observe the signed cordon before opening a + // replacement; a fresh heartbeat does not invalidate that read window. + tokio::time::timeout_at(deadline(), async { + loop { + let selected = boots[0] + .directory + .select_readers( + inputs.target.cell_id(), + session(0), + inputs.description.code, + 1, + clock()?, + 128, + ) + .await?; + if selected.len() == 1 && selected[0].session() == session(2) { + return Ok::<_, JournalError>(()); + } + tokio::time::sleep(Duration::from_millis(10)).await; + } + }) + .await??; + inputs + .peer + .activate( + &inputs.target, + &boots[0].directory, + spare, + inputs.description, + ) + .await?; + let current = journal.load_snapshot(scope()).await?; + let operation = current + .head() + .maintenance() + .ok_or_else(|| invalid("reader maintenance operation disappeared"))?; + let capture = inputs.managers[1] + .evacuate(&original, operation, &inputs.peer, deadline()) + .await?; + let publication = FleetReaderEvacuationPublication::publish( + &capture, + journal.as_ref(), + &inputs.verifier, + deadline(), + clock, + ) + .await?; + let record = publication.record()?; + let completed = loop { + let report = driver.reconcile_once(clock, deadline()).await?; + checked_report(&report)?; + for blocker in report.blockers { + if !blockers.contains(&blocker) { + blockers.push(blocker); + } + } + if report.snapshot.head().maintenance().is_some_and(|current| { + current.id() == operation.id() && current.phase() == MaintenancePhase::Completed + }) { + break report.snapshot; + } + if Instant::now() >= wall_deadline { + return Err(invalid( + "reader maintenance did not complete before its deadline", + )); + } + }; + let evidence = completed + .head() + .maintenance() + .and_then(MaintenanceOperation::drain_evidence) + .ok_or_else(|| invalid("reader maintenance completion evidence is missing"))?; + if !evidence.relocated + || !evidence.readers_settled + || !evidence.followers_settled + || evidence.remaining_cells != 0 + || evidence.unresolved_attempts != 0 + || !completed.head().attempts().is_empty() + || nodes[1].state() != NodeState::Stopped + || !boots[1].directory.is_withdrawn(session(1)).await? + { + return Err(invalid( + "reader maintenance completed without native shutdown and withdrawal", + )); + } + let retired = journal + .load_enrollment(scope(), original.spec().key()?) + .await? + .ok_or_else(|| invalid("original reader history is missing"))?; + if retired != *record.retired() + || retired.status() != EnrollmentStatus::Retired + || original_reader + .query::(None, 0) + .await + .is_ok() + { + return Err(invalid("original reader was not joined and retired")); + } + let replacement = inputs.managers[2].resolve(inputs.target.clone()).await?; + if replacement + .query::(Some(capture.minimum()), 0) + .await? + .output + != 29 + { + return Err(invalid("replacement reader lost acknowledged state")); + } + // The original writer is unchanged; verify the durable request record too, + // rather than equating successful reader queries with mutation durability. + let acknowledged = &inputs.acknowledged; + if acknowledged + .source + .resolve(acknowledged.identity, acknowledged.digest, clock()?, 64) + .await? + != Resolution::Committed(acknowledged.outcome.clone()) + { + return Err(invalid( + "reader evacuation changed the original mutation receipt", + )); + } + let roster = FleetRoster::collect(journal.as_ref(), &completed, deadline()).await?; + let final_counts = observation::complete_counts(&fleet, &roster, deadline()) + .await? + .ok_or_else(|| invalid("reader maintenance final observation is incomplete"))?; + if final_counts != [1, 0, 0] { + return Err(invalid("reader maintenance changed writer ownership")); + } + Ok(ScenarioSummary { + released: 0, + activated: 0, + retired: 0, + receipt_checks: 2, + max_inflight: 0, + max_restore_bytes: 0, + joined_nodes: 0, + boot_retirements: 0, + receiver_nodes: 1, + lost_release_replies: 0, + controller_epoch: completed + .head() + .controller() + .ok_or_else(|| invalid("controller is absent"))? + .epoch, + expired_receiver_cleanups: 0, + blockers, + final_counts, + maintenance_completed: true, + maintenance_boot_withdrawn: true, + receiver_process_closures: 0, + lost_activation_replies: 0, + routed_activation_replays: 0, + }) +} +fn checked_report(report: &cellule_host::fleet::FleetReconcileReport) -> JournalResult<()> { + if let Some(failure) = report.failures.first() { + return Err(Box::new(failure.error.clone())); + } + if let Some(failure) = &report.maintenance_failure { + return Err(Box::new(failure.clone())); + } + Ok(()) +} diff --git a/crates/cellule-host/minion/scenario/reader_maintenance/setup.rs b/crates/cellule-host/minion/scenario/reader_maintenance/setup.rs new file mode 100644 index 00000000..f1013ff6 --- /dev/null +++ b/crates/cellule-host/minion/scenario/reader_maintenance/setup.rs @@ -0,0 +1,215 @@ +//! Register every partially constructed owner before fallible startup work. +use super::*; + +pub(super) async fn initialize( + root: &tempfile::TempDir, + journal: &Arc, + nodes: &mut Vec>, + boots: &mut Vec, +) -> JournalResult { + let app = application::compile()?; + let code = *app + .registry() + .module_digests() + .first() + .ok_or_else(|| invalid("reader maintenance module is absent"))?; + let layout = CellStorageLayout::new( + Store::new(Arc::new(InMemory::new())), + ObjectPath::from("reader-maintenance-cells"), + [3; 16], + ); + let directory = NodeDirectory::new( + layout.clone(), + scope().fleet, + Digest::from_bytes([31; 32]), + app.registry().release_digest(), + ); + let limits = Limits { + max_database_bytes: 64 << 20, + max_capture_bytes: 16 << 20, + ..Limits::default() + }; + let target = CellTarget::new( + TenantId::from_bytes([1; 16]), + scope().application, + application::NAMESPACE, + &[1], + )?; + let incarnation = IncarnationId::from_bytes([1; 16]); + let catalog = CellCatalog::new(layout.clone(), target.tenant()) + .provision(CatalogEntry::new(&target, CatalogRole::Sql, code, 1)?) + .await?; + let authority = CellAuthority::new(layout.clone()); + let initial = authority + .create_initial(&catalog, incarnation, owner(0)) + .await?; + let replica = CellReplica::new( + layout.clone(), + *target.cell_id().as_bytes(), + *incarnation.as_bytes(), + limits, + )?; + let records = Arc::new(HashMap::from([( + target.cell_id(), + Record { + target: target.clone(), + incarnation, + catalog: catalog.clone(), + replica: replica.clone(), + authority: authority.clone(), + }, + )])); + let mut managers = Vec::new(); + for index in 0..3 { + let intent = journal + .register_initial_intent(&NodeIntent::initial( + scope(), + node_id(index), + session(index), + )?) + .await?; + let node = Arc::new( + CellNodeBuilder::new(app.clone()) + .with_runtime( + SqlWorkerPool::new(2, 8)?.with_native_memory_limit(128 << 20)?, + 16 << 20, + ) + .with_replica_host(Host::default().with_local_disk_budget(DiskBudget::new(8 << 30))) + .with_session(session(index)) + .with_fleet_startup_intent(intent.clone()) + .build()?, + ); + nodes.push(node.clone()); + node.install_task_group(CancellationToken::new(), CancellationToken::new())?; + node.install_fleet_actions( + scope(), + node_id(index), + journal.clone(), + Arc::new(adapters::Cells { + records: records.clone(), + local: index, + root: root.path().into(), + receiver_directory: None, + }), + )?; + let manager = node.install_read_replicas( + layout.clone(), + directory.clone(), + root.path().join(format!("readers-{index}")), + limits, + )?; + node.install_fleet_reader_enrollment(scope(), node_id(index), journal.clone())?; + managers.push(manager); + let ad = startup::advertisement(index, &node, &intent).await?; + let spec = startup::spec(&intent)?; + boots.push(startup::BootOwner { + node: node.clone(), + directory: directory.clone(), + spec: spec.clone(), + advertisement: ad.clone(), + guard: None, + }); + let original = startup::enroll(journal, &directory, &spec, ad.clone(), clock()?).await?; + let guard = NodeLeaseGuard::new(clock()?, ad.expires_at_ms())?; + node.install_node_lease_for_startup(guard.clone())?; + boots[index].guard = Some(guard); + node.confirm_fleet_startup(journal.as_ref(), spec.key()?) + .await?; + let observed = directory + .load(session(index), clock()?) + .await? + .ok_or_else(|| invalid("original reader maintenance boot is absent"))?; + node.install_fleet_boot_withdrawal(directory.clone(), observed, original, journal.clone())?; + node.start()?; + } + let handle = nodes[0] + .runtime() + .bootstrap( + catalog, + replica, + authority, + initial, + root.path().join("source.sqlite"), + |tx| { + tx.execute_batch( + "CREATE TABLE counter(value INTEGER); INSERT INTO counter VALUES (17)", + )?; + Ok(()) + }, + ) + .await?; + let now = clock()?; + let identity = MutationIdentity { + request_id: RequestId::from_bytes([1; 16]), + issued_at_ms: now, + expires_at_ms: now + .checked_add(120_000) + .ok_or_else(|| invalid("reader maintenance receipt deadline overflow"))?, + }; + let digest = Digest::from_bytes([1; 32]); + let outcome = handle + .execute(identity, digest, now, 64, 64, |tx| { + tx.execute_batch("UPDATE counter SET value = 29")?; + Ok(HandlerOutcome::Success(29i64.to_be_bytes().to_vec())) + }) + .await?; + let acknowledged = Acknowledged { + identity, + digest, + outcome, + value: 29, + source: handle, + }; + let peer = ReplicaPeerClient::new( + app.registry(), + Arc::new(PeerSigner::new( + session(0), + app.registry().release_digest(), + SigningKey::from_bytes(&[1; 32]), + )), + PeerPrincipal { + issuer: "managed-owner".into(), + subject: "live-owner".into(), + actions: vec!["replica-maintenance".into()], + }, + Arc::new(native_peers::NativePeers::new( + nodes, + &managers, + &layout, + directory.clone(), + )), + ); + let description = CellDescription { + cell: target.cell_id(), + incarnation, + code, + schema: 1, + }; + managers[1] + .set_target(&target, 0, 1) + .await? + .ok_or_else(|| invalid("reader policy is absent"))?; + let initial_reader = boots[1] + .refresh_capacity(1, journal.as_ref(), deadline()) + .await?; + peer.activate(&target, &directory, initial_reader, description) + .await?; + let version = journal.load_snapshot(scope()).await?.registry(); + let version = journal.bootstrap_registry(version).await?; + journal.set_scheduling(version, true).await?; + let verifier = FleetReaderEvacuationVerifier::new( + directory, + CellAuthority::new(layout.clone()), + ReadPolicyStore::new(layout), + peer.clone(), + ); + Ok(Inputs { + records, + managers, + target, + description, + peer, + verifier, + acknowledged, + }) +} diff --git a/crates/cellule-host/minion/scenario/reader_maintenance/tests.rs b/crates/cellule-host/minion/scenario/reader_maintenance/tests.rs new file mode 100644 index 00000000..e6ece1be --- /dev/null +++ b/crates/cellule-host/minion/scenario/reader_maintenance/tests.rs @@ -0,0 +1,17 @@ +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn executable_reader_maintenance_preserves_receipts_and_joins_every_owner() { + let summary = super::super::maintenance_reader().await.unwrap(); + assert!(summary.maintenance_completed); + assert!(summary.maintenance_boot_withdrawn); + assert_eq!(summary.final_counts, [1, 0, 0]); + assert_eq!(summary.receipt_checks, 2); + assert_eq!(summary.joined_nodes, 3); + assert_eq!(summary.boot_retirements, 3); + assert_eq!( + (summary.released, summary.activated, summary.retired), + (0, 0, 0) + ); + assert_eq!(summary.max_inflight, 0); + assert_eq!(summary.max_restore_bytes, 0); + assert!(!summary.blockers.is_empty()); +} diff --git a/crates/cellule-host/minion/scenario/reader_tests/evacuation/fixture.rs b/crates/cellule-host/minion/scenario/reader_tests/evacuation/fixture.rs index dd0dfcb3..3b0f8cc3 100644 --- a/crates/cellule-host/minion/scenario/reader_tests/evacuation/fixture.rs +++ b/crates/cellule-host/minion/scenario/reader_tests/evacuation/fixture.rs @@ -120,6 +120,7 @@ impl Fixture { node_id(index), journal.clone(), Arc::new(super::super::super::adapters::Cells { + receiver_directory: None, records: records.clone(), local: index, root: root.path().into(), @@ -369,6 +370,7 @@ impl Fixture { directory, nodes, managers, + records, boots, handle, target, diff --git a/crates/cellule-host/minion/scenario/reader_tests/evacuation/mod.rs b/crates/cellule-host/minion/scenario/reader_tests/evacuation/mod.rs index 598e3e49..8d4747ce 100644 --- a/crates/cellule-host/minion/scenario/reader_tests/evacuation/mod.rs +++ b/crates/cellule-host/minion/scenario/reader_tests/evacuation/mod.rs @@ -18,7 +18,7 @@ mod inventory_tests; mod persisted; mod source; mod tests; -mod transport; +use crate::scenario::native_peers as transport; struct Fixture { layout: CellStorageLayout, @@ -28,6 +28,7 @@ struct Fixture { directory: NodeDirectory, nodes: Vec>, managers: Vec, + records: Arc>, boots: Vec, handle: CellHandle, target: CellTarget, diff --git a/crates/cellule-host/minion/scenario/reader_tests/evacuation/persisted/observation.rs b/crates/cellule-host/minion/scenario/reader_tests/evacuation/persisted/observation.rs index 031acd8b..5fa08b0c 100644 --- a/crates/cellule-host/minion/scenario/reader_tests/evacuation/persisted/observation.rs +++ b/crates/cellule-host/minion/scenario/reader_tests/evacuation/persisted/observation.rs @@ -2,7 +2,7 @@ use super::*; use cellule_host::fleet::{ FleetActionCompletion, FleetAdapterFuture, FleetMaintenanceEnrollments, FleetObservation, - FleetObserver, FleetReaderEvacuationCheck, FleetRoster, FleetTransport, + FleetObserver, FleetReaderEvacuationCheck, FleetReconciler, FleetRoster, FleetTransport, }; use cellule_runtime::fleet::operations::{ FleetAction, FleetInspectionObservation, FleetInspectionRequest, @@ -13,6 +13,137 @@ use std::sync::{ atomic::{AtomicUsize, Ordering}, }; +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn reference_observer_covers_a_live_reader_but_keeps_replacement_policy_open() { + let fixture = Fixture::new().await; + let snapshot = fixture.journal.load_snapshot(scope()).await.unwrap(); + let roster = FleetRoster::collect(fixture.journal.as_ref(), &snapshot, deadline()) + .await + .unwrap(); + let fleet = Arc::new(crate::scenario::adapters::LocalFleet { + nodes: fixture.nodes.clone(), + journal: fixture.journal.clone(), + boots: fixture.boots.clone(), + records: fixture.records.clone(), + reader_verifier: None, + capture_sequence: std::sync::atomic::AtomicU64::new(0), + lose_release_replies: false, + lost_release_replies: AtomicUsize::new(0), + drop_closed_finalize_replies: std::sync::atomic::AtomicUsize::new(0), + expired_receiver_cleanups: AtomicUsize::new(0), + }); + + let observation = fleet + .observe(&roster, clock().unwrap(), deadline()) + .await + .unwrap(); + let graph = observation.role_coverage().unwrap(); + assert_eq!(graph.native_boots(), 3); + assert_eq!(graph.physical_nodes(), 3); + assert_eq!(graph.pending_enrollments(), 0); + let policies = observation.maintenance_policy_coverage().unwrap(); + assert!(!policies.is_complete()); + assert_eq!(policies.progress().required, 1); + assert_eq!(policies.progress().established, 1); + + fixture.finish().await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn reference_observer_reconciles_reader_only_maintenance_to_completion() { + let fixture = Fixture::new().await; + let fleet = Arc::new(crate::scenario::adapters::LocalFleet { + nodes: fixture.nodes.clone(), + journal: fixture.journal.clone(), + boots: fixture.boots.clone(), + records: fixture.records.clone(), + reader_verifier: Some(fixture.verifier()), + capture_sequence: std::sync::atomic::AtomicU64::new(0), + lose_release_replies: false, + lost_release_replies: AtomicUsize::new(0), + drop_closed_finalize_replies: std::sync::atomic::AtomicUsize::new(0), + expired_receiver_cleanups: AtomicUsize::new(0), + }); + + fixture.spare().await; + let capture = fixture.evacuate().await.unwrap(); + let publication = fixture.publish(&capture).await; + let snapshot = fixture.journal.load_snapshot(scope()).await.unwrap(); + let roster = FleetRoster::collect(fixture.journal.as_ref(), &snapshot, deadline()) + .await + .unwrap(); + let observation = fleet + .observe(&roster, clock().unwrap(), deadline()) + .await + .unwrap(); + assert_eq!(observation.role_coverage().unwrap().native_boots(), 3); + let policies = observation.maintenance_policy_coverage().unwrap(); + assert!(policies.is_complete()); + assert_eq!(policies.progress().required, 1); + assert_eq!(policies.progress().checked, 1); + assert_eq!( + policies.obligations()[0].status(), + cellule_host::fleet::FleetMaintenancePolicyStatus::Reader( + publication.record().unwrap().digest().unwrap() + ) + ); + + enabled(&fixture).await; + let driver = FleetReconciler::new( + scope(), + SessionId::from_bytes([206; 16]), + FleetProfile::default(), + fixture.journal.clone(), + fleet.clone(), + fleet.clone(), + ) + .unwrap(); + let mut report = None; + for _ in 0..3 { + let next = driver.reconcile_once(clock, deadline()).await.unwrap(); + let completed = next.snapshot.head().maintenance().is_some_and(|operation| { + operation.phase() == cellule_runtime::fleet::operations::MaintenancePhase::Completed + }); + report = Some(next); + if completed { + break; + } + } + let report = report.unwrap(); + assert_eq!( + report.snapshot.head().maintenance().unwrap().phase(), + cellule_runtime::fleet::operations::MaintenancePhase::Completed + ); + assert_eq!(fixture.nodes[1].state(), NodeState::Stopped); + assert!(fixture.directory.is_withdrawn(session(1)).await.unwrap()); + assert_eq!( + fixture.original_row().await.status(), + EnrollmentStatus::Retired + ); + assert!( + fixture + .reader + .query::(None, 0) + .await + .is_err() + ); + let replacement = fixture.managers[2] + .resolve(fixture.target.clone()) + .await + .unwrap(); + assert_eq!( + replacement + .query::(Some(capture.minimum()), 0) + .await + .unwrap() + .output, + 17 + ); + + drop((capture, driver, fleet)); + fixture.finish().await; +} + pub(super) async fn advertisements(directory: &NodeDirectory) -> Vec { let mut nodes = Vec::new(); for index in 0..3 { diff --git a/crates/cellule-host/minion/scenario/reader_tests/evacuation/persisted/tests.rs b/crates/cellule-host/minion/scenario/reader_tests/evacuation/persisted/tests.rs index 4fd515d5..84c6719f 100644 --- a/crates/cellule-host/minion/scenario/reader_tests/evacuation/persisted/tests.rs +++ b/crates/cellule-host/minion/scenario/reader_tests/evacuation/persisted/tests.rs @@ -13,6 +13,9 @@ async fn reader_policy_publication_can_refresh_redundancy_during_closing_without let record = publication.record().unwrap(); let retired = fixture.original_row().await; let snapshot = fixture.journal.load_snapshot(scope()).await.unwrap(); + crate::scenario::commit_test_role_settlement(&fixture.journal, clock().unwrap()) + .await + .unwrap(); let snapshot = fixture .journal .compare_exchange( diff --git a/crates/cellule-host/minion/scenario/receiver_loss/adapters.rs b/crates/cellule-host/minion/scenario/receiver_loss/adapters.rs new file mode 100644 index 00000000..86ffb793 --- /dev/null +++ b/crates/cellule-host/minion/scenario/receiver_loss/adapters.rs @@ -0,0 +1,199 @@ +//! Fixed in-process failure adapters shared by the command and fault tests. +use super::*; +use cellule_host::fleet::{ + FleetAdapterFuture, FleetFailedBootProcessEvidence, FleetFailedBootProcessRequest, + FleetFailedBootProcesses, FleetFailedBootRetirement, FleetObservation, FleetObserver, + FleetRoster, FleetTransport, +}; +use cellule_runtime::fleet::operations::{ + EnrollmentStatus, FleetAction, FleetActionKind, FleetInspectionRequest, FleetOutcome, + MovementAction, +}; + +/// Reference in-process provider: joined CellNode shutdown supplies retained +/// lifetime evidence for this finite constructor. It does not qualify OS +/// process supervision or replace an application's external-work provider. +#[derive(Clone)] +pub(in crate::scenario) struct StoppedNodeProcesses { + pub(in crate::scenario) node: Arc, +} + +impl FleetFailedBootProcesses for StoppedNodeProcesses { + fn confirm_stopped<'a>( + &'a self, + request: &'a FleetFailedBootProcessRequest, + ) -> FleetAdapterFuture<'a, FleetFailedBootProcessEvidence> { + Box::pin(async move { + let endpoint = request.boot().spec().target; + let stats = self.node.stats(); + if endpoint.node != node_id(1) + || endpoint.session != session(1) + || self.node.state() != NodeState::Stopped + || stats.active_cells() != 0 + || stats.worker_jobs() != 0 + || stats.retained_bytes() != 0 + || stats.resident_bytes() != 0 + || stats.file_descriptors() != 0 + || stats.local_disk_reserved_bytes() != 0 + || stats.primitive_jobs() != 0 + || stats.hydration_jobs() != 0 + || stats.io_slots() != 0 + || stats.blocking_jobs() != 0 + || stats.recovery_jobs() != 0 + { + return Err(invalid("receiver process has not joined shutdown")); + } + let mut hash = blake3::Hasher::new(); + hash.update(b"cellule.example-test-stopped-node.v1\0"); + hash.update(request.digest().as_bytes()); + Ok(FleetFailedBootProcessEvidence::new( + request, + Digest::from_bytes(*hash.finalize().as_bytes()), + )?) + }) + } +} + +pub(in crate::scenario) async fn close_receiver( + fleet: Arc, +) -> JournalResult> { + let node = fleet + .nodes + .get(1) + .ok_or_else(|| invalid("receiver-loss node absent"))?; + let boot = fleet + .boots + .get(1) + .ok_or_else(|| invalid("receiver-loss boot absent"))?; + // This finite constructor omits normal host withdrawal on the failed + // receiver. Join its runtime, then publish the independently checked typed + // closure under its original enrollment; directory absence is insufficient. + boot.guard + .as_ref() + .ok_or_else(|| invalid("receiver-loss guard absent"))? + .fence(); + node.shutdown().await?; + if node.state() != NodeState::Stopped { + return Err(invalid("receiver-loss shutdown did not stop")); + } + let now = clock()?; + let observed = boot + .directory + .load_if_live(session(1), now) + .await? + .ok_or_else(|| invalid("receiver-loss original advertisement absent"))?; + boot.directory.withdraw_after_drain(&observed, now).await?; + let snapshot = fleet.journal.load_snapshot(scope()).await?; + let roster = FleetRoster::collect( + fleet.journal.as_ref(), + &snapshot, + Instant::now() + Duration::from_secs(5), + ) + .await?; + let original = roster + .enrollments() + .iter() + .find(|row| { + row.spec().target.node == node_id(1) + && row.spec().target.session == session(1) + && row.status() == EnrollmentStatus::Established + }) + .ok_or_else(|| invalid("receiver-loss original enrollment absent"))?; + let processes = StoppedNodeProcesses { node: node.clone() }; + let retirement = FleetFailedBootRetirement::capture( + fleet.journal.as_ref(), + &boot.directory, + &roster, + original, + session(0), + Instant::now() + Duration::from_secs(5), + clock, + ) + .await?; + let request = retirement.request().clone(); + retirement + .publish( + fleet.journal.as_ref(), + &boot.directory, + &processes, + session(0), + Instant::now() + Duration::from_secs(5), + clock, + ) + .await? + .confirmed()?; + Ok(Arc::new(ClosedBootObserver { + fleet, + request, + processes, + })) +} + +pub(in crate::scenario) struct ClosedBootObserver { + pub(in crate::scenario) fleet: Arc, + pub(in crate::scenario) request: FleetFailedBootProcessRequest, + pub(in crate::scenario) processes: StoppedNodeProcesses, +} + +impl FleetObserver for ClosedBootObserver { + fn observe<'a>( + &'a self, + roster: &'a FleetRoster, + _: i64, + deadline: Instant, + ) -> FleetAdapterFuture<'a, FleetObservation> { + Box::pin(observation::observe_with_failed_boot_closure( + &self.fleet, + roster, + deadline, + &self.request, + &self.processes, + session(0), + )) + } +} + +pub(in crate::scenario) struct LoseRoutedActivationReply { + pub(in crate::scenario) inner: Arc, + pub(in crate::scenario) lost: std::sync::atomic::AtomicBool, +} + +impl FleetTransport for LoseRoutedActivationReply { + fn dispatch<'a>( + &'a self, + action: &'a FleetAction, + deadline: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + Box::pin(async move { + let completion = self.inner.dispatch(action, deadline).await?; + let routed_activation = action.receiver_route().is_some() + && matches!( + action.kind(), + FleetActionKind::Movement { + action: MovementAction::Activate, + .. + } + ); + if routed_activation && !self.lost.swap(true, std::sync::atomic::Ordering::AcqRel) { + if !completion.committed + || !matches!(&completion.outcome.outcome, FleetOutcome::Activated(_)) + { + return Err(invalid( + "routed activation did not commit before reply loss", + )); + } + return Err(invalid("injected loss after committed routed activation")); + } + Ok(completion) + }) + } + + fn inspect<'a>( + &'a self, + request: &'a FleetInspectionRequest, + deadline: Instant, + ) -> FleetAdapterFuture<'a, Arc> + { + self.inner.inspect(request, deadline) + } +} diff --git a/crates/cellule-host/minion/scenario/receiver_loss/mod.rs b/crates/cellule-host/minion/scenario/receiver_loss/mod.rs new file mode 100644 index 00000000..9cc0123f --- /dev/null +++ b/crates/cellule-host/minion/scenario/receiver_loss/mod.rs @@ -0,0 +1,166 @@ +//! Finite receiver loss and durable controller reconstruction. +use super::adapters::LocalFleet; +use super::*; +use cellule_host::fleet::{FleetActionAcceptance, FleetActionJournal, FleetTransport}; +use cellule_runtime::fleet::operations::*; + +pub(super) mod adapters; +mod movement; +#[cfg(test)] +mod tests; + +struct ReleasedScenario { + records: Arc>, + acknowledged: HashMap, + attempt: MoveAttempt, +} + +pub(super) fn run<'a>( + root: &'a tempfile::TempDir, + journal: Arc, + nodes: &'a mut Vec>, + boots: &'a mut Vec, + profile: FleetProfile, +) -> FleetAdapterFuture<'a, ScenarioSummary> { + Box::pin(run_inner(root, journal, nodes, boots, profile)) +} + +async fn run_inner( + root: &tempfile::TempDir, + journal: Arc, + nodes: &mut Vec>, + boots: &mut Vec, + profile: FleetProfile, +) -> JournalResult { + let (records, acknowledged) = + initialize_inner(root, &journal, nodes, boots, 60_000, Some(1), false).await?; + let fleet = Arc::new(LocalFleet { + nodes: nodes.clone(), + journal: journal.clone(), + boots: boots.clone(), + records: records.clone(), + reader_verifier: None, + capture_sequence: std::sync::atomic::AtomicU64::new(0), + lose_release_replies: false, + lost_release_replies: std::sync::atomic::AtomicUsize::new(0), + drop_closed_finalize_replies: std::sync::atomic::AtomicUsize::new(0), + expired_receiver_cleanups: std::sync::atomic::AtomicUsize::new(0), + }); + let driver = FleetReconciler::new( + scope(), + SessionId::from_bytes([206; 16]), + profile, + journal.clone(), + fleet.clone(), + fleet.clone(), + )?; + let first = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await?; + if !first.snapshot.head().attempts().is_empty() { + return Err(invalid( + "receiver-loss startup unexpectedly planned movement", + )); + } + let page = nodes[0].runtime().fleet_cells_page(None, 128).await?; + let Some(CellInventoryEntry::Owned(row)) = page.entries().first() else { + return Err(invalid("receiver-loss source writer absent")); + }; + let spec = MoveAttemptSpec { + id: AttemptId { + operation: OperationId::from_bytes([220; 16])?, + sequence: 1, + }, + target: row.target.clone(), + incarnation: row.incarnation, + source_node: node_id(0), + source: session(0), + generation: row.generation, + source_epoch: row + .position + .as_ref() + .ok_or_else(|| invalid("receiver-loss source has no position"))? + .epoch, + destination_node: node_id(1), + destination: session(1), + cost: row + .cost + .ok_or_else(|| invalid("receiver-loss source cost absent"))?, + snapshot_digest: Digest::from_bytes([221; 32]), + deadline_ms: clock()? + 60_000, + }; + drop(page); + journal + .compare_exchange( + &first.snapshot, + first + .snapshot + .head() + .controller() + .ok_or_else(|| invalid("receiver-loss controller absent"))? + .epoch, + clock()?, + &JournalTransition::Allocate(spec.clone()), + ) + .await?; + let version = journal.load_snapshot(scope()).await?.registry(); + journal.set_scheduling(version, false).await?; + let prepared = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await?; + if !prepared.failures.is_empty() { + return Err(failure::endpoints( + "receiver-loss prepare", + 0, + prepared.failures, + )); + } + if prepared + .snapshot + .head() + .attempts() + .first() + .is_none_or(|attempt| attempt.phase() != AttemptPhase::Reserved) + { + return Err(invalid("receiver-loss preparation did not reserve")); + } + let release = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await?; + if !release.failures.is_empty() { + return Err(failure::endpoints( + "receiver-loss release", + 0, + release.failures, + )); + } + if release.released != 1 { + return Err(invalid("receiver-loss source release was not proved")); + } + let released_attempt = release + .snapshot + .head() + .attempts() + .first() + .ok_or_else(|| invalid("receiver-loss release attempt absent"))? + .clone(); + if released_attempt.released().is_none() + || nodes[1].stats().local_disk_reserved_bytes() != spec.cost.disk_bytes + { + return Err(invalid("receiver-loss release or original credit differs")); + } + let observer = adapters::close_receiver(fleet.clone()).await?; + movement::continue_released( + root, + journal, + fleet, + observer, + ReleasedScenario { + records, + acknowledged, + attempt: released_attempt, + }, + profile, + ) + .await +} diff --git a/crates/cellule-host/minion/scenario/receiver_loss/movement.rs b/crates/cellule-host/minion/scenario/receiver_loss/movement.rs new file mode 100644 index 00000000..1823ba41 --- /dev/null +++ b/crates/cellule-host/minion/scenario/receiver_loss/movement.rs @@ -0,0 +1,360 @@ +//! Adopt one exact release and retained routed result across a real lease expiry. +use super::*; +use adapters::{ClosedBootObserver, LoseRoutedActivationReply}; +use cellule_host::fleet::FleetReconcileReport; + +pub(super) fn continue_released<'a>( + root: &'a tempfile::TempDir, + journal: Arc, + fleet: Arc, + observer: Arc, + state: ReleasedScenario, + profile: FleetProfile, +) -> FleetAdapterFuture<'a, ScenarioSummary> { + Box::pin(continue_inner( + root, journal, fleet, observer, state, profile, + )) +} + +async fn continue_inner( + root: &tempfile::TempDir, + journal: Arc, + fleet: Arc, + observer: Arc, + state: ReleasedScenario, + profile: FleetProfile, +) -> JournalResult { + let released = state + .attempt + .released() + .ok_or_else(|| invalid("receiver-loss release evidence absent"))?; + let loss = Arc::new(LoseRoutedActivationReply { + inner: fleet.clone(), + lost: std::sync::atomic::AtomicBool::new(false), + }); + let driver = FleetReconciler::new( + scope(), + SessionId::from_bytes([206; 16]), + profile, + journal.clone(), + observer.clone(), + loss.clone(), + )?; + let mut summary = ScenarioSummary { + released: 1, + activated: 0, + retired: 0, + receipt_checks: 0, + max_inflight: 1, + max_restore_bytes: state.attempt.spec().cost.disk_bytes, + joined_nodes: 0, + boot_retirements: 0, + receiver_nodes: 1, + lost_release_replies: 0, + controller_epoch: 1, + expired_receiver_cleanups: 0, + blockers: Vec::new(), + final_counts: vec![0; 3], + maintenance_completed: false, + maintenance_boot_withdrawn: false, + receiver_process_closures: 1, + lost_activation_replies: 0, + routed_activation_replays: 0, + }; + let mut report = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await?; + accumulate(&mut summary, &report, profile)?; + for _ in 0..8 { + if loss.lost.load(std::sync::atomic::Ordering::Acquire) { + break; + } + report = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await?; + accumulate(&mut summary, &report, profile)?; + } + if !loss.lost.load(std::sync::atomic::Ordering::Acquire) { + return Err(failure::endpoints( + "receiver-loss replacement did not reach committed reply loss", + 8, + report.failures, + )); + } + let accepted = routed_acceptance(&journal, &state.attempt, MovementAction::Activate).await?; + let replay = fleet + .dispatch(accepted.action(), Instant::now() + Duration::from_secs(5)) + .await?; + let FleetOutcome::Activated(ref activation) = replay.outcome.outcome else { + return Err(invalid("receiver-loss replay did not confirm activation")); + }; + if !replay.committed + || fleet.nodes[2].stats().active_cells() != 1 + || activation.position.root != released.root + || activation.position.epoch != released.epoch + 1 + { + return Err(invalid( + "receiver-loss replay changed root, epoch or actor count", + )); + } + summary.lost_activation_replies = 1; + summary.routed_activation_replays = 1; + let expires = report + .snapshot + .head() + .controller() + .ok_or_else(|| invalid("receiver-loss controller lease absent"))? + .expires_at_ms; + let wait = u64::try_from(expires.saturating_sub(clock()?).max(0))? + .checked_add(50) + .ok_or_else(|| invalid("receiver-loss lease wait overflow"))?; + tokio::time::sleep(Duration::from_millis(wait)).await; + let reopened = Arc::new( + SqliteJournal::open( + root.path().join("fleet-journal.sqlite"), + scope(), + profile, + clock()?, + ) + .await?, + ); + // Own this independently reconstructed client through every resume exit. + // The outer scenario still owns and joins all native nodes and the original + // journal before its private directory can be dropped. + let result = async { + let summary = Box::pin(resume( + &reopened, &fleet, observer, &state, profile, report, summary, + )) + .await?; + let before = reopened.load_snapshot(scope()).await?; + let error = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(1)) + .await + .err() + .ok_or_else(|| invalid("receiver-loss old controller renewed successor lease"))?; + if !matches!( + &error, + cellule_runtime::Error::Facility { name: "fleet-journal", source } + if matches!(source.downcast_ref::(), Some(OperationError::Fenced)) + ) { + return Err(error.into()); + } + if reopened.load_snapshot(scope()).await? != before { + return Err(invalid( + "receiver-loss old controller changed successor journal", + )); + } + Ok::<_, JournalError>(summary) + } + .await; + let close = reopened.close().await; + match result { + Ok(summary) => { + close?; + Ok(summary) + } + Err(error) => { + if let Err(cleanup) = close { + eprintln!("additional receiver-loss journal cleanup failure: {cleanup:?}"); + } + Err(error) + } + } +} + +async fn resume( + journal: &Arc, + fleet: &Arc, + observer: Arc, + state: &ReleasedScenario, + profile: FleetProfile, + mut report: FleetReconcileReport, + mut summary: ScenarioSummary, +) -> JournalResult { + let driver = FleetReconciler::new( + scope(), + SessionId::from_bytes([207; 16]), + profile, + journal.clone(), + observer, + fleet.clone(), + )?; + let mut serving = None; + for _ in 0..12 { + if report.snapshot.head().attempts().is_empty() { + break; + } + report = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await?; + accumulate(&mut summary, &report, profile)?; + if let Some(attempt) = report.snapshot.head().attempts().first() { + if attempt.spec() != state.attempt.spec() + || attempt.released() != state.attempt.released() + { + return Err(invalid( + "receiver-loss reconstruction changed original release", + )); + } + serving = serving.or_else(|| attempt.activated().cloned()); + } + } + if !report.snapshot.head().attempts().is_empty() + || report.snapshot.head().reserved_restore_bytes() != 0 + || summary.released != 1 + || summary.activated != 1 + || summary.retired != 1 + || summary.controller_epoch != 2 + { + return Err(failure::endpoints( + "receiver-loss retained work did not settle", + 12, + report.failures, + )); + } + let serving = serving.ok_or_else(|| invalid("receiver-loss successor evidence absent"))?; + let released = state + .attempt + .released() + .ok_or_else(|| invalid("receiver-loss release disappeared"))?; + if (serving.node, serving.session) != (node_id(2), session(2)) + || serving.position.root != released.root + || serving.position.epoch != released.epoch + 1 + { + return Err(invalid("receiver-loss successor differs from release")); + } + for effect in [MovementAction::Activate, MovementAction::Cancel] { + routed_acceptance(journal, &state.attempt, effect).await?; + } + for (cell, record) in state.records.iter() { + let original = state + .acknowledged + .get(cell) + .ok_or_else(|| invalid("receiver-loss acknowledged Cell missing"))?; + let observed = record + .authority + .load(*cell) + .await? + .ok_or_else(|| invalid("receiver-loss authority absent"))?; + let owner = observed + .value() + .owner + .as_ref() + .ok_or_else(|| invalid("receiver-loss current owner absent"))?; + let index = (0..3) + .find(|index| session(*index) == owner.session) + .ok_or_else(|| invalid("receiver-loss current owner unbound"))?; + let handle = fleet.nodes[index] + .runtime() + .local_handle(record.catalog.clone(), &observed) + .await? + .ok_or_else(|| invalid("receiver-loss native actor absent"))?; + if handle + .resolve(original.identity, original.digest, clock()?, 64) + .await? + != Resolution::Committed(original.outcome.clone()) + { + return Err(invalid("receiver-loss command receipt changed")); + } + let bytes = handle + .query(64, 64, |connection| { + let value: i64 = + connection.query_row("SELECT value FROM counter", [], |row| row.get(0))?; + Ok(value.to_be_bytes().to_vec()) + }) + .await?; + if bytes != original.value.to_be_bytes() { + return Err(invalid("receiver-loss acknowledged state changed")); + } + summary.receipt_checks += 1; + } + let original = state + .acknowledged + .get(&state.attempt.spec().target.cell_id()) + .ok_or_else(|| invalid("receiver-loss original handle absent"))?; + if original + .source + .query(64, 64, |_| Ok(Vec::new())) + .await + .is_ok() + { + return Err(invalid("receiver-loss released source still serves")); + } + let roster = cellule_host::fleet::FleetRoster::collect( + journal.as_ref(), + &report.snapshot, + Instant::now() + Duration::from_secs(5), + ) + .await?; + summary.final_counts = + observation::complete_counts(fleet, &roster, Instant::now() + Duration::from_secs(5)) + .await? + .ok_or_else(|| invalid("receiver-loss final count coverage incomplete"))?; + if summary.final_counts != [CELL_COUNT - 1, 0, 1] || summary.receipt_checks != CELL_COUNT { + return Err(invalid("receiver-loss final placement differs")); + } + Ok(summary) +} + +async fn routed_acceptance( + journal: &SqliteJournal, + attempt: &MoveAttempt, + effect: MovementAction, +) -> JournalResult { + let accepted = journal + .load_movement_actions(scope(), attempt, effect) + .await? + .into_iter() + .find_map(|value| match value { + FleetActionAcceptance::New(accepted) + | FleetActionAcceptance::Existing { accepted, .. } + if accepted.action().receiver_endpoint() == Some((node_id(2), session(2))) => + { + Some(accepted) + } + _ => None, + }) + .ok_or_else(|| invalid("receiver-loss routed action absent"))?; + if accepted + .action() + .receiver_route() + .and_then(|route| route.latest_handoff()) + .is_none_or(|hop| hop.previous() != (node_id(1), session(1))) + { + return Err(invalid("receiver-loss closed handoff differs")); + } + Ok(accepted) +} + +fn accumulate( + summary: &mut ScenarioSummary, + report: &FleetReconcileReport, + profile: FleetProfile, +) -> JournalResult<()> { + summary.released += report.released; + summary.activated += report.activated; + summary.retired += report.retired; + summary.max_inflight = summary + .max_inflight + .max(report.snapshot.head().attempts().len()); + summary.max_restore_bytes = summary + .max_restore_bytes + .max(report.snapshot.head().reserved_restore_bytes()); + summary.controller_epoch = report + .snapshot + .head() + .controller() + .ok_or_else(|| invalid("receiver-loss report has no controller"))? + .epoch; + for blocker in &report.blockers { + if !summary.blockers.contains(blocker) { + summary.blockers.push(*blocker); + } + } + if summary.max_inflight > profile.max_inflight + || summary.max_restore_bytes > profile.max_restore_bytes + { + return Err(invalid("receiver-loss exceeded shared movement budget")); + } + Ok(()) +} diff --git a/crates/cellule-host/minion/scenario/receiver_loss/tests.rs b/crates/cellule-host/minion/scenario/receiver_loss/tests.rs new file mode 100644 index 00000000..42159086 --- /dev/null +++ b/crates/cellule-host/minion/scenario/receiver_loss/tests.rs @@ -0,0 +1,21 @@ +use super::*; + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn receiver_loss_command_preserves_every_receipt_and_joins_all_owners() { + let summary = crate::scenario::receiver_loss().await.unwrap(); + assert_eq!(summary.released, 1); + assert_eq!(summary.activated, 1); + assert_eq!(summary.retired, 1); + assert_eq!(summary.receipt_checks, CELL_COUNT); + assert_eq!(summary.max_inflight, 1); + assert!(summary.max_restore_bytes > 0); + assert!(summary.max_restore_bytes <= FleetProfile::default().max_restore_bytes); + assert_eq!(summary.controller_epoch, 2); + assert_eq!(summary.receiver_process_closures, 1); + assert_eq!(summary.lost_activation_replies, 1); + assert_eq!(summary.routed_activation_replays, 1); + assert_eq!(summary.joined_nodes, 3); + assert_eq!(summary.boot_retirements, 3); + assert_eq!(summary.final_counts, [CELL_COUNT - 1, 0, 1]); + assert_eq!(summary.blockers, vec![DrainBlocker::OutcomeUnknown]); +} diff --git a/crates/cellule-host/minion/scenario/recovered_followers/failed_boot/confirmation.rs b/crates/cellule-host/minion/scenario/recovered_followers/failed_boot/confirmation.rs index 6d9adb2f..11550c20 100644 --- a/crates/cellule-host/minion/scenario/recovered_followers/failed_boot/confirmation.rs +++ b/crates/cellule-host/minion/scenario/recovered_followers/failed_boot/confirmation.rs @@ -3,7 +3,16 @@ use super::*; use cellule_host::fleet::FleetFailedBootClosure; pub(super) async fn settled(now: i64, retire_boot: bool) -> Fixture { - let fixture = Fixture::with_recovery_inputs_at(true, Vec::new(), Vec::new(), now).await; + settled_with_profile(now, retire_boot, FleetProfile::default()).await +} + +pub(super) async fn settled_with_profile( + now: i64, + retire_boot: bool, + profile: FleetProfile, +) -> Fixture { + let fixture = + Fixture::with_recovery_inputs_at_profile(true, Vec::new(), Vec::new(), now, profile).await; let checked = now + 10_005; let proof = retire_recovered_members(fixture.transport.clone(), &fixture.sealed) .await diff --git a/crates/cellule-host/minion/scenario/recovered_followers/failed_boot/observation.rs b/crates/cellule-host/minion/scenario/recovered_followers/failed_boot/observation.rs index 5361e771..97fef5dd 100644 --- a/crates/cellule-host/minion/scenario/recovered_followers/failed_boot/observation.rs +++ b/crates/cellule-host/minion/scenario/recovered_followers/failed_boot/observation.rs @@ -2,13 +2,16 @@ //! This fixture has no writers and uses a joined process lifetime stand-in. use super::*; use cellule_host::fleet::{ - FleetActionCompletion, FleetFailedBootClosure, FleetFollowerReferences, FleetNodeInventory, - FleetNodeInventoryScan, FleetObservation, FleetObserver, FleetRoleCoverage, - FleetSnapshotRequest, FleetSnapshotSubject, FleetTransport, + FleetActionCompletion, FleetFailedBootClosure, FleetFollowerReferences, + FleetMaintenanceEnrollments, FleetNodeInventory, FleetNodeInventoryScan, FleetObservation, + FleetObserver, FleetRecoveredFollowerClosure, FleetRecoveredFollowerRetirement, + FleetRoleCoverage, FleetRoster, FleetSnapshotRequest, FleetSnapshotSubject, FleetTransport, }; use cellule_runtime::fleet::operations::{ - FleetAction, FleetInspectionObservation, FleetInspectionRequest, + FleetAction, FleetInspectionObservation, FleetInspectionRequest, JournalTransition, + MaintenanceEvent, MaintenanceOperation, OperationId, }; +use std::sync::atomic::AtomicU64; struct Observed { base: Fixture, @@ -17,7 +20,16 @@ struct Observed { } impl Observed { async fn new() -> Self { - let base = confirmation::settled(crate::scenario::clock().unwrap() - 10_005, true).await; + Self::new_with_profile(FleetProfile::default()).await + } + + async fn new_with_profile(profile: FleetProfile) -> Self { + let base = confirmation::settled_with_profile( + crate::scenario::clock().unwrap() - 10_005, + true, + profile, + ) + .await; let before = base.journal.load_snapshot(scope()).await.unwrap(); base.journal .set_scheduling(before.registry(), true) @@ -72,6 +84,7 @@ impl Observed { node_id(index), base.journal.clone(), Arc::new(crate::scenario::adapters::Cells { + receiver_directory: None, records: Arc::new(HashMap::new()), local: index, root: base._root.path().into(), @@ -130,6 +143,32 @@ impl Observed { ) .await } + async fn recovered_closure( + &self, + roster: &FleetRoster, + ) -> cellule_runtime::Result { + let retirement = FleetRecoveredFollowerRetirement::capture( + self.base.journal.as_ref(), + &self.base.directory, + roster, + &self.base.sealed, + session(1), + deadline(), + crate::scenario::clock, + ) + .await?; + Ok(retirement + .publish( + self.base.journal.as_ref(), + &self.base.directory, + session(1), + deadline(), + crate::scenario::clock, + ) + .await? + .confirmed()? + .clone()) + } async fn page( &self, roster: &FleetRoster, @@ -252,6 +291,48 @@ impl Observed { .with_role_coverage(graph)? .with_failed_boot_closures(vec![closure]) } + async fn observe_maintenance( + &self, + roster: &FleetRoster, + ) -> cellule_runtime::Result { + let closure = self.closure(roster).await?; + let recovered = self.recovered_closure(roster).await?; + let original = FleetMaintenanceEnrollments::collect( + self.base.journal.as_ref(), + roster, + deadline(), + crate::scenario::clock, + ) + .await?; + let graph = self.graph(roster).await; + let started = [ + closure.interval().0, + recovered.interval().0, + original.interval().0, + graph.interval().0, + ] + .into_iter() + .min() + .ok_or(cellule_runtime::Error::Deadline)?; + roster + .confirm(self.base.journal.as_ref(), deadline()) + .await?; + let observation = FleetObservation::new( + scope(), + roster.snapshot().registry(), + roster.snapshot().registry().revision(), + started, + crate::scenario::clock()?, + true, + self.advertisements().await, + Vec::new(), + )? + .with_role_coverage(graph)? + .with_failed_boot_closures(vec![closure])? + .with_recovered_follower_closures(vec![recovered])? + .with_maintenance_enrollments(original)?; + observation.check_maintenance_policies(roster, crate::scenario::clock()?) + } async fn finish(&self) { for node in &self.nodes { node.shutdown().await.unwrap(); @@ -261,6 +342,196 @@ impl Observed { self.base.journal.close().await.unwrap(); } } + +struct MaintenanceObserver(Arc); +impl FleetObserver for MaintenanceObserver { + fn observe<'a>( + &'a self, + roster: &'a FleetRoster, + _: i64, + _: Instant, + ) -> FleetAdapterFuture<'a, FleetObservation> { + Box::pin(async move { Ok(self.0.observe_maintenance(roster).await?) }) + } +} + +async fn begin_failed_owner_maintenance(fixture: &Observed) -> MaintenanceOperation { + let before = fixture.base.journal.load_snapshot(scope()).await.unwrap(); + let roster = FleetRoster::collect(fixture.base.journal.as_ref(), &before, deadline()) + .await + .unwrap(); + let intent_revision = roster + .intents() + .iter() + .find(|intent| intent.node() == node_id(0)) + .unwrap() + .revision() + + 1; + let now = crate::scenario::clock().unwrap(); + let operation = MaintenanceOperation::new( + OperationId::from_bytes([235; 16]).unwrap(), + Digest::from_bytes([236; 32]), + node_id(0), + session(0), + intent_revision, + now, + now + 60_000, + ) + .unwrap(); + let mut snapshot = fixture + .base + .journal + .compare_exchange( + &before, + before.head().controller().unwrap().epoch, + now, + &JournalTransition::BeginMaintenance(operation.clone()), + ) + .await + .unwrap(); + for transition in [ + JournalTransition::Maintenance(MaintenanceEvent::Cordoned), + JournalTransition::Maintenance(MaintenanceEvent::BeginEvacuation), + ] { + snapshot = fixture + .base + .journal + .compare_exchange( + &snapshot, + snapshot.head().controller().unwrap().epoch, + crate::scenario::clock().unwrap(), + &transition, + ) + .await + .unwrap(); + } + operation +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn failed_owner_maintenance_settles_and_finalizes_only_from_fresh_process_closure() { + let profile = FleetProfile { + controller_lease_ms: 2_500, + reconcile_interval_ms: 1_250, + ..FleetProfile::default() + }; + let fixture = Arc::new(Observed::new_with_profile(profile).await); + let operation = begin_failed_owner_maintenance(&fixture).await; + let transport = Arc::new(crate::scenario::adapters::LocalFleet { + nodes: Vec::new(), + journal: fixture.base.journal.clone(), + boots: Vec::new(), + records: Arc::new(HashMap::new()), + reader_verifier: None, + capture_sequence: AtomicU64::new(0), + lose_release_replies: false, + lost_release_replies: AtomicUsize::new(0), + drop_closed_finalize_replies: AtomicUsize::new(1), + expired_receiver_cleanups: AtomicUsize::new(0), + }); + let first_controller = FleetReconciler::new( + scope(), + session(1), + profile, + fixture.base.journal.clone(), + Arc::new(MaintenanceObserver(fixture.clone())), + transport.clone(), + ) + .unwrap(); + let closing = first_controller + .reconcile_once(crate::scenario::clock, deadline()) + .await + .unwrap(); + assert!(closing.maintenance_failure.is_none()); + assert!(closing.blockers.is_empty()); + assert_eq!( + closing.snapshot.head().maintenance().unwrap().phase(), + cellule_runtime::fleet::operations::MaintenancePhase::Closing + ); + let stopped_reply_lost = first_controller + .reconcile_once(crate::scenario::clock, deadline()) + .await + .unwrap(); + assert!(stopped_reply_lost.maintenance_failure.is_some()); + assert_eq!( + stopped_reply_lost + .snapshot + .head() + .maintenance() + .unwrap() + .phase(), + cellule_runtime::fleet::operations::MaintenancePhase::Closing + ); + assert!( + stopped_reply_lost + .blockers + .contains(&cellule_runtime::fleet::operations::DrainBlocker::OutcomeUnknown) + ); + drop(first_controller); + + // The first controller published Stopped and lost its reply. Wait through + // its actual journal lease, then prove a different controller can replay + // the accepted result from the same SQLite journal. + let controller_expires_at = stopped_reply_lost + .snapshot + .head() + .controller() + .unwrap() + .expires_at_ms; + let wait_ms = controller_expires_at + .saturating_sub(crate::scenario::clock().unwrap()) + .max(0) as u64 + + 10; + tokio::time::sleep(std::time::Duration::from_millis(wait_ms)).await; + let after_expiry = crate::scenario::clock().unwrap(); + assert!(after_expiry > controller_expires_at); + assert!(after_expiry < operation.deadline_ms()); + let restarted_controller = FleetReconciler::new( + scope(), + session(2), + profile, + fixture.base.journal.clone(), + Arc::new(MaintenanceObserver(fixture.clone())), + transport.clone(), + ) + .unwrap(); + let completed = restarted_controller + .reconcile_once(crate::scenario::clock, deadline()) + .await + .unwrap(); + let terminal = completed.snapshot.head().maintenance().unwrap(); + assert_eq!(terminal.id(), operation.id()); + assert_eq!( + terminal.phase(), + cellule_runtime::fleet::operations::MaintenancePhase::Completed + ); + assert!(completed.maintenance_failure.is_none()); + assert!(completed.blockers.is_empty()); + assert!(completed.dispatched >= 1); + assert_eq!( + completed.snapshot.head().controller().unwrap().claimant, + session(2) + ); + assert_eq!( + completed.snapshot.head().controller().unwrap().epoch, + stopped_reply_lost + .snapshot + .head() + .controller() + .unwrap() + .epoch + + 1 + ); + assert_eq!( + transport + .drop_closed_finalize_replies + .load(std::sync::atomic::Ordering::SeqCst), + 0 + ); + drop(restarted_controller); + drop(transport); + fixture.finish().await; +} struct Observer { fixture: Arc, stale: Mutex>, diff --git a/crates/cellule-host/minion/scenario/recovered_followers/failed_boot/writer_tests/successors/observation.rs b/crates/cellule-host/minion/scenario/recovered_followers/failed_boot/writer_tests/successors/observation.rs index 4ff2b92c..7e25b919 100644 --- a/crates/cellule-host/minion/scenario/recovered_followers/failed_boot/writer_tests/successors/observation.rs +++ b/crates/cellule-host/minion/scenario/recovered_followers/failed_boot/writer_tests/successors/observation.rs @@ -122,6 +122,16 @@ impl SuccessorFixture { cells: Vec, complete: bool, ) -> FleetObservation { + self.try_observation(inventory, nodes, cells, complete) + .unwrap() + } + fn try_observation( + &self, + inventory: &FleetOriginalWriterSuccessorInventory, + nodes: Vec, + cells: Vec, + complete: bool, + ) -> cellule_runtime::Result { FleetObservation::new( scope(), inventory.original().snapshot().registry(), @@ -132,7 +142,6 @@ impl SuccessorFixture { nodes, cells, ) - .unwrap() } } @@ -197,7 +206,7 @@ async fn original_writer_observation_retains_all_applications_and_refuses_duplic #[tokio::test] async fn original_writer_observation_refuses_changed_native_rows_and_complete_omission() { let fixture = fixture().await; - for field in 0..8 { + for field in 0..14 { let inventory = fixture.collect_current().await.unwrap(); let (nodes, mut cells) = fixture.observation_parts(&inventory).await; let row = &mut cells[0]; @@ -215,15 +224,35 @@ async fn original_writer_observation_refuses_changed_native_rows_and_complete_om 4 => row.observation.incarnation = IncarnationId::from_bytes([99; 16]), 5 => row.observation.code = Digest::from_bytes([99; 32]), 6 => row.observation.schema += 1, - _ => row.observation.role = CatalogRole::Blob, + 7 => row.observation.role = CatalogRole::Blob, + 8 => row.observation.owner_fence.epoch += 1, + 9 => row.observation.owner_fence.epoch = 0, + 10 => row.observation.owner_fence.incarnation = IncarnationId::from_bytes([99; 16]), + 11 => row.observation.position.as_mut().unwrap().epoch += 1, + 12 => { + row.observation.owner_fence.epoch += 1; + row.observation.position.as_mut().unwrap().epoch += 1; + } + _ => { + let incarnation = IncarnationId::from_bytes([99; 16]); + row.observation.incarnation = incarnation; + row.observation.owner_fence.incarnation = incarnation; + row.observation.position.as_mut().unwrap().incarnation = incarnation; + } } + // Inconsistent native identity is refused by construction. Coherent + // but substituted writer identity must still fail attachment against + // the original independently collected serving proof. + let expected = if matches!(field, 4 | 8..=11) { + "fleet ownership observation identity mismatch" + } else { + "original writer successor row differs" + }; + let refused = fixture + .try_observation(&inventory, nodes, cells, false) + .and_then(|observation| observation.with_original_writer_successors(inventory)); assert!( - matches!( - fixture - .observation(&inventory, nodes, cells, false) - .with_original_writer_successors(inventory), - Err(Error::Node("original writer successor row differs")) - ), + matches!(refused, Err(Error::Node(actual)) if actual == expected), "field {field}" ); } diff --git a/crates/cellule-host/minion/scenario/recovered_followers/members.rs b/crates/cellule-host/minion/scenario/recovered_followers/members.rs new file mode 100644 index 00000000..9e8d60bf --- /dev/null +++ b/crates/cellule-host/minion/scenario/recovered_followers/members.rs @@ -0,0 +1,64 @@ +//! The same authenticated native member transport for recovery fault fixtures. +use super::*; + +pub(in crate::scenario) struct Members { + pub(super) peers: Vec<(NodeId, LocalRecoveredFollowerTransport)>, + pub(super) retirements: AtomicUsize, +} +impl Members { + pub(in crate::scenario) fn new(peers: Vec<(NodeId, LocalRecoveredFollowerTransport)>) -> Self { + Self { + peers, + retirements: AtomicUsize::new(0), + } + } + + fn peer(&self, member: NodeId) -> &LocalRecoveredFollowerTransport { + &self + .peers + .iter() + .find(|(node, _)| *node == member) + .unwrap() + .1 + } +} +impl NodeLogTransport for Members { + fn append<'a>( + &'a self, + member: NodeId, + request: AppendRequest, + ) -> BoxFuture<'a, cellule_runtime::Result> { + self.peer(member).append(member, request) + } + fn seal<'a>( + &'a self, + member: NodeId, + request: SealRequest, + ) -> BoxFuture<'a, cellule_runtime::Result> { + self.peer(member).seal(member, request) + } + fn retire<'a>( + &'a self, + member: NodeId, + request: RetireRequest, + ) -> BoxFuture<'a, cellule_runtime::Result> { + self.peer(member).retire(member, request) + } + fn tail<'a>( + &'a self, + member: NodeId, + request: TailRequest, + ) -> BoxFuture<'a, cellule_runtime::Result>> { + self.peer(member).tail(member, request) + } +} +impl RecoveredNodeLogTransport for Members { + fn retire_recovered<'a>( + &'a self, + member: NodeId, + request: RecoveredRetireRequest, + ) -> BoxFuture<'a, cellule_runtime::Result> { + self.retirements.fetch_add(1, Ordering::AcqRel); + self.peer(member).retire_recovered(member, request) + } +} diff --git a/crates/cellule-host/minion/scenario/recovered_followers/mod.rs b/crates/cellule-host/minion/scenario/recovered_followers/mod.rs index 7705c341..2f453deb 100644 --- a/crates/cellule-host/minion/scenario/recovered_followers/mod.rs +++ b/crates/cellule-host/minion/scenario/recovered_followers/mod.rs @@ -31,7 +31,9 @@ use std::sync::atomic::{AtomicUsize, Ordering}; #[cfg(unix)] mod failed_boot; +mod members; mod tests; +pub(in crate::scenario) use members::Members; const NOW: i64 = 1_000_000; const CHECK: i64 = NOW + 10_005; @@ -39,61 +41,6 @@ fn deadline() -> Instant { Instant::now() + Duration::from_secs(5) } -struct Members { - peers: Vec<(NodeId, LocalRecoveredFollowerTransport)>, - retirements: AtomicUsize, -} -impl Members { - fn peer(&self, member: NodeId) -> &LocalRecoveredFollowerTransport { - &self - .peers - .iter() - .find(|(node, _)| *node == member) - .unwrap() - .1 - } -} -impl NodeLogTransport for Members { - fn append<'a>( - &'a self, - member: NodeId, - request: AppendRequest, - ) -> BoxFuture<'a, cellule_runtime::Result> { - self.peer(member).append(member, request) - } - fn seal<'a>( - &'a self, - member: NodeId, - request: SealRequest, - ) -> BoxFuture<'a, cellule_runtime::Result> { - self.peer(member).seal(member, request) - } - fn retire<'a>( - &'a self, - member: NodeId, - request: RetireRequest, - ) -> BoxFuture<'a, cellule_runtime::Result> { - self.peer(member).retire(member, request) - } - fn tail<'a>( - &'a self, - member: NodeId, - request: TailRequest, - ) -> BoxFuture<'a, cellule_runtime::Result>> { - self.peer(member).tail(member, request) - } -} -impl RecoveredNodeLogTransport for Members { - fn retire_recovered<'a>( - &'a self, - member: NodeId, - request: RecoveredRetireRequest, - ) -> BoxFuture<'a, cellule_runtime::Result> { - self.retirements.fetch_add(1, Ordering::AcqRel); - self.peer(member).retire_recovered(member, request) - } -} - struct Fixture { directory: NodeDirectory, journal: Arc, @@ -136,6 +83,16 @@ impl Fixture { Self::with_recovered_boot_at(observe, cells, frames, now, 0, 1).await } + async fn with_recovery_inputs_at_profile( + observe: bool, + cells: Vec, + frames: Vec, + now: i64, + profile: FleetProfile, + ) -> Self { + Self::with_recovered_boot_at_profile(observe, cells, frames, now, 0, 1, profile).await + } + async fn with_recovered_boot_at( observe: bool, cells: Vec, @@ -143,6 +100,27 @@ impl Fixture { now: i64, leader: usize, claimant: usize, + ) -> Self { + Self::with_recovered_boot_at_profile( + observe, + cells, + frames, + now, + leader, + claimant, + FleetProfile::default(), + ) + .await + } + + async fn with_recovered_boot_at_profile( + observe: bool, + cells: Vec, + frames: Vec, + now: i64, + leader: usize, + claimant: usize, + profile: FleetProfile, ) -> Self { assert!(leader < 3 && claimant < 3 && leader != claimant); let check = now + 10_005; @@ -152,7 +130,7 @@ impl Fixture { let root = tempfile::tempdir().unwrap(); let path = root.path().join("journal.sqlite"); let journal = Arc::new( - SqliteJournal::open(path.clone(), scope(), FleetProfile::default(), now) + SqliteJournal::open(path.clone(), scope(), profile, now) .await .unwrap(), ); diff --git a/crates/cellule-host/minion/scenario/startup/mod.rs b/crates/cellule-host/minion/scenario/startup/mod.rs index a67bd484..0e998b22 100644 --- a/crates/cellule-host/minion/scenario/startup/mod.rs +++ b/crates/cellule-host/minion/scenario/startup/mod.rs @@ -239,6 +239,10 @@ impl BootOwner { loop { if let Some(sample) = self.node.runtime().operational_sample()? && previous.is_none_or(|old| sample.sequence > old.sequence) + // A newer cached tick may still precede the intent change. + // Publish only a sample that reflects the current role gate; + // otherwise remote placement can keep selecting this donor. + && sample.mode == self.node.runtime().node_admission().mode()? { return Ok::<_, cellule_runtime::Error>(sample); } diff --git a/crates/cellule-host/minion/scenario/startup/tests/mod.rs b/crates/cellule-host/minion/scenario/startup/tests/mod.rs index 924173ad..a54bfe7b 100644 --- a/crates/cellule-host/minion/scenario/startup/tests/mod.rs +++ b/crates/cellule-host/minion/scenario/startup/tests/mod.rs @@ -486,6 +486,7 @@ async fn cordon_racing_accepted_boot_and_draining_reboot_keep_management_without node_id(0), fixture.journal.clone(), Arc::new(super::super::adapters::Cells { + receiver_directory: None, records: Arc::new(HashMap::new()), local: 0, root: fixture.root.path().into(), diff --git a/crates/cellule-host/minion/scenario/startup/tests/withdrawal.rs b/crates/cellule-host/minion/scenario/startup/tests/withdrawal.rs index 9bfa9071..f0bbe907 100644 --- a/crates/cellule-host/minion/scenario/startup/tests/withdrawal.rs +++ b/crates/cellule-host/minion/scenario/startup/tests/withdrawal.rs @@ -1,5 +1,70 @@ //! Native closing owns canonical withdrawal and the original durable boot row. use super::*; +use cellule_host::fleet::{ + FleetActionCompletion, FleetAdapterFuture, FleetCellInputs, FleetCellProvider, + FleetObservation, FleetObserver, FleetRecoveryInputs, FleetRoster, FleetTransport, +}; + +struct NoCells; + +impl FleetCellProvider for NoCells { + fn cell_inputs<'a>( + &'a self, + _: &'a MoveAttemptSpec, + ) -> FleetAdapterFuture<'a, FleetCellInputs> { + Box::pin(async { Err(invalid("Finalize fixture has no movable Cells")) }) + } + + fn recovery_inputs<'a>( + &'a self, + _: &'a MoveAttemptSpec, + ) -> FleetAdapterFuture<'a, FleetRecoveryInputs> { + Box::pin(async { Err(invalid("Finalize fixture has no recovery work")) }) + } +} + +struct UnusedObserver; + +impl FleetObserver for UnusedObserver { + fn observe<'a>( + &'a self, + _: &'a FleetRoster, + _: i64, + _: Instant, + ) -> FleetAdapterFuture<'a, FleetObservation> { + Box::pin(async { Err(invalid("closing test does not observe Cells")) }) + } +} + +struct NodeTransport { + node: Arc, + actions: std::sync::Mutex>, +} + +impl FleetTransport for NodeTransport { + fn dispatch<'a>( + &'a self, + action: &'a FleetAction, + _: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + Box::pin(async move { + self.actions.lock().unwrap().push(action.clone()); + self.node + .apply_fleet_action(action.clone(), clock()?) + .await + .map_err(|error| Box::new(error) as JournalError) + }) + } + + fn inspect<'a>( + &'a self, + _: &'a cellule_runtime::fleet::operations::FleetInspectionRequest, + _: Instant, + ) -> FleetAdapterFuture<'a, Arc> + { + Box::pin(async { Err(invalid("closing test does not inspect Cells")) }) + } +} async fn prepare(fixture: &Fixture) -> EnrollmentRecord { let original = enroll( @@ -31,6 +96,15 @@ async fn prepare(fixture: &Fixture) -> EnrollmentRecord { fixture.journal.clone(), ) .unwrap(); + fixture + .node + .install_fleet_actions( + scope(), + node_id(0), + fixture.journal.clone(), + Arc::new(NoCells), + ) + .unwrap(); original } @@ -54,6 +128,136 @@ async fn retired(fixture: &Fixture, original: &EnrollmentRecord) -> EnrollmentRe record } +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn fleet_finalize_joins_the_node_and_publishes_only_after_exact_withdrawal() { + let fixture = fixture().await; + let original = prepare(&fixture).await; + fixture.node.start().unwrap(); + + let registry = fixture + .journal + .load_snapshot(scope()) + .await + .unwrap() + .registry(); + let registry = fixture.journal.bootstrap_registry(registry).await.unwrap(); + fixture + .journal + .set_scheduling(registry, true) + .await + .unwrap(); + + let snapshot = maintenance(&fixture).await; + let now = clock().unwrap(); + let cordon = snapshot + .head() + .maintenance_action(MaintenanceAction::Cordon, now) + .unwrap(); + let cordoned = fixture.node.apply_fleet_action(cordon, now).await.unwrap(); + assert!(cordoned.committed); + assert_eq!(cordoned.outcome.outcome, FleetOutcome::Cordoned); + + let snapshot = fixture + .journal + .compare_exchange( + &snapshot, + snapshot.head().controller().unwrap().epoch, + clock().unwrap(), + &JournalTransition::Maintenance(MaintenanceEvent::Cordoned), + ) + .await + .unwrap(); + let snapshot = fixture + .journal + .compare_exchange( + &snapshot, + snapshot.head().controller().unwrap().epoch, + clock().unwrap(), + &JournalTransition::Maintenance(MaintenanceEvent::BeginEvacuation), + ) + .await + .unwrap(); + let operation = snapshot.head().maintenance().unwrap(); + crate::scenario::commit_test_role_settlement(&fixture.journal, clock().unwrap()) + .await + .unwrap(); + let snapshot = fixture + .journal + .compare_exchange( + &snapshot, + snapshot.head().controller().unwrap().epoch, + clock().unwrap(), + &JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(DrainEvidence { + node: operation.node(), + session: operation.session(), + remaining_cells: 0, + unresolved_attempts: 0, + relocated: true, + readers_settled: true, + followers_settled: true, + facilities_closed: false, + stopped: false, + withdrawn: false, + })), + ) + .await + .unwrap(); + assert_eq!( + snapshot.head().maintenance().unwrap().phase(), + MaintenancePhase::Closing + ); + let registry = fixture + .journal + .load_snapshot(scope()) + .await + .unwrap() + .registry(); + fixture + .journal + .set_scheduling(registry, false) + .await + .unwrap(); + + let transport = Arc::new(NodeTransport { + node: Arc::clone(&fixture.node), + actions: std::sync::Mutex::new(Vec::new()), + }); + let reconciler = FleetReconciler::new( + scope(), + SessionId::from_bytes([206; 16]), + FleetProfile::default(), + fixture.journal.clone(), + Arc::new(UnusedObserver), + Arc::clone(&transport) as Arc, + ) + .unwrap(); + let report = reconciler + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await + .unwrap(); + assert_eq!(report.dispatched, 1); + assert_eq!( + report.snapshot.head().maintenance().unwrap().phase(), + MaintenancePhase::Completed + ); + assert_eq!(fixture.node.state(), NodeState::Stopped); + retired(&fixture, &original).await; + + let finalize = transport.actions.lock().unwrap()[0].clone(); + let replay = fixture + .node + .apply_fleet_action(finalize, clock().unwrap()) + .await + .unwrap(); + assert!(replay.committed); + assert!(replay.execution_error.is_none()); + let FleetOutcome::Stopped(evidence) = replay.outcome.outcome else { + panic!("Finalize replay did not prove terminal shutdown: {replay:?}") + }; + assert!(evidence.facilities_closed && evidence.stopped && evidence.withdrawn); + close(fixture).await; +} + #[tokio::test] async fn native_shutdown_withdraws_and_retires_original_boot_before_stopped() { let fixture = fixture().await; diff --git a/crates/cellule-host/minion/scenario/successor_tests.rs b/crates/cellule-host/minion/scenario/successor_tests.rs index c5c34fe2..2002c031 100644 --- a/crates/cellule-host/minion/scenario/successor_tests.rs +++ b/crates/cellule-host/minion/scenario/successor_tests.rs @@ -1,6 +1,7 @@ use super::*; use cellule_host::fleet::FleetActionJournal; use cellule_runtime::fleet::operations::*; +mod routed; #[tokio::test(flavor = "multi_thread", worker_threads = 2)] async fn driver_adopts_ordinary_winner_and_joins_original_receiver_before_retirement() { @@ -26,9 +27,11 @@ async fn driver_adopts_ordinary_winner_and_joins_original_receiver_before_retire journal: journal.clone(), boots: boots.clone(), records: records.clone(), + reader_verifier: None, capture_sequence: std::sync::atomic::AtomicU64::new(0), lose_release_replies: false, lost_release_replies: std::sync::atomic::AtomicUsize::new(0), + drop_closed_finalize_replies: std::sync::atomic::AtomicUsize::new(0), expired_receiver_cleanups: std::sync::atomic::AtomicUsize::new(0), }); let driver = FleetReconciler::new( diff --git a/crates/cellule-host/minion/scenario/successor_tests/routed/fixture.rs b/crates/cellule-host/minion/scenario/successor_tests/routed/fixture.rs new file mode 100644 index 00000000..8f51e795 --- /dev/null +++ b/crates/cellule-host/minion/scenario/successor_tests/routed/fixture.rs @@ -0,0 +1,307 @@ +use super::*; + +pub(super) struct Fixture { + _root: tempfile::TempDir, + pub(super) profile: FleetProfile, + pub(super) journal: Arc, + pub(super) nodes: Vec>, + boots: Vec, + pub(super) records: Arc>, + pub(super) acknowledged: HashMap, + pub(super) fleet: Arc, + pub(super) spec: MoveAttemptSpec, + pub(super) released: PublishedPosition, + pub(super) released_attempt: MoveAttempt, +} + +impl Fixture { + pub(super) async fn released() -> Self { + Self::released_with_recovery(false).await + } + pub(super) async fn released_with_recovery(recover_receiver: bool) -> Self { + let root = tempfile::tempdir().unwrap(); + let profile = FleetProfile { + controller_lease_ms: 3_000, + reconcile_interval_ms: 500, + ..FleetProfile::default() + }; + let journal = Arc::new( + SqliteJournal::open( + root.path().join("closed-receiver-journal.sqlite"), + scope(), + profile, + clock().unwrap(), + ) + .await + .unwrap(), + ); + let mut nodes = Vec::new(); + let mut boots = Vec::new(); + let (records, acknowledged) = initialize_without_boot_withdrawal( + &root, + &journal, + &mut nodes, + &mut boots, + 60_000, + 1, + recover_receiver, + ) + .await + .unwrap(); + let fleet = Arc::new(adapters::LocalFleet { + nodes: nodes.clone(), + journal: journal.clone(), + boots: boots.clone(), + records: records.clone(), + reader_verifier: None, + capture_sequence: std::sync::atomic::AtomicU64::new(0), + lose_release_replies: false, + lost_release_replies: std::sync::atomic::AtomicUsize::new(0), + drop_closed_finalize_replies: std::sync::atomic::AtomicUsize::new(0), + expired_receiver_cleanups: std::sync::atomic::AtomicUsize::new(0), + }); + let controller = SessionId::from_bytes([206; 16]); + let driver = FleetReconciler::new( + scope(), + controller, + profile, + journal.clone(), + fleet.clone(), + fleet.clone(), + ) + .unwrap(); + let first = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await + .unwrap(); + assert!(first.snapshot.head().attempts().is_empty()); + let page = nodes[0] + .runtime() + .fleet_cells_page(None, 128) + .await + .unwrap(); + let CellInventoryEntry::Owned(row) = &page.entries()[0] else { + panic!("fixture writer missing") + }; + let spec = MoveAttemptSpec { + id: AttemptId { + operation: OperationId::from_bytes([220; 16]).unwrap(), + sequence: 1, + }, + target: row.target.clone(), + incarnation: row.incarnation, + source_node: node_id(0), + source: session(0), + generation: row.generation, + source_epoch: row.position.as_ref().unwrap().epoch, + destination_node: node_id(1), + destination: session(1), + cost: row.cost.unwrap(), + snapshot_digest: Digest::from_bytes([221; 32]), + deadline_ms: clock().unwrap() + 60_000, + }; + journal + .compare_exchange( + &first.snapshot, + first.snapshot.head().controller().unwrap().epoch, + clock().unwrap(), + &JournalTransition::Allocate(spec.clone()), + ) + .await + .unwrap(); + let version = journal.load_snapshot(scope()).await.unwrap().registry(); + journal.set_scheduling(version, false).await.unwrap(); + drop(page); + let preparing = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await + .unwrap(); + assert_eq!( + preparing.snapshot.head().attempts()[0].phase(), + AttemptPhase::Reserved + ); + let release = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await + .unwrap(); + assert_eq!(release.released, 1); + let released = release.snapshot.head().attempts()[0] + .released() + .unwrap() + .clone(); + let released_attempt = release.snapshot.head().attempts()[0].clone(); + assert_eq!( + nodes[1].stats().local_disk_reserved_bytes(), + spec.cost.disk_bytes + ); + + Self { + _root: root, + profile, + journal, + nodes, + boots, + records, + acknowledged, + fleet, + spec, + released, + released_attempt, + } + } + + pub(super) fn scratch(&self, name: &str) -> PathBuf { + self._root.path().join(name) + } + pub(super) fn directory(&self) -> cellule_runtime::node::NodeDirectory { + self.boots[0].directory.clone() + } + + pub(super) fn driver( + &self, + claimant: SessionId, + observer: Arc, + transport: Arc, + ) -> FleetReconciler { + FleetReconciler::new( + scope(), + claimant, + self.profile, + self.journal.clone(), + observer, + transport, + ) + .unwrap() + } + + pub(super) async fn reopen_journal(&self) -> Arc { + Arc::new( + SqliteJournal::open( + self._root.path().join("closed-receiver-journal.sqlite"), + scope(), + self.profile, + clock().unwrap(), + ) + .await + .unwrap(), + ) + } + + pub(super) async fn accept_original_activation(&self) -> AcceptedFleetAction { + let snapshot = self.journal.load_snapshot(scope()).await.unwrap(); + let activating = self + .journal + .compare_exchange( + &snapshot, + snapshot.head().controller().unwrap().epoch, + clock().unwrap(), + &JournalTransition::Attempt { + id: self.spec.id, + event: AttemptEvent::BeginActivate, + }, + ) + .await + .unwrap(); + let action = activating + .head() + .movement_action(self.spec.id, MovementAction::Activate, clock().unwrap()) + .unwrap(); + match self + .journal + .accept_action(&action, node_id(1), session(1), clock().unwrap()) + .await + .unwrap() + { + FleetActionAcceptance::New(accepted) + | FleetActionAcceptance::Existing { accepted, .. } => accepted, + } + } + + pub(super) async fn claim_without_actor(&self, state: ControlState) { + let accepted = self.accept_original_activation().await; + let record = &self.records[&self.spec.target.cell_id()]; + let idle = record + .authority + .load(self.spec.target.cell_id()) + .await + .unwrap() + .unwrap(); + let basis = + AcquisitionBasis::new(accepted, idle.value().clone(), clock().unwrap()).unwrap(); + self.journal.record_acquisition_basis(&basis).await.unwrap(); + // Model a crash after the canonical CAS and before actor/result + // publication. These are real validated authority transitions. + let recovering = record + .authority + .transition( + &idle, + idle.value().takeover(owner(1)).unwrap(), + Transition::Takeover, + ) + .await + .unwrap(); + if state == ControlState::Serving { + let mut serving = recovering.value().clone(); + serving.revision += 1; + serving.progress += 1; + serving.state = ControlState::Serving; + record + .authority + .transition(&recovering, serving, Transition::Activate) + .await + .unwrap(); + } else { + assert_eq!(state, ControlState::Recovering); + } + } + + pub(super) async fn close_receiver(&self) -> Arc { + crate::scenario::receiver_loss::adapters::close_receiver(self.fleet.clone()) + .await + .unwrap() + } + + pub(super) async fn acquire_rolled_back_replacement(&self) { + let record = &self.records[&self.spec.target.cell_id()]; + let idle = record + .authority + .load(record.target.cell_id()) + .await + .unwrap() + .unwrap(); + assert_eq!(idle.value().state, ControlState::Idle); + self.nodes[2] + .runtime() + .acquire_idle_restored( + record.catalog.clone(), + record.replica.clone(), + record.authority.clone(), + idle, + self._root.path().join("ordinary-receiver-resume.sqlite"), + owner(2), + ) + .await + .unwrap(); + } + + pub(super) async fn shutdown(&self) { + let nodes = &self.nodes; + let boots = &self.boots; + let journal = &self.journal; + for node in nodes { + node.shutdown().await.unwrap(); + assert_eq!(node.state(), NodeState::Stopped); + let stats = node.stats(); + assert_eq!(stats.active_cells(), 0); + assert_eq!(stats.worker_jobs(), 0); + assert_eq!(stats.retained_bytes(), 0); + assert_eq!(stats.resident_bytes(), 0); + assert_eq!(stats.file_descriptors(), 0); + assert_eq!(stats.local_disk_reserved_bytes(), 0); + } + for boot in boots { + boot.withdraw(journal).await.unwrap(); + } + journal.close().await.unwrap(); + } +} diff --git a/crates/cellule-host/minion/scenario/successor_tests/routed/mod.rs b/crates/cellule-host/minion/scenario/successor_tests/routed/mod.rs new file mode 100644 index 00000000..e74a1555 --- /dev/null +++ b/crates/cellule-host/minion/scenario/successor_tests/routed/mod.rs @@ -0,0 +1,8 @@ +//! Closed-receiver continuation through the real host and SQLite journal. +use super::*; +use crate::scenario::receiver_loss::adapters::{ClosedBootObserver, LoseRoutedActivationReply}; +use cellule_host::fleet::{FleetActionAcceptance, FleetObserver, FleetTransport}; +use cellule_runtime::control::{ControlState, Transition}; +mod fixture; +mod recovery_faults; +mod tests; diff --git a/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/fixture.rs b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/fixture.rs new file mode 100644 index 00000000..449c8066 --- /dev/null +++ b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/fixture.rs @@ -0,0 +1,279 @@ +use super::*; + +pub(super) struct FaultFixture { + pub(super) native: super::super::fixture::Fixture, + pub(super) observer: Arc, + transport: Arc, + pub(super) accepted: AcceptedFleetAction, + pub(super) original: Control, + resume: tokio::sync::oneshot::Sender<()>, + pass: tokio::task::JoinHandle>, +} + +impl FaultFixture { + pub(super) async fn paused( + write: RecoveryWrite, + boundary: RecoveryWriteBoundary, + fail: bool, + ) -> Self { + let native = super::super::fixture::Fixture::released_with_recovery(true).await; + native.claim_without_actor(ControlState::Serving).await; + // Keep the shared finite constructor out of every caller's inline + // future; nested native observation exceeds normal debug test stacks. + Box::pin(Self::pause_claim(native, write, boundary, fail)).await + } + + pub(super) async fn pause_claim( + native: super::super::fixture::Fixture, + write: RecoveryWrite, + boundary: RecoveryWriteBoundary, + fail: bool, + ) -> Self { + let original = native.records[&native.spec.target.cell_id()] + .authority + .load(native.spec.target.cell_id()) + .await + .unwrap() + .unwrap() + .value() + .clone(); + let observer = native.close_receiver().await; + let transport = Arc::new(CapturedTransport { + fleet: native.fleet.clone(), + activation: Mutex::new(None), + }); + let (entered, resume) = native + .journal + .pause_receiver_recovery_write(write, boundary, fail); + let driver = native.driver( + SessionId::from_bytes([206; 16]), + observer.clone(), + transport.clone(), + ); + let pass = tokio::spawn(async move { + driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await + }); + tokio::time::timeout(Duration::from_secs(5), entered) + .await + .unwrap() + .unwrap(); + let accepted = native + .journal + .load_movement_actions(scope(), &native.released_attempt, MovementAction::Activate) + .await + .unwrap() + .into_iter() + .find_map(|value| match value { + FleetActionAcceptance::New(accepted) + | FleetActionAcceptance::Existing { accepted, .. } + if accepted.action().receiver_endpoint() == Some((node_id(2), session(2))) => + { + Some(accepted) + } + _ => None, + }) + .unwrap(); + assert!(accepted.action().receiver_route().is_some()); + Self { + native, + observer, + transport, + accepted, + original, + resume, + pass, + } + } + + pub(super) async fn current(&self) -> Control { + self.native.records[&self.native.spec.target.cell_id()] + .authority + .load(self.native.spec.target.cell_id()) + .await + .unwrap() + .unwrap() + .value() + .clone() + } + + pub(super) async fn release(self, cancel_waiter: bool) -> CompletedFixture { + if cancel_waiter { + self.pass.abort(); + assert!(self.pass.await.unwrap_err().is_cancelled()); + self.resume.send(()).unwrap(); + } else { + self.resume.send(()).unwrap(); + let report = self.pass.await.unwrap().unwrap(); + assert_eq!( + report.snapshot.head().reserved_restore_bytes(), + self.native.spec.cost.disk_bytes + ); + } + CompletedFixture { + native: self.native, + observer: self.observer, + accepted: self.accepted, + original: self.original, + completion: self.transport.activation.lock().unwrap().clone(), + } + } +} + +pub(super) struct CompletedFixture { + pub(super) native: super::super::fixture::Fixture, + observer: Arc, + pub(super) accepted: AcceptedFleetAction, + pub(super) original: Control, + pub(super) completion: Option>, +} +impl CompletedFixture { + pub(super) async fn current(&self) -> cellule_runtime::control::authority::VersionedControl { + self.native.records[&self.native.spec.target.cell_id()] + .authority + .load(self.native.spec.target.cell_id()) + .await + .unwrap() + .unwrap() + } + pub(super) async fn replay(&self) -> Arc { + self.native.nodes[2] + .apply_fleet_action(self.accepted.action().clone(), clock().unwrap()) + .await + .unwrap() + } + pub(super) async fn finish(self) { + Box::pin(self.finish_inner(None, false)).await; + } + pub(super) async fn finish_value(self, expected_value: Option) { + Box::pin(self.finish_inner(expected_value, false)).await; + } + pub(super) async fn finish_after_controller_loss(self, expected_value: i64) { + Box::pin(self.finish_inner(Some(expected_value), true)).await; + } + async fn finish_inner(self, expected_value: Option, replace_controller: bool) { + let client = self.native.reopen_journal().await; + let basis = self + .native + .journal + .load_receiver_recovery_basis(&self.accepted) + .await + .unwrap() + .unwrap(); + let evidence = self + .native + .journal + .load_receiver_recovery_evidence(&self.accepted) + .await + .unwrap() + .unwrap(); + assert_eq!(basis.control(), &self.original); + assert_eq!(evidence.basis(), &basis); + assert_eq!( + client + .load_receiver_recovery_basis(&self.accepted) + .await + .unwrap(), + Some(basis) + ); + assert_eq!( + client + .load_receiver_recovery_evidence(&self.accepted) + .await + .unwrap(), + Some(evidence) + ); + let old_snapshot = client.load_snapshot(scope()).await.unwrap(); + let old_lease = old_snapshot.head().controller().unwrap(); + let old_epoch = old_lease.epoch; + let old_driver = self.native.driver( + SessionId::from_bytes([206; 16]), + self.observer.clone(), + self.native.fleet.clone(), + ); + if replace_controller { + let wait = old_lease.expires_at_ms.saturating_sub(clock().unwrap()) + 1; + tokio::time::sleep(Duration::from_millis(u64::try_from(wait.max(0)).unwrap())).await; + } + let claimant = SessionId::from_bytes([if replace_controller { 207 } else { 206 }; 16]); + let driver = FleetReconciler::new( + scope(), + claimant, + self.native.profile, + client.clone(), + self.observer.clone(), + self.native.fleet.clone(), + ) + .unwrap(); + let mut report = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await + .unwrap(); + if replace_controller { + assert!(report.snapshot.head().controller().unwrap().epoch > old_epoch); + assert_eq!( + report.snapshot.head().controller().unwrap().claimant, + claimant + ); + let before = client.load_snapshot(scope()).await.unwrap(); + let error = old_driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(1)) + .await + .unwrap_err(); + assert!( + matches!(&error, cellule_runtime::Error::Facility { name: "fleet-journal", source } + if matches!(source.downcast_ref::(), Some(OperationError::Fenced))), + "{error:?}" + ); + assert_eq!(client.load_snapshot(scope()).await.unwrap(), before); + } + for _ in 0..12 { + if report.snapshot.head().attempts().is_empty() { + break; + } + report = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await + .unwrap(); + } + assert!(report.snapshot.head().attempts().is_empty(), "{report:?}"); + assert_eq!(report.snapshot.head().reserved_restore_bytes(), 0); + let record = &self.native.records[&self.native.spec.target.cell_id()]; + let current = record + .authority + .load(record.target.cell_id()) + .await + .unwrap() + .unwrap(); + let handle = self.native.nodes[2] + .runtime() + .local_handle(record.catalog.clone(), ¤t) + .await + .unwrap() + .unwrap(); + let original = &self.native.acknowledged[&record.target.cell_id()]; + assert_eq!( + handle + .resolve(original.identity, original.digest, clock().unwrap(), 64) + .await + .unwrap(), + Resolution::Committed(original.outcome.clone()) + ); + let bytes = handle + .query(64, 64, |connection| { + let value: i64 = + connection.query_row("SELECT value FROM counter", [], |row| row.get(0))?; + Ok(value.to_be_bytes().to_vec()) + }) + .await + .unwrap(); + assert_eq!( + bytes, + expected_value.unwrap_or(original.value).to_be_bytes() + ); + assert_eq!(self.native.nodes[2].stats().active_cells(), 1); + self.native.shutdown().await; + client.close().await.unwrap(); + } +} diff --git a/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/inherited/mod.rs b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/inherited/mod.rs new file mode 100644 index 00000000..9cfd8654 --- /dev/null +++ b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/inherited/mod.rs @@ -0,0 +1,36 @@ +//! Actual sealed suffix inherited through an interrupted native receiver claim. +use super::*; +use crate::scenario::recovered_followers::Members; +use bytes::Bytes; +use cellule_host::fleet::{ + FleetEnrollmentAcceptance, FleetRecoveredFollowerRetirement, FleetRoster, +}; +use cellule_runtime::fleet::operations::{EnrollmentEndpoint, EnrollmentRole, EnrollmentSpec}; +use cellule_runtime::{ + Error, + follower::FollowerStore, + node::{ + NodeAdvertisement, + log_recovery::{ + NodeLogRecovery, RecoveryCell, RecoveryCoordinator, + retirement::retire_recovered_members, + }, + log_transport::{ + AppendRequest, LocalFollowerTransport, LocalRecoveredFollowerTransport, + NodeLogTransport, + }, + }, + recovery::manifest::RecoveryManifestStore, +}; +use fixture::FaultFixture; + +mod setup; +mod tests; + +struct Inherited { + native: super::super::fixture::Fixture, + path: ObjectPath, + manifest: Bytes, + original: Control, + value: i64, +} diff --git a/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/inherited/setup.rs b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/inherited/setup.rs new file mode 100644 index 00000000..c37de3e9 --- /dev/null +++ b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/inherited/setup.rs @@ -0,0 +1,454 @@ +use super::*; + +impl Inherited { + pub(super) async fn new() -> Self { + let native = super::super::super::fixture::Fixture::released_with_recovery(true).await; + let record = &native.records[&native.spec.target.cell_id()]; + let early_session = session(9); + let early_owner = owner(9); + let early = CellNodeBuilder::new(application::compile().unwrap()) + .with_session(early_session) + .with_runtime(SqlWorkerPool::new(2, 8).unwrap(), 128 << 20) + .with_replica_host(Host::default().with_local_disk_budget(DiskBudget::new(8 << 30))) + .build() + .unwrap(); + early + .install_task_group(CancellationToken::new(), CancellationToken::new()) + .unwrap(); + let guard = NodeLeaseGuard::new(clock().unwrap(), clock().unwrap() + 60_000).unwrap(); + early.install_node_lease(guard.clone()).unwrap(); + let idle = record + .authority + .load(record.target.cell_id()) + .await + .unwrap() + .unwrap(); + let handle = early + .runtime() + .acquire_idle_restored( + record.catalog.clone(), + record.replica.clone(), + record.authority.clone(), + idle, + native.scratch("earlier-owner.sqlite"), + early_owner, + ) + .await + .unwrap(); + let observed = record + .authority + .load(record.target.cell_id()) + .await + .unwrap() + .unwrap(); + assert_eq!(observed.value().epoch, native.released.epoch + 1); + let predecessor = observed.value().ltx_root().unwrap(); + let frames = frames(&native, &observed).await; + let directory = native.directory(); + native + .journal + .register_initial_intent( + &NodeIntent::initial(scope(), node_id(9), early_session).unwrap(), + ) + .await + .unwrap(); + // Original member requests are journaled before canonical enrollment. + // These native in-process stores and their complete retired enrollment + // set are joined before the later receiver's process closure. + for index in [1, 2] { + let current = directory + .load_if_live(session(index), clock().unwrap()) + .await + .unwrap() + .unwrap(); + let ad = current.advertisement(); + let previous = ad.operational_sample().unwrap(); + let sample = tokio::time::timeout(Duration::from_secs(3), async { + loop { + if let Some(sample) = + native.nodes[index].runtime().operational_sample().unwrap() + && sample.sequence > previous.sequence + { + break sample; + } + tokio::time::sleep(Duration::from_millis(10)).await; + } + }) + .await + .unwrap(); + let mut capacity = ad.capacity(); + assert!( + capacity.free_memory_bytes != 0 + && capacity.free_disk_bytes != 0 + && capacity.job_credits != 0, + "inherited follower node={index} capacity={capacity:?}" + ); + capacity.follower_free_bytes = capacity.free_disk_bytes.min(64 << 20); + capacity.log_protocol = cellule_runtime::node::NODE_LOG_PROTOCOL_VERSION; + let next = NodeAdvertisement::sign( + ad.node(), + ad.session(), + ad.endpoint().into(), + ad.fleet(), + ad.certificate(), + ad.image(), + ad.release(), + &ed25519_dalek::SigningKey::from_bytes(&[index as u8 + 1; 32]), + ad.progress(), + clock().unwrap(), + clock().unwrap() + 30_000, + ad.module_digests().to_vec(), + ad.peer_versions().to_vec(), + ad.failure_domain().clone(), + capacity, + ) + .unwrap() + .with_operational_placement( + ad.placement_capacity().unwrap(), + sample, + &ed25519_dalek::SigningKey::from_bytes(&[index as u8 + 1; 32]), + ) + .unwrap(); + directory + .refresh(¤t, next, clock().unwrap()) + .await + .unwrap(); + } + let reference = directory + .load_if_live(session(0), clock().unwrap()) + .await + .unwrap() + .unwrap(); + let ad = reference.advertisement(); + let now = clock().unwrap(); + let expires = now + 1_000; + let leader = directory + .create( + NodeAdvertisement::sign( + node_id(9), + early_session, + owner(9).endpoint, + ad.fleet(), + ad.certificate(), + ad.image(), + ad.release(), + &ed25519_dalek::SigningKey::from_bytes(&[10; 32]), + 1, + now, + expires, + ad.module_digests().to_vec(), + ad.peer_versions().to_vec(), + ad.failure_domain().clone(), + ad.capacity(), + ) + .unwrap(), + now, + ) + .await + .unwrap(); + let prepared = directory + .prepare_log_enrollment(&leader, 7, 1, 128, clock().unwrap()) + .await + .unwrap() + .unwrap(); + let enrollment = directory + .prepare_log_enrollment_attempt(&prepared, clock().unwrap()) + .await + .unwrap(); + let snapshot = native.journal.load_snapshot(scope()).await.unwrap(); + let roster = FleetRoster::collect( + native.journal.as_ref(), + &snapshot, + Instant::now() + Duration::from_secs(5), + ) + .await + .unwrap(); + for (index, member) in prepared.followers().iter().enumerate() { + let intent = roster + .intents() + .iter() + .find(|intent| intent.node() == member.node()) + .unwrap(); + let accepted = native + .journal + .accept_enrollment( + &EnrollmentSpec { + scope: scope(), + request: Digest::from_bytes([index as u8 + 248; 32]), + source: Some(EnrollmentEndpoint { + node: node_id(9), + session: early_session, + intent_revision: 1, + }), + target: EnrollmentEndpoint { + node: member.node(), + session: member.session(), + intent_revision: intent.revision(), + }, + role: EnrollmentRole::Follower { log_epoch: 7 }, + }, + clock().unwrap(), + ) + .await + .unwrap(); + assert!(matches!(accepted, FleetEnrollmentAcceptance::New(_))); + } + let enrolled = directory + .commit_log_enrollment(&enrollment, clock().unwrap()) + .await + .unwrap() + .enrollment() + .clone(); + let mut peers = Vec::new(); + for member in enrolled.advertisement().log().unwrap().members() { + let store = FollowerStore::open( + native.scratch(&format!("inherited-follower-{member:?}")), + record.replica.limits(), + DiskBudget::new(8 << 30), + ) + .unwrap(); + let local = LocalFollowerTransport::new(*member, store); + local + .append( + *member, + AppendRequest { + leader_session: early_session, + log_epoch: 7, + frames: frames.clone(), + covered_through: 0, + }, + ) + .await + .unwrap(); + peers.push(( + *member, + LocalRecoveredFollowerTransport::new(local, directory.clone(), session(1), clock) + .unwrap(), + )); + } + let transport = Arc::new(Members::new(peers)); + // Activate only after every original member's first append has fsynced. + directory + .activate_log(&enrolled, clock().unwrap()) + .await + .unwrap(); + guard.fence(); + assert!(matches!( + handle.query(1, 1, |_| Ok(Vec::new())).await, + Err(Error::Fenced) | Err(Error::CellDraining) + )); + let remaining = (expires.saturating_sub(clock().unwrap()) + 1).max(0); + tokio::time::sleep(Duration::from_millis(u64::try_from(remaining).unwrap())).await; + let fenced = directory + .claim_expired(early_session, session(1), clock().unwrap()) + .await + .unwrap(); + let manifests = + RecoveryManifestStore::new(record.authority.layout().clone(), record.replica.limits()); + let completed = RecoveryCoordinator::new( + NodeLogRecovery::from_fenced(transport.clone(), &fenced, record.replica.limits()) + .unwrap(), + manifests.clone(), + ) + .recover_and_seal( + &directory, + fenced, + vec![RecoveryCell { + application: scope().application, + authority: record.authority.clone(), + observed, + }], + clock().unwrap(), + ) + .await + .unwrap(); + let retirement = retire_recovered_members(transport, &completed.sealed) + .await + .unwrap() + .confirmed() + .unwrap(); + directory + .retire_recovered_log(&retirement, session(1), clock().unwrap()) + .await + .unwrap(); + let snapshot = native.journal.load_snapshot(scope()).await.unwrap(); + let roster = FleetRoster::collect( + native.journal.as_ref(), + &snapshot, + Instant::now() + Duration::from_secs(5), + ) + .await + .unwrap(); + FleetRecoveredFollowerRetirement::capture( + native.journal.as_ref(), + &directory, + &roster, + &completed.sealed, + session(1), + Instant::now() + Duration::from_secs(5), + clock, + ) + .await + .unwrap() + .publish( + native.journal.as_ref(), + &directory, + session(1), + Instant::now() + Duration::from_secs(5), + clock, + ) + .await + .unwrap() + .confirmed() + .unwrap(); + let original = completed.controls[0].value().clone(); + let overlay = original.recovery.as_ref().unwrap(); + assert_eq!(overlay.predecessor.digest.as_bytes(), &predecessor.digest); + let path = record.authority.layout().node_log_recovery_path( + early_session.as_bytes(), + 7, + overlay.manifest_digest.as_bytes(), + ); + let manifest = record + .authority + .layout() + .store() + .get_with_etag(&path) + .await + .unwrap() + .0; + record + .authority + .layout() + .store() + .delete(&path) + .await + .unwrap(); + let failure = native.nodes[1] + .runtime() + .takeover_restored( + record.catalog.clone(), + record.replica.clone(), + record.authority.clone(), + completed.controls[0].clone(), + completed.takeover, + manifests, + native.scratch("interrupted-preferred-receiver.sqlite"), + owner(1), + ) + .await + .err() + .unwrap(); + assert!(matches!(failure, Error::Storage(_)), "{failure:?}"); + let claimed = record + .authority + .load(record.target.cell_id()) + .await + .unwrap() + .unwrap(); + assert_eq!(claimed.value().epoch, original.epoch + 1); + assert_eq!(claimed.value().state, ControlState::Recovering); + assert_eq!(claimed.value().recovery, original.recovery); + assert_eq!(claimed.value().owner.as_ref().unwrap().session, session(1)); + assert!( + record + .authority + .acquisition_record(original.cell, original.incarnation, claimed.value().epoch) + .await + .unwrap() + .is_none() + ); + // The original prepared receiver still owns its one admitted slot; + // the interrupted ordinary claim admitted no writer alongside it. + assert_eq!(native.nodes[1].stats().active_cells(), 1); + let inventory = native.nodes[1] + .runtime() + .fleet_cells_page(None, 128) + .await + .unwrap(); + assert!( + inventory + .entries() + .iter() + .all(|row| !matches!(row, CellInventoryEntry::Owned(_))) + ); + drop(inventory); + early.shutdown().await.unwrap(); + assert_eq!(early.stats().active_cells(), 0); + assert_eq!(early.stats().worker_jobs(), 0); + assert_eq!(early.stats().resident_bytes(), 0); + assert_eq!(early.stats().retained_bytes(), 0); + assert_eq!(early.stats().file_descriptors(), 0); + assert_eq!(early.stats().blocking_jobs(), 0); + assert_eq!(early.stats().recovery_jobs(), 0); + assert_eq!(early.stats().io_slots(), 0); + assert_eq!(early.stats().local_disk_reserved_bytes(), 0); + record + .authority + .layout() + .store() + .create_strict(&path, manifest.clone()) + .await + .unwrap(); + let value = native.acknowledged[&record.target.cell_id()].value + 1; + Self { + native, + path, + manifest, + original, + value, + } + } +} + +async fn frames( + native: &super::super::super::fixture::Fixture, + observed: &cellule_runtime::control::authority::VersionedControl, +) -> Vec { + let record = &native.records[&native.spec.target.cell_id()]; + let predecessor = observed.value().ltx_root().unwrap(); + let path = native.scratch("inherited-actual-tail.sqlite"); + let writable = record + .replica + .open_root(&predecessor) + .await + .unwrap() + .paged() + .prepare_writable(&path) + .await + .unwrap(); + let mut writer = writable.open_writable(&path).unwrap(); + writer.transaction(|tx| { + tx.execute("UPDATE counter SET value = value + 1", [])?; + tx.execute("UPDATE sys_meta SET commit_sequence = commit_sequence + 1, logical_time_ms = logical_time_ms + 1 WHERE singleton = 1", [])?; + Ok(()) + }).unwrap(); + let capture = writer.capture().unwrap(); + let frames = capture + .segments + .iter() + .enumerate() + .map(|(index, segment)| { + cellule_runtime::ltx::encode_node_frame( + cellule_runtime::ltx::NodeFrameScope { + leader_session: *session(9).as_bytes(), + log_epoch: 7, + node_sequence: index as u64 + 1, + application: *scope().application.as_bytes(), + cell: *record.target.cell_id().as_bytes(), + incarnation: *record.incarnation.as_bytes(), + cell_epoch: observed.value().epoch, + commit_sequence: predecessor.commit_sequence + 1, + }, + segment.info().clone(), + Bytes::from(std::fs::read(segment.path()).unwrap()), + record.replica.limits(), + ) + .unwrap() + .encoded() + .clone() + }) + .collect(); + writer.close().unwrap(); + frames +} diff --git a/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/inherited/tests.rs b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/inherited/tests.rs new file mode 100644 index 00000000..87cae560 --- /dev/null +++ b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/inherited/tests.rs @@ -0,0 +1,265 @@ +use super::*; + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn routed_inherited_suffix_survives_missing_evidence_and_idle_reacquisition() { + let inherited = Inherited::new().await; + let original = inherited.original.clone(); + let value = inherited.value; + let fixture = FaultFixture::pause_claim( + inherited.native, + RecoveryWrite::Evidence, + RecoveryWriteBoundary::BeforeCommit, + true, + ) + .await + .release(false) + .await; + assert_original_error(fixture.completion.as_ref().unwrap()); + assert_eq!(fixture.original.epoch, original.epoch + 1); + assert_eq!(fixture.original.recovery, original.recovery); + let idle = fixture.current().await; + assert_eq!(idle.value().state, ControlState::Idle); + assert!(idle.value().recovery.is_none()); + assert_eq!(idle.value().epoch, original.epoch + 2); + assert_ne!(idle.value().root, original.root); + let completion = fixture.replay().await; + assert!( + completion.committed && completion.execution_error.is_none(), + "{completion:?}" + ); + assert!(matches!( + completion.outcome.outcome, + FleetOutcome::Activated(_) + )); + assert_eq!( + fixture.current().await.value().epoch, + idle.value().epoch + 1 + ); + let record = &fixture.native.records[&idle.value().cell]; + let overlay = original.recovery.as_ref().unwrap(); + let manifest = + RecoveryManifestStore::new(record.authority.layout().clone(), record.replica.limits()) + .load_manifest( + overlay.leader_session, + overlay.log_epoch, + overlay.manifest_digest, + ) + .await + .unwrap(); + let required = &manifest.cells()[0]; + assert_eq!(required.cell_epoch, original.epoch); + let current = fixture.current().await; + let proof = fixture.native.nodes[2] + .runtime() + .verify_recovered_prefix( + &record.catalog, + &record.authority, + record.replica.clone(), + required, + current.value().ltx_root().unwrap(), + 128, + ) + .await + .unwrap(); + assert_eq!(proof.acquisition_epoch(), idle.value().epoch); + assert_eq!(proof.required().cell_epoch, original.epoch); + assert!( + record + .authority + .acquisition_record(original.cell, original.incarnation, original.epoch + 1) + .await + .unwrap() + .is_none() + ); + fixture.finish_value(Some(value)).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn routed_inherited_manifest_faults_keep_idle_charged_then_a_new_controller_joins() { + for corrupt in [false, true] { + let inherited = Inherited::new().await; + let fixture = FaultFixture::pause_claim( + inherited.native, + RecoveryWrite::Evidence, + RecoveryWriteBoundary::BeforeCommit, + true, + ) + .await + .release(false) + .await; + let idle = fixture.current().await.value().clone(); + let record = &fixture.native.records[&idle.cell]; + let layout = record.authority.layout(); + layout.store().delete(&inherited.path).await.unwrap(); + if corrupt { + layout + .store() + .create_strict( + &inherited.path, + Bytes::from_static(b"corrupt inherited routed manifest"), + ) + .await + .unwrap(); + } + let refused = fixture.replay().await; + assert!(refused.committed && matches!(refused.outcome.outcome, FleetOutcome::Unknown)); + let error = refused.execution_error.as_ref().unwrap().as_ref(); + if corrupt { + assert!(matches!(error, Error::Node(_)), "{error:?}"); + } else { + assert!(matches!(error, Error::Storage(_)), "{error:?}"); + } + assert_eq!(fixture.current().await.value(), &idle); + assert_eq!(fixture.native.nodes[2].stats().active_cells(), 0); + assert_eq!(fixture.native.nodes[2].stats().file_descriptors(), 0); + assert_eq!( + fixture + .native + .journal + .load_snapshot(scope()) + .await + .unwrap() + .head() + .reserved_restore_bytes(), + fixture.native.spec.cost.disk_bytes + ); + let evidence = fixture + .native + .journal + .load_receiver_recovery_evidence(&fixture.accepted) + .await + .unwrap() + .unwrap(); + assert_eq!( + evidence.basis().control().recovery, + inherited.original.recovery + ); + assert_eq!( + evidence.basis().control().epoch, + inherited.original.epoch + 1 + ); + if corrupt { + layout.store().delete(&inherited.path).await.unwrap(); + } + layout + .store() + .create_strict(&inherited.path, inherited.manifest) + .await + .unwrap(); + let completion = fixture.replay().await; + assert!( + completion.committed && completion.execution_error.is_none(), + "{completion:?}" + ); + assert_eq!( + fixture + .native + .journal + .load_receiver_recovery_evidence(&fixture.accepted) + .await + .unwrap(), + Some(evidence) + ); + fixture.finish_after_controller_loss(inherited.value).await; + } +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn routed_inherited_claim_resumes_without_cas_after_manifest_outage_and_cancelled_waiter() { + let inherited = Inherited::new().await; + let fixture = FaultFixture::pause_claim( + inherited.native, + RecoveryWrite::Basis, + RecoveryWriteBoundary::AfterCommit, + true, + ) + .await + .release(false) + .await; + let before = fixture.current().await; + let record = &fixture.native.records[&before.value().cell]; + let layout = record.authority.layout(); + layout.store().delete(&inherited.path).await.unwrap(); + let failed = fixture.replay().await; + assert!(failed.committed && matches!(failed.outcome.outcome, FleetOutcome::Unknown)); + assert!(matches!( + failed.execution_error.as_ref().unwrap().as_ref(), + Error::Storage(_) + )); + let claimed = fixture.current().await.value().clone(); + assert_eq!(claimed.state, ControlState::Recovering); + assert_eq!(claimed.epoch, before.value().epoch + 1); + assert_eq!(claimed.recovery, before.value().recovery); + assert_eq!(fixture.native.nodes[2].stats().active_cells(), 0); + assert!( + record + .authority + .acquisition_record(claimed.cell, claimed.incarnation, claimed.epoch) + .await + .unwrap() + .is_none() + ); + layout + .store() + .create_strict(&inherited.path, inherited.manifest) + .await + .unwrap(); + let (entered, resume) = fixture.native.journal.pause_receiver_recovery_write( + RecoveryWrite::Evidence, + RecoveryWriteBoundary::BeforeCommit, + false, + ); + let node = fixture.native.nodes[2].clone(); + let action = fixture.accepted.action().clone(); + let waiter = tokio::spawn(async move { + node.apply_fleet_action(action, clock().unwrap()) + .await + .unwrap() + }); + tokio::time::timeout(Duration::from_secs(5), entered) + .await + .unwrap() + .unwrap(); + let canonical = record + .authority + .acquisition_record(claimed.cell, claimed.incarnation, claimed.epoch) + .await + .unwrap() + .unwrap(); + assert_eq!(canonical.input(), &fixture.original); + assert_eq!(canonical.materialized().epoch, claimed.epoch); + assert!(canonical.materialized().recovery.is_none()); + let inventory = fixture.native.nodes[2] + .runtime() + .fleet_cells_page(None, 128) + .await + .unwrap(); + assert!( + inventory + .entries() + .iter() + .all(|row| !matches!(row, CellInventoryEntry::Owned(_))) + ); + drop(inventory); + waiter.abort(); + assert!(waiter.await.unwrap_err().is_cancelled()); + resume.send(()).unwrap(); + let completed = fixture.replay().await; + assert!( + completed.committed && completed.execution_error.is_none(), + "{completed:?}" + ); + assert_eq!(fixture.current().await.value().epoch, claimed.epoch); + assert_eq!( + fixture + .native + .journal + .load_receiver_recovery_evidence(&fixture.accepted) + .await + .unwrap() + .unwrap() + .restored(), + canonical.materialized() + ); + fixture.finish_after_controller_loss(inherited.value).await; +} diff --git a/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/mod.rs b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/mod.rs new file mode 100644 index 00000000..3cd1b22e --- /dev/null +++ b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/mod.rs @@ -0,0 +1,60 @@ +//! Real receiver takeover faults at the durable journal boundaries. +use super::*; +use crate::journal::{RecoveryWrite, RecoveryWriteBoundary}; +use cellule_host::fleet::{FleetActionCompletion, FleetAdapterFuture, FleetReconcileReport}; +use cellule_runtime::control::Control; +use std::sync::Mutex; + +mod fixture; +mod inherited; +mod tests; + +struct CapturedTransport { + fleet: Arc, + activation: Mutex>>, +} +impl FleetTransport for CapturedTransport { + fn dispatch<'a>( + &'a self, + action: &'a FleetAction, + deadline: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + Box::pin(async move { + let result = self.fleet.dispatch(action, deadline).await?; + if action.receiver_route().is_some() + && matches!( + action.kind(), + FleetActionKind::Movement { + action: MovementAction::Activate, + .. + } + ) + { + *self.activation.lock().unwrap() = Some(result.clone()); + } + Ok(result) + }) + } + fn inspect<'a>( + &'a self, + request: &'a FleetInspectionRequest, + deadline: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + self.fleet.inspect(request, deadline) + } +} + +fn assert_original_error(completion: &FleetActionCompletion) { + let error = completion.execution_error.as_ref().unwrap(); + let mut source: &dyn std::error::Error = error; + loop { + if let Some(io) = source.downcast_ref::() { + assert_eq!( + io.to_string(), + "injected receiver recovery write reply failure" + ); + break; + } + source = source.source().expect("original journal error was lost"); + } +} diff --git a/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/tests/continuation.rs b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/tests/continuation.rs new file mode 100644 index 00000000..c5e0fe54 --- /dev/null +++ b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/tests/continuation.rs @@ -0,0 +1,236 @@ +use super::*; + +#[derive(Clone)] +pub(super) enum CanonicalFault { + Missing, + Corrupt, + Substituted { body: bytes::Bytes, input: Control }, +} + +pub(super) async fn evidence_failure( + boundary: RecoveryWriteBoundary, + fault: Option, + ordinary_winner: bool, +) { + let fixture = FaultFixture::paused(RecoveryWrite::Evidence, boundary, true).await; + let basis = fixture + .native + .journal + .load_receiver_recovery_basis(&fixture.accepted) + .await + .unwrap() + .unwrap(); + assert_eq!(basis.control(), &fixture.original); + let evidence = fixture + .native + .journal + .load_receiver_recovery_evidence(&fixture.accepted) + .await + .unwrap(); + assert_eq!( + evidence.is_some(), + boundary == RecoveryWriteBoundary::AfterCommit + ); + assert_reserved_without_actor(&fixture.native).await; + let fixture = fixture.release(false).await; + let failed = fixture.completion.as_ref().unwrap(); + assert!(failed.committed && matches!(failed.outcome.outcome, FleetOutcome::Unknown)); + assert_original_error(failed); + let idle = fixture.current().await; + assert_eq!(idle.value().state, ControlState::Idle); + assert!(idle.value().owner.is_none()); + assert_eq!(idle.value().epoch, fixture.original.epoch + 1); + assert_eq!(idle.value().root, fixture.original.root); + assert_eq!(fixture.native.nodes[2].stats().active_cells(), 0); + assert_eq!(fixture.native.nodes[2].stats().worker_jobs(), 0); + assert_eq!(fixture.native.nodes[2].stats().file_descriptors(), 0); + assert_eq!( + fixture.native.nodes[2].stats().local_disk_reserved_bytes(), + 0 + ); + let record = &fixture.native.records[&fixture.native.spec.target.cell_id()]; + let canonical = record + .authority + .acquisition_record( + record.target.cell_id(), + record.incarnation, + idle.value().epoch, + ) + .await + .unwrap() + .unwrap(); + assert_eq!(canonical.input(), &fixture.original); + if let Some(evidence) = &evidence { + assert_eq!(evidence.restored(), canonical.materialized()); + } + let snapshot = fixture.native.journal.load_snapshot(scope()).await.unwrap(); + let inspection = FleetInspectionRequest::new( + snapshot + .head() + .movement_action( + fixture.native.spec.id, + MovementAction::Inspect, + clock().unwrap(), + ) + .unwrap(), + snapshot.registry(), + Digest::from_bytes([229; 32]), + node_id(2), + session(2), + clock().unwrap() + 1_000, + ) + .unwrap(); + assert!( + fixture.native.nodes[2] + .inspect_fleet_action(inspection) + .await + .is_err() + ); + assert_eq!(fixture.current().await.value(), idle.value()); + assert_eq!( + fixture + .native + .journal + .load_receiver_recovery_evidence(&fixture.accepted) + .await + .unwrap(), + evidence + ); + assert!( + fixture + .native + .journal + .load_acquisition_basis(&fixture.accepted) + .await + .unwrap() + .is_none() + ); + assert_eq!( + fixture + .native + .journal + .load_snapshot(scope()) + .await + .unwrap() + .head() + .reserved_restore_bytes(), + fixture.native.spec.cost.disk_bytes + ); + // Replay normally resumes the safe rollback root itself. Also race an + // ordinary acquisition winner against replay without rewriting history. + if ordinary_winner { + fixture.native.acquire_rolled_back_replacement().await; + } + if let Some(fault) = fault { + let layout = record.authority.layout(); + let path = layout.acquisition_record_path( + record.target.cell_id().as_bytes(), + record.incarnation.as_bytes(), + canonical.materialized().epoch, + ); + let (original, _) = layout.store().get_with_etag(&path).await.unwrap(); + let before = fixture.current().await.value().clone(); + layout.store().delete(&path).await.unwrap(); + match &fault { + CanonicalFault::Missing => {} + CanonicalFault::Corrupt => { + layout + .store() + .create_strict( + &path, + bytes::Bytes::from_static(b"corrupt-receiver-acquisition"), + ) + .await + .unwrap(); + } + CanonicalFault::Substituted { body, input } => { + assert_ne!(input, &fixture.original); + layout + .store() + .create_strict(&path, body.clone()) + .await + .unwrap(); + assert_eq!( + record + .authority + .acquisition_record( + record.target.cell_id(), + record.incarnation, + canonical.materialized().epoch + ) + .await + .unwrap() + .unwrap() + .input(), + input + ); + } + } + let refused = fixture.replay().await; + assert!(refused.committed && matches!(refused.outcome.outcome, FleetOutcome::Unknown)); + let error = refused.execution_error.as_ref().unwrap().as_ref(); + match &fault { + CanonicalFault::Missing => assert!( + matches!(error, cellule_runtime::Error::AcquisitionHistoryIncomplete { epoch, .. } if *epoch == canonical.materialized().epoch) + ), + CanonicalFault::Corrupt => assert!(matches!(error, cellule_runtime::Error::Control(_))), + CanonicalFault::Substituted { .. } => assert!(matches!( + error, + cellule_runtime::Error::Control( + "receiver recovery input differs from canonical acquisition" + ) + )), + } + assert_eq!(fixture.current().await.value(), &before); + assert_eq!( + fixture + .native + .journal + .load_receiver_recovery_evidence(&fixture.accepted) + .await + .unwrap(), + evidence + ); + assert_eq!( + fixture + .native + .journal + .load_snapshot(scope()) + .await + .unwrap() + .head() + .reserved_restore_bytes(), + fixture.native.spec.cost.disk_bytes + ); + assert_eq!( + fixture.native.nodes[2].stats().active_cells(), + usize::from(ordinary_winner) + ); + if !matches!(fault, CanonicalFault::Missing) { + layout.store().delete(&path).await.unwrap(); + } + layout.store().create_strict(&path, original).await.unwrap(); + } + let activated = fixture.replay().await; + assert!( + activated.committed && activated.execution_error.is_none(), + "{activated:?}" + ); + assert!(matches!( + activated.outcome.outcome, + FleetOutcome::Activated(_) + )); + let retained = fixture + .native + .journal + .load_receiver_recovery_evidence(&fixture.accepted) + .await + .unwrap() + .unwrap(); + assert_eq!(retained.basis(), &basis); + assert_eq!(retained.restored(), canonical.materialized()); + if let Some(evidence) = evidence { + assert_eq!(retained, evidence); + } + fixture.finish().await; +} diff --git a/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/tests/evidence.rs b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/tests/evidence.rs new file mode 100644 index 00000000..fd1ba075 --- /dev/null +++ b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/tests/evidence.rs @@ -0,0 +1,175 @@ +use super::*; +use continuation::evidence_failure; + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn routed_original_basis_resumes_its_exact_claim_without_another_epoch() { + let fixture = FaultFixture::paused( + RecoveryWrite::Basis, + RecoveryWriteBoundary::AfterCommit, + true, + ) + .await + .release(false) + .await; + let basis = fixture + .native + .journal + .load_receiver_recovery_basis(&fixture.accepted) + .await + .unwrap() + .unwrap(); + let record = &fixture.native.records[&fixture.native.spec.target.cell_id()]; + let before = fixture.current().await; + assert_eq!(before.value(), basis.control()); + // A real canonical CAS models the interruption window after independently + // confirming the retained basis and before materialization/actor admission. + // No native acquisition or serving evidence is synthesized. + let claimed = record + .authority + .transition( + &before, + before.value().takeover(owner(2)).unwrap(), + Transition::Takeover, + ) + .await + .unwrap(); + assert_eq!(claimed.value().epoch, fixture.original.epoch + 1); + assert_eq!(fixture.native.nodes[2].stats().active_cells(), 0); + assert!( + record + .authority + .acquisition_record( + record.target.cell_id(), + record.incarnation, + claimed.value().epoch + ) + .await + .unwrap() + .is_none() + ); + let completion = fixture.replay().await; + assert!( + completion.committed && completion.execution_error.is_none(), + "{completion:?}" + ); + assert!(matches!( + completion.outcome.outcome, + FleetOutcome::Activated(_) + )); + assert_eq!(fixture.current().await.value().epoch, claimed.value().epoch); + assert_eq!( + fixture + .native + .journal + .load_receiver_recovery_basis(&fixture.accepted) + .await + .unwrap(), + Some(basis) + ); + fixture.finish().await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn lost_receiver_evidence_reply_keeps_idle_root_and_retained_proof_until_native_serving() { + evidence_failure(RecoveryWriteBoundary::AfterCommit, None, false).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn interrupted_receiver_evidence_write_reconstructs_only_from_canonical_acquisition() { + evidence_failure(RecoveryWriteBoundary::BeforeCommit, None, false).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn ordinary_acquisition_winner_retains_the_original_receiver_recovery() { + for boundary in [ + RecoveryWriteBoundary::BeforeCommit, + RecoveryWriteBoundary::AfterCommit, + ] { + evidence_failure(boundary, None, true).await; + } +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn interrupted_reconstructed_evidence_write_prevents_idle_reacquisition() { + for boundary in [ + RecoveryWriteBoundary::BeforeCommit, + RecoveryWriteBoundary::AfterCommit, + ] { + let fixture = FaultFixture::paused( + RecoveryWrite::Evidence, + RecoveryWriteBoundary::BeforeCommit, + true, + ) + .await + .release(false) + .await; + let idle = fixture.current().await.value().clone(); + assert_eq!(idle.state, ControlState::Idle); + let (entered, resume) = fixture.native.journal.pause_receiver_recovery_write( + RecoveryWrite::Evidence, + boundary, + true, + ); + let node = fixture.native.nodes[2].clone(); + let action = fixture.accepted.action().clone(); + let waiter = tokio::spawn(async move { + node.apply_fleet_action(action, clock().unwrap()) + .await + .unwrap() + }); + tokio::time::timeout(Duration::from_secs(5), entered) + .await + .unwrap() + .unwrap(); + assert_eq!(fixture.current().await.value(), &idle); + assert_eq!(fixture.native.nodes[2].stats().active_cells(), 0); + let retained = fixture + .native + .journal + .load_receiver_recovery_evidence(&fixture.accepted) + .await + .unwrap(); + assert_eq!( + retained.is_some(), + boundary == RecoveryWriteBoundary::AfterCommit + ); + resume.send(()).unwrap(); + let failed = waiter.await.unwrap(); + assert!(failed.committed && matches!(failed.outcome.outcome, FleetOutcome::Unknown)); + assert_original_error(&failed); + assert_eq!(fixture.current().await.value(), &idle); + assert_eq!( + fixture + .native + .journal + .load_snapshot(scope()) + .await + .unwrap() + .head() + .reserved_restore_bytes(), + fixture.native.spec.cost.disk_bytes + ); + let activated = fixture.replay().await; + assert!( + activated.committed && activated.execution_error.is_none(), + "{activated:?}" + ); + assert!(matches!( + activated.outcome.outcome, + FleetOutcome::Activated(_) + )); + assert_eq!(fixture.current().await.value().epoch, idle.epoch + 1); + if let Some(retained) = retained { + assert_eq!( + fixture + .native + .journal + .load_receiver_recovery_evidence(&fixture.accepted) + .await + .unwrap(), + Some(retained) + ); + } + fixture.finish().await; + } +} diff --git a/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/tests/history.rs b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/tests/history.rs new file mode 100644 index 00000000..d1e992f5 --- /dev/null +++ b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/tests/history.rs @@ -0,0 +1,68 @@ +use super::*; +use continuation::{CanonicalFault, evidence_failure}; + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn missing_or_corrupt_receiver_acquisition_history_cannot_reconstruct_evidence() { + for fault in [CanonicalFault::Missing, CanonicalFault::Corrupt] { + for ordinary in [false, true] { + evidence_failure( + RecoveryWriteBoundary::BeforeCommit, + Some(fault.clone()), + ordinary, + ) + .await; + } + } +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn retained_receiver_evidence_cannot_replace_missing_native_history() { + for ordinary in [false, true] { + evidence_failure( + RecoveryWriteBoundary::AfterCommit, + Some(CanonicalFault::Missing), + ordinary, + ) + .await; + } +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn valid_but_substituted_receiver_acquisition_cannot_authorize_continuation() { + // Obtain valid canonical bytes from a separate, fully joined recovery. + // Its scope/epoch decode correctly, but its exact original input differs. + let foreign = FaultFixture::paused( + RecoveryWrite::Evidence, + RecoveryWriteBoundary::AfterCommit, + false, + ) + .await + .release(false) + .await; + let record = &foreign.native.records[&foreign.native.spec.target.cell_id()]; + let path = record.authority.layout().acquisition_record_path( + record.target.cell_id().as_bytes(), + record.incarnation.as_bytes(), + foreign.original.epoch + 1, + ); + let (body, _) = record + .authority + .layout() + .store() + .get_with_etag(&path) + .await + .unwrap(); + let input = foreign.original.clone(); + foreign.finish().await; + for ordinary in [false, true] { + evidence_failure( + RecoveryWriteBoundary::BeforeCommit, + Some(CanonicalFault::Substituted { + body: body.clone(), + input: input.clone(), + }), + ordinary, + ) + .await; + } +} diff --git a/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/tests/mod.rs b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/tests/mod.rs new file mode 100644 index 00000000..615fe6ce --- /dev/null +++ b/crates/cellule-host/minion/scenario/successor_tests/routed/recovery_faults/tests/mod.rs @@ -0,0 +1,131 @@ +use super::*; +use fixture::FaultFixture; + +async fn assert_reserved_without_actor(native: &super::super::fixture::Fixture) { + // active_cells includes the held affine activation credit. The actual + // native writer inventory must remain empty before the journal confirms. + assert_eq!(native.nodes[2].stats().active_cells(), 1); + assert_eq!(native.nodes[2].stats().worker_jobs(), 0); + let page = native.nodes[2] + .runtime() + .fleet_cells_page(None, 128) + .await + .unwrap(); + assert!( + !page + .entries() + .iter() + .any(|row| matches!(row, CellInventoryEntry::Owned(_))) + ); +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn interrupted_receiver_basis_write_prevents_cas_and_replays_exact_input() { + for boundary in [ + RecoveryWriteBoundary::BeforeCommit, + RecoveryWriteBoundary::AfterCommit, + ] { + let fixture = FaultFixture::paused(RecoveryWrite::Basis, boundary, true).await; + assert_eq!(fixture.current().await, fixture.original); + assert_reserved_without_actor(&fixture.native).await; + let basis = fixture + .native + .journal + .load_receiver_recovery_basis(&fixture.accepted) + .await + .unwrap(); + assert_eq!( + basis.is_some(), + boundary == RecoveryWriteBoundary::AfterCommit + ); + if let Some(basis) = &basis { + assert_eq!(basis.control(), &fixture.original); + } + let fixture = fixture.release(false).await; + let failed = fixture.completion.as_ref().unwrap(); + assert!(failed.committed); + assert!(matches!(failed.outcome.outcome, FleetOutcome::Unknown)); + assert_original_error(failed); + assert_eq!(fixture.current().await.value(), &fixture.original); + let activated = fixture.replay().await; + assert!( + activated.committed && activated.execution_error.is_none(), + "{activated:?}" + ); + assert!(matches!( + activated.outcome.outcome, + FleetOutcome::Activated(_) + )); + if let Some(basis) = basis { + assert_eq!( + fixture + .native + .journal + .load_receiver_recovery_basis(&fixture.accepted) + .await + .unwrap(), + Some(basis) + ); + } + fixture.finish().await; + } +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn cancelled_receiver_waiter_keeps_the_original_basis_and_takeover_owned() { + for write in [RecoveryWrite::Basis, RecoveryWrite::Evidence] { + let fixture = FaultFixture::paused(write, RecoveryWriteBoundary::BeforeCommit, false).await; + let current = fixture.current().await; + assert_eq!( + current.epoch, + fixture.original.epoch + u64::from(write == RecoveryWrite::Evidence) + ); + assert_reserved_without_actor(&fixture.native).await; + let basis = fixture + .native + .journal + .load_receiver_recovery_basis(&fixture.accepted) + .await + .unwrap(); + assert_eq!(basis.is_some(), write == RecoveryWrite::Evidence); + assert!( + fixture + .native + .journal + .load_receiver_recovery_evidence(&fixture.accepted) + .await + .unwrap() + .is_none() + ); + let fixture = fixture.release(true).await; + let completion = fixture.replay().await; + assert!( + completion.committed && completion.execution_error.is_none(), + "{completion:?}" + ); + assert!(matches!( + completion.outcome.outcome, + FleetOutcome::Activated(_) + )); + assert_eq!( + fixture.current().await.value().epoch, + fixture.original.epoch + 1 + ); + if let Some(basis) = basis { + assert_eq!( + fixture + .native + .journal + .load_receiver_recovery_basis(&fixture.accepted) + .await + .unwrap(), + Some(basis) + ); + } + fixture.finish().await; + } +} + +mod continuation; +mod evidence; +mod history; diff --git a/crates/cellule-host/minion/scenario/successor_tests/routed/tests.rs b/crates/cellule-host/minion/scenario/successor_tests/routed/tests.rs new file mode 100644 index 00000000..5a42e568 --- /dev/null +++ b/crates/cellule-host/minion/scenario/successor_tests/routed/tests.rs @@ -0,0 +1,397 @@ +use super::*; +use fixture::Fixture; + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn driver_routes_activation_and_cleanup_after_receiver_boot_closure() { + route_idle_receiver(false).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn accepted_activation_without_claim_continues_after_receiver_boot_closure() { + route_idle_receiver(true).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn closed_receiver_recovering_claim_requires_canonical_recovery() { + refuse_claimed_receiver(ControlState::Recovering).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn closed_receiver_serving_claim_requires_canonical_recovery() { + refuse_claimed_receiver(ControlState::Serving).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn closed_recovering_receiver_recovers_through_canonical_takeover() { + route_receiver(false, Some(ControlState::Recovering)).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn closed_serving_receiver_recovers_through_canonical_takeover() { + route_receiver(false, Some(ControlState::Serving)).await; +} + +async fn refuse_claimed_receiver(state: ControlState) { + let fixture = Fixture::released().await; + fixture.claim_without_actor(state).await; + let record = &fixture.records[&fixture.spec.target.cell_id()]; + let claimed = record + .authority + .load(fixture.spec.target.cell_id()) + .await + .unwrap() + .unwrap(); + assert_eq!(claimed.value().state, state); + assert_eq!(claimed.value().owner.as_ref().unwrap().session, session(1)); + let observer = fixture.close_receiver().await; + let driver = fixture.driver( + SessionId::from_bytes([206; 16]), + observer, + fixture.fleet.clone(), + ); + for _ in 0..3 { + let report = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await + .unwrap(); + assert_eq!(report.activated, 0); + assert_eq!(report.retired, 0); + assert_eq!(report.failures.len(), 1); + assert_eq!(report.failures[0].attempt, fixture.spec.id); + let attempt = &report.snapshot.head().attempts()[0]; + assert_eq!(attempt.phase(), AttemptPhase::Activating); + assert_eq!(attempt.released(), Some(&fixture.released)); + assert!(!attempt.receiver_resources_settled()); + assert_eq!( + report.snapshot.head().reserved_restore_bytes(), + fixture.spec.cost.disk_bytes + ); + assert_eq!(fixture.nodes[2].stats().active_cells(), 0); + assert_eq!(fixture.nodes[2].stats().local_disk_reserved_bytes(), 0); + assert_eq!( + record + .authority + .load(fixture.spec.target.cell_id()) + .await + .unwrap() + .unwrap() + .value(), + claimed.value() + ); + } + let actions = fixture + .journal + .load_movement_actions(scope(), &fixture.released_attempt, MovementAction::Activate) + .await + .unwrap(); + let (accepted, result) = actions + .into_iter() + .find_map(|acceptance| match acceptance { + FleetActionAcceptance::Existing { accepted, result } + if accepted.action().receiver_endpoint() == Some((node_id(2), session(2))) => + { + Some((accepted, result)) + } + _ => None, + }) + .expect("checked routed refusal missing"); + assert!(matches!(result.unwrap().outcome, FleetOutcome::Unknown)); + assert!( + fixture + .journal + .load_acquisition_basis(&accepted) + .await + .unwrap() + .is_none() + ); + + // Receiver-credit cleanup is not permission to erase an unresolved + // ownership claim, even after the original process has joined shutdown. + let snapshot = fixture.journal.load_snapshot(scope()).await.unwrap(); + fixture + .journal + .compare_exchange( + &snapshot, + snapshot.head().controller().unwrap().epoch, + clock().unwrap(), + &JournalTransition::Attempt { + id: fixture.spec.id, + event: AttemptEvent::BeginCancel, + }, + ) + .await + .unwrap(); + let report = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await + .unwrap(); + let attempt = &report.snapshot.head().attempts()[0]; + assert_eq!(attempt.phase(), AttemptPhase::CleaningReceiver); + assert!(!attempt.receiver_resources_settled()); + assert_eq!(report.retired, 0); + assert_eq!( + report.snapshot.head().reserved_restore_bytes(), + fixture.spec.cost.disk_bytes + ); + fixture.shutdown().await; +} + +async fn route_idle_receiver(accepted_before_shutdown: bool) { + route_receiver(accepted_before_shutdown, None).await; +} + +async fn route_receiver(accepted_before_shutdown: bool, failed_state: Option) { + let fixture = Fixture::released_with_recovery(failed_state.is_some()).await; + if let Some(state) = failed_state { + fixture.claim_without_actor(state).await; + } + if accepted_before_shutdown { + fixture.accept_original_activation().await; + } + let observer = fixture.close_receiver().await; + let journal = &fixture.journal; + let nodes = &fixture.nodes; + let fleet = &fixture.fleet; + let profile = fixture.profile; + let spec = &fixture.spec; + let released_attempt = &fixture.released_attempt; + let released = &fixture.released; + let acknowledged = &fixture.acknowledged; + let controller = SessionId::from_bytes([206; 16]); + let new_controller_observer = Arc::new(ClosedBootObserver { + fleet: fleet.clone(), + request: observer.request.clone(), + processes: observer.processes.clone(), + }); + let loss_transport = Arc::new(LoseRoutedActivationReply { + inner: fleet.clone(), + lost: std::sync::atomic::AtomicBool::new(false), + }); + let driver = FleetReconciler::new( + scope(), + controller, + profile, + journal.clone(), + observer, + loss_transport.clone(), + ) + .unwrap(); + let mut report = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await + .unwrap(); + for _ in 0..8 { + if loss_transport + .lost + .load(std::sync::atomic::Ordering::Acquire) + { + break; + } + report = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await + .unwrap(); + } + assert!( + loss_transport + .lost + .load(std::sync::atomic::Ordering::Acquire) + ); + assert!( + report + .failures + .iter() + .any(|failure| failure.attempt == spec.id) + ); + let accepted_activation = journal + .load_movement_actions(scope(), released_attempt, MovementAction::Activate) + .await + .unwrap() + .into_iter() + .find_map(|acceptance| match acceptance { + FleetActionAcceptance::New(accepted) + | FleetActionAcceptance::Existing { accepted, .. } + if accepted.action().receiver_endpoint() == Some((node_id(2), session(2))) => + { + Some(accepted) + } + _ => None, + }) + .expect("routed Activate acceptance missing after reply loss"); + assert_eq!( + accepted_activation.action().receiver_endpoint(), + Some((node_id(2), session(2))) + ); + let duplicate = fleet + .dispatch( + accepted_activation.action(), + Instant::now() + Duration::from_secs(5), + ) + .await + .unwrap(); + assert!(duplicate.committed); + assert!(matches!( + &duplicate.outcome.outcome, + FleetOutcome::Activated(_) + )); + assert_eq!(nodes[2].stats().active_cells(), 1); + if failed_state.is_some() { + assert!( + journal + .load_acquisition_basis(&accepted_activation) + .await + .unwrap() + .is_none() + ); + let evidence = journal + .load_receiver_recovery_evidence(&accepted_activation) + .await + .unwrap() + .unwrap(); + assert_eq!(evidence.basis().control().state, failed_state.unwrap()); + assert_eq!( + evidence.basis().control().owner.as_ref().unwrap().session, + session(1) + ); + assert_eq!(evidence.basis().control().epoch, released.epoch + 1); + assert_eq!(evidence.restored().epoch, released.epoch + 2); + assert_eq!(evidence.restored().root.as_ref(), Some(&released.root)); + assert_eq!( + journal + .record_receiver_recovery_evidence(&evidence) + .await + .unwrap(), + evidence + ); + assert_eq!( + journal + .record_receiver_recovery_basis(evidence.basis()) + .await + .unwrap(), + *evidence.basis() + ); + } + + // Let the original lease expire. The new claimant must adopt the exact + // route and committed result already retained by the application journal. + tokio::time::sleep(Duration::from_millis( + u64::try_from(profile.controller_lease_ms).unwrap() + 50, + )) + .await; + let reopened_journal = fixture.reopen_journal().await; + let driver = FleetReconciler::new( + scope(), + SessionId::from_bytes([207; 16]), + profile, + reopened_journal.clone(), + new_controller_observer, + fleet.clone(), + ) + .unwrap(); + let mut activated = None; + for _ in 0..12 { + if report.snapshot.head().attempts().is_empty() { + break; + } + report = driver + .reconcile_once(clock, Instant::now() + Duration::from_secs(5)) + .await + .unwrap(); + activated = activated.or_else(|| { + report + .snapshot + .head() + .attempts() + .first()? + .activated() + .cloned() + }); + } + assert!(report.snapshot.head().attempts().is_empty()); + assert_eq!(report.retired, 1); + assert_eq!(report.snapshot.head().reserved_restore_bytes(), 0); + let activated = activated.expect("routed replacement activation was not observed"); + assert_eq!( + (activated.node, activated.session), + (node_id(2), session(2)) + ); + assert_eq!(activated.position.root, released.root); + assert_eq!( + activated.position.epoch, + released.epoch + 1 + u64::from(failed_state.is_some()) + ); + if failed_state.is_some() { + let original = journal + .load_receiver_recovery_evidence(&accepted_activation) + .await + .unwrap() + .unwrap(); + assert_eq!( + reopened_journal + .load_receiver_recovery_evidence(&accepted_activation) + .await + .unwrap(), + Some(original) + ); + } + + for effect in [MovementAction::Activate, MovementAction::Cancel] { + let actions = journal + .load_movement_actions(scope(), released_attempt, effect) + .await + .unwrap(); + let accepted = actions.into_iter().find_map(|acceptance| match acceptance { + FleetActionAcceptance::New(accepted) + | FleetActionAcceptance::Existing { accepted, .. } + if accepted.action().receiver_endpoint() == Some((node_id(2), session(2))) => + { + Some(accepted) + } + _ => None, + }); + let accepted = accepted.expect("routed receiver action acceptance missing"); + assert_eq!( + accepted.action().receiver_endpoint(), + Some((node_id(2), session(2))) + ); + let route = accepted.action().receiver_route().unwrap(); + assert_eq!( + route.latest_handoff().unwrap().previous(), + (node_id(1), session(1)) + ); + } + + let record = &fixture.records[&spec.target.cell_id()]; + let serving = record + .authority + .load(spec.target.cell_id()) + .await + .unwrap() + .unwrap(); + let successor = nodes[2] + .runtime() + .local_handle(record.catalog.clone(), &serving) + .await + .unwrap() + .unwrap(); + let original = &acknowledged[&spec.target.cell_id()]; + assert_eq!( + successor + .resolve(original.identity, original.digest, clock().unwrap(), 64) + .await + .unwrap(), + Resolution::Committed(original.outcome.clone()) + ); + let bytes = successor + .query(64, 64, |connection| { + let value: i64 = + connection.query_row("SELECT value FROM counter", [], |row| row.get(0))?; + Ok(value.to_be_bytes().to_vec()) + }) + .await + .unwrap(); + assert_eq!(bytes, original.value.to_be_bytes()); + + fixture.shutdown().await; + reopened_journal.close().await.unwrap(); +} diff --git a/crates/cellule-host/minion/scenario/tests.rs b/crates/cellule-host/minion/scenario/tests.rs index 74865e84..5d087b47 100644 --- a/crates/cellule-host/minion/scenario/tests.rs +++ b/crates/cellule-host/minion/scenario/tests.rs @@ -28,3 +28,19 @@ async fn new_controller_adopts_lost_releases_after_real_expiry_and_joins_receive assert_eq!(summary.boot_retirements, 3); assert_eq!(summary.receiver_nodes, 2); } + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn maintenance_moves_every_cell_then_settles_roles_and_withdraws_the_node() { + let summary = super::maintenance().await.unwrap(); + assert_eq!(summary.released, 12); + assert_eq!(summary.activated, 12); + assert_eq!(summary.retired, 12); + assert_eq!(summary.receipt_checks, 12); + assert_eq!(summary.final_counts[0], 0); + assert_eq!(summary.final_counts.iter().sum::(), 12); + assert_eq!(summary.receiver_nodes, 2); + assert!(summary.maintenance_completed); + assert!(summary.maintenance_boot_withdrawn); + assert_eq!(summary.joined_nodes, 3); + assert_eq!(summary.boot_retirements, 3); +} diff --git a/crates/cellule-host/src/fleet/actions/mod.rs b/crates/cellule-host/src/fleet/actions/mod.rs index 35f51ff1..944d8760 100644 --- a/crates/cellule-host/src/fleet/actions/mod.rs +++ b/crates/cellule-host/src/fleet/actions/mod.rs @@ -11,7 +11,7 @@ use cellule_runtime::identity::{Digest, NodeId, SessionId}; use tokio::{sync::watch, task::JoinHandle}; use super::snapshot::{FleetNodeSnapshot, FleetSnapshotRequest, SnapshotOwners}; -use super::{FleetActionAcceptance, FleetActionJournal, FleetCellProvider}; +use super::{FleetActionAcceptance, FleetActionJournal, FleetCellProvider, FleetRoleSettlement}; mod work; pub use work::{ @@ -29,8 +29,9 @@ pub struct FleetActionCompletion { pub committed: bool, /// Original runtime error, independently of result publication failure. pub execution_error: Option>, - /// Original result-publication error. Retry publishes the same retained - /// evidence and never executes the source release again. + /// Original result-publication error. Native effects retry the same retained + /// evidence without reexecution. A failed retry of joined read-only role + /// evidence retires that local proof; a fresh full capture must replace it. pub journal_error: Option>, } @@ -39,6 +40,15 @@ pub(super) struct ActionResult { pub(super) error: Option, } +/// Result of durably accepting a host Finalize action outside the ordinary +/// finite action bank. +pub(crate) enum FinalizeActionAdmission { + /// The local terminal owner must run or resume the canonical host drain. + Accepted(AcceptedFleetAction), + /// The exact action already has a durable non-Unknown result. + Completed(Arc), +} + impl ActionResult { pub(super) fn checked(outcome: FleetOutcome) -> Self { Self { @@ -68,6 +78,7 @@ enum JobRequest { Effect { action: FleetAction, now_ms: i64, + settlement: Option>, }, Inspection(FleetInspectionRequest), Snapshot { @@ -105,9 +116,16 @@ impl JobRequest { } fn validate_replay(&self, other: &Self) -> Result<(), OperationError> { match (self, other) { - (Self::Effect { action, .. }, Self::Effect { action: replay, .. }) => { - action.validate_replay(replay) - } + ( + Self::Effect { + action, settlement, .. + }, + Self::Effect { + action: replay, + settlement: replay_settlement, + .. + }, + ) if settlement == replay_settlement => action.validate_replay(replay), (Self::Inspection(original), Self::Inspection(replay)) if original == replay => Ok(()), ( Self::Snapshot { @@ -239,6 +257,13 @@ impl FleetActionExecutor { }) } + pub(crate) fn reserve_finalize_retention( + &self, + ) -> cellule_runtime::Result { + self.runtime + .try_reserve_node_bytes(3 * MAX_RECORD_BYTES as usize) + } + pub(crate) async fn apply( self: &Arc, action: FleetAction, @@ -266,7 +291,13 @@ impl FleetActionExecutor { "fleet action time regressed", )))); } - let completion = self.submit(JobRequest::Effect { action, now_ms }).await?; + let completion = self + .submit(JobRequest::Effect { + action, + now_ms, + settlement: None, + }) + .await?; match completion.as_ref() { JobCompletion::Effect(result) => Ok(Arc::clone(result)), _ => Err(Arc::new(Error::Control( @@ -275,6 +306,163 @@ impl FleetActionExecutor { } } + pub(crate) async fn apply_role_settlement( + self: &Arc, + action: FleetAction, + settlement: FleetRoleSettlement, + now_ms: i64, + ) -> EffectCompletion { + if action.scope() != self.scope + || !matches!( + action.kind(), + FleetActionKind::Maintenance { + action: MaintenanceAction::SettleRoles, + .. + } + ) + { + return Err(Arc::new(Error::Fenced)); + } + action + .validate_endpoint(self.node, self.session) + .map_err(operation) + .map_err(Arc::new)?; + settlement.validate_for(&action).map_err(Arc::new)?; + if now_ms < action.issued_at_ms() { + return Err(Arc::new(operation(OperationError::Invalid( + "fleet action time regressed", + )))); + } + let completion = self + .submit(JobRequest::Effect { + action, + now_ms, + settlement: Some(Arc::new(settlement)), + }) + .await?; + match completion.as_ref() { + JobCompletion::Effect(result) => Ok(Arc::clone(result)), + _ => Err(Arc::new(Error::Control( + "fleet role settlement completion kind mismatch", + ))), + } + } + + /// Accepts the terminal action without adding it to the bank that the + /// canonical node drain joins. Closing admission after acceptance makes + /// every previously admitted sibling part of that drain's join set. + pub(crate) async fn accept_finalize( + &self, + action: &FleetAction, + now_ms: i64, + ) -> Result> { + let maintenance = match action.kind() { + FleetActionKind::Maintenance { + action: MaintenanceAction::Finalize, + operation, + } => operation, + _ => { + return Err(Arc::new(Error::Control( + "host finalization requires a Finalize maintenance action", + ))); + } + }; + let evidence = maintenance.drain_evidence().ok_or_else(|| { + Arc::new(operation(OperationError::Invalid( + "fleet Finalize lacks committed drain evidence", + ))) + })?; + if action.scope() != self.scope + || maintenance.node() != self.node + || maintenance.session() != self.session + || maintenance.phase() != cellule_runtime::fleet::operations::MaintenancePhase::Closing + || evidence.node != self.node + || evidence.session != self.session + || evidence.remaining_cells != 0 + || evidence.unresolved_attempts != 0 + || !evidence.relocated + || !evidence.readers_settled + || !evidence.followers_settled + { + return Err(Arc::new(operation(OperationError::Invalid( + "fleet Finalize lacks exact ready-to-close evidence", + )))); + } + action + .validate_endpoint(self.node, self.session) + .map_err(operation) + .map_err(Arc::new)?; + if now_ms < action.issued_at_ms() { + return Err(Arc::new(operation(OperationError::Invalid( + "fleet Finalize time regressed", + )))); + } + + let acceptance = self + .journal + .accept_action(action, self.node, self.session, now_ms) + .await + .map_err(journal_error) + .map_err(Arc::new)?; + let accepted = match acceptance { + FleetActionAcceptance::New(accepted) => { + accepted + .validate_replay(action, self.node, self.session) + .map_err(operation) + .map_err(Arc::new)?; + if accepted.action() != action || accepted.accepted_at_ms() != now_ms { + return Err(Arc::new(operation(OperationError::Conflict))); + } + accepted + } + FleetActionAcceptance::Existing { accepted, result } => { + accepted + .validate_replay(action, self.node, self.session) + .map_err(operation) + .map_err(Arc::new)?; + if let Some(result) = result { + accepted + .validate_result(&result) + .map_err(operation) + .map_err(Arc::new)?; + if !matches!(&result.outcome, FleetOutcome::Unknown) { + return Ok(FinalizeActionAdmission::Completed(Arc::new( + FleetActionCompletion { + accepted, + outcome: *result, + committed: true, + execution_error: None, + journal_error: None, + }, + ))); + } + } + accepted + } + }; + self.close_admission().map_err(Arc::new)?; + Ok(FinalizeActionAdmission::Accepted(accepted)) + } + + fn close_admission(&self) -> cellule_runtime::Result<()> { + let mut bank = self + .bank + .lock() + .map_err(|_| Error::Control("fleet action bank poisoned"))?; + bank.draining = true; + Ok(()) + } + + /// Publishes terminal evidence retained by the host Finalize owner. + pub(crate) async fn publish_terminal( + &self, + accepted: AcceptedFleetAction, + outcome: FleetActionOutcome, + execution_error: Option>, + ) -> FleetActionCompletion { + self.publish(accepted, outcome, execution_error).await + } + pub(crate) async fn observe( self: &Arc, request: FleetInspectionRequest, @@ -377,8 +565,12 @@ impl FleetActionExecutor { let issued = request.clone(); let task = tokio::spawn(async move { let result = match issued { - JobRequest::Effect { action, now_ms } => executor - .execute(action, now_ms) + JobRequest::Effect { + action, + now_ms, + settlement, + } => executor + .execute(action, now_ms, settlement) .await .map(JobCompletion::Effect), JobRequest::Inspection(request) => executor @@ -430,7 +622,12 @@ impl FleetActionExecutor { } } - async fn execute(&self, action: FleetAction, now_ms: i64) -> EffectCompletion { + async fn execute( + &self, + action: FleetAction, + now_ms: i64, + settlement: Option>, + ) -> EffectCompletion { let acceptance = self .journal .accept_action(&action, self.node, self.session, now_ms) @@ -443,8 +640,21 @@ impl FleetActionExecutor { .validate_replay(&action, self.node, self.session) .map_err(operation) .map_err(Arc::new)?; + let refresh_roles = matches!( + accepted.action().kind(), + FleetActionKind::Maintenance { + action: MaintenanceAction::SettleRoles, + .. + } + ) && result.as_ref().is_some_and(|result| { + matches!( + &result.outcome, + FleetOutcome::RolesSettled { .. } | FleetOutcome::RolesSettledAt { .. } + ) + }); if let Some(result) = result && !matches!(result.outcome, FleetOutcome::Unknown) + && !(refresh_roles && settlement.is_some()) { accepted .validate_result(&result) @@ -458,7 +668,12 @@ impl FleetActionExecutor { journal_error: None, })); } - let result = self.inspect_accepted(&accepted).await; + let result = if let Some(settlement) = settlement { + self.perform_action(&accepted, &action, Some(settlement)) + .await + } else { + self.inspect_accepted(&accepted).await + }; (accepted, result) } FleetActionAcceptance::New(accepted) => { @@ -469,7 +684,7 @@ impl FleetActionExecutor { if accepted.action() != &action || accepted.accepted_at_ms() != now_ms { return Err(Arc::new(operation(OperationError::Conflict))); } - let result = self.perform_action(&accepted).await; + let result = self.perform_action(&accepted, &action, settlement).await; (accepted, result) } }; @@ -640,12 +855,30 @@ impl FleetActionExecutor { && let JobCompletion::Effect(completion) = completion.as_ref() { if !completion.committed { - // Retry the retained proof, never the canonical effect. Keep - // this receipt if publication or resource settlement still fails. - self.journal + // Retry the retained proof, never the canonical effect. Native + // effect receipts stay retained across every publication failure. + if let Err(source) = self + .journal .publish_action_result(&completion.accepted, &completion.outcome) .await - .map_err(journal_error)?; + { + if matches!(&job.request, JobRequest::Effect { + action, settlement: Some(_), .. + } if matches!(action.kind(), FleetActionKind::Maintenance { + action: MaintenanceAction::SettleRoles, .. + })) { + // SettleRoles only validates a captured observation; it + // starts no physical effect. A joined proof cannot be + // restamped after its journal barrier changes. Return the + // publication error and retire this local read so the + // next pass can collect fresh opaque evidence. Its durable + // acceptance remains unresolved until that proof commits. + // Movement/native-effect receipts keep their original + // evidence and retained owner on every publication error. + self.remove_job(job)?; + } + return Err(journal_error(source)); + } } self.retire_receiver_receipt(completion).await?; } @@ -670,7 +903,7 @@ impl ActionBank { } } -pub(super) fn wall_time_ms() -> cellule_runtime::Result { +pub(crate) fn wall_time_ms() -> cellule_runtime::Result { let elapsed = std::time::SystemTime::now() .duration_since(std::time::UNIX_EPOCH) .map_err(|source| Error::Facility { diff --git a/crates/cellule-host/src/fleet/cells.rs b/crates/cellule-host/src/fleet/cells.rs index 91fd5c0e..9f070b4a 100644 --- a/crates/cellule-host/src/fleet/cells.rs +++ b/crates/cellule-host/src/fleet/cells.rs @@ -1,7 +1,7 @@ use super::FleetAdapterFuture; use cellule_runtime::cell::catalog::CatalogProof; -use cellule_runtime::control::{Owner, authority::CellAuthority}; -use cellule_runtime::fleet::operations::MoveAttemptSpec; +use cellule_runtime::control::{Control, Owner, authority::CellAuthority}; +use cellule_runtime::fleet::operations::{AcceptedFleetAction, MoveAttemptSpec}; use cellule_runtime::ltx::CellReplica; use std::path::PathBuf; @@ -52,4 +52,16 @@ pub trait FleetCellProvider: Send + Sync + 'static { &'a self, spec: &'a MoveAttemptSpec, ) -> FleetAdapterFuture<'a, FleetRecoveryInputs>; + + /// Read-only lookup of canonical recovery prerequisites for the exact + /// failed receiver control under a committed routed activation. None means + /// no configured/proven recovery, never permission to discard its claim. + /// This lookup must not fence, recover a log, change authority or admit work. + fn receiver_recovery_inputs<'a>( + &'a self, + _accepted: &'a AcceptedFleetAction, + _control: &'a Control, + ) -> FleetAdapterFuture<'a, Option> { + Box::pin(async { Ok(None) }) + } } diff --git a/crates/cellule-host/src/fleet/controller.rs b/crates/cellule-host/src/fleet/controller.rs index 4731b866..9010d185 100644 --- a/crates/cellule-host/src/fleet/controller.rs +++ b/crates/cellule-host/src/fleet/controller.rs @@ -69,6 +69,12 @@ pub trait FleetJournal: FleetActionJournal + FleetEnrollmentJournal { /// loss of history after session adoption cannot become another first capture. /// For Retire, publish the exact progress page with permit retirement; a /// failed CAS cannot expose committed history or release either budget. + /// For ReadyToClose, require the exact operation/node/session's accepted + /// SettleRoles action and its committed RolesSettledAt result for the exact + /// current head revision and RegistryVersion in this same transaction. + /// Validate its original operation identity and result time; an accepted + /// action, stale head/registry, Unknown result, or caller-supplied drain + /// flags alone cannot enter Closing. /// Finalization must compare the observed registry version again here. /// ResolveUnaccepted must prove no accepted record exists for its exact /// effect/attempt/endpoint in this same transaction before changing the head. diff --git a/crates/cellule-host/src/fleet/journal.rs b/crates/cellule-host/src/fleet/journal.rs index 763bbbfc..e40d2b81 100644 --- a/crates/cellule-host/src/fleet/journal.rs +++ b/crates/cellule-host/src/fleet/journal.rs @@ -2,7 +2,8 @@ use std::{future::Future, pin::Pin}; use cellule_runtime::fleet::operations::{ AcceptedFleetAction, AcquisitionBasis, AttemptId, FleetAction, FleetActionOutcome, - FleetInspectionRequest, FleetScope, MovementAction, RecoveryBasis, RecoveryEvidence, + FleetInspectionRequest, FleetScope, MoveAttempt, MovementAction, ReceiverRecoveryBasis, + ReceiverRecoveryEvidence, RecoveryBasis, RecoveryEvidence, }; use cellule_runtime::identity::{NodeId, SessionId}; @@ -29,9 +30,11 @@ pub enum FleetActionAcceptance { /// /// First acceptance must linearize its fresh head, controller epoch, permit, /// intent revision, endpoint and deadline checks with record publication. -/// Use `AcceptedFleetAction::new` inside that transaction. An unconditional -/// write following a separate head read is insufficient. Compare complete -/// execution inputs with `validate_replay`, not only the stable action key. +/// Use `AcceptedFleetAction::new_with_registry` inside that transaction. This +/// also binds receiver continuations to the current registry revision. An +/// unconditional write following a separate head read is insufficient. +/// Compare complete execution inputs with `validate_replay`, not only the +/// stable action key. /// Applications authenticate the caller before invoking the node. pub trait FleetActionJournal: Send + Sync + 'static { /// Checks a native page request against its full current head/registry and @@ -62,6 +65,26 @@ pub trait FleetActionJournal: Send + Sync + 'static { now_ms: i64, ) -> FleetAdapterFuture<'a, FleetActionAcceptance>; + /// Atomically first-accepts a routed action after checking its final hop + /// against this fresh typed closed-boot proof and the same current snapshot. + /// The default refuses; adapters must implement the proof comparison before + /// any remote effect can start. Replays of an already accepted action use + /// `accept_action` and retain the original acceptance. + fn accept_closed_receiver_action<'a>( + &'a self, + _action: &'a FleetAction, + _node: NodeId, + _session: SessionId, + _now_ms: i64, + _closure: &'a super::FleetFailedBootClosure, + ) -> FleetAdapterFuture<'a, FleetActionAcceptance> { + Box::pin(async { + Err(Box::new(std::io::Error::other( + "journal does not support closed receiver continuations", + )) as Box) + }) + } + /// Publishes a checked result bound to the original acceptance. /// /// Identical terminal results are idempotent; incompatible terminal results @@ -69,6 +92,8 @@ pub trait FleetActionJournal: Send + Sync + 'static { /// evidence. A successful return means the result is durably reachable, /// including after backend/client reconstruction. Retiring a fleet permit /// remains a separate controller transition requiring all cleanup evidence. + /// `RolesSettledAt` must compare its head revision and registry version with + /// the exact current snapshot in this same publication transaction. fn publish_action_result<'a>( &'a self, accepted: &'a AcceptedFleetAction, @@ -87,6 +112,32 @@ pub trait FleetActionJournal: Send + Sync + 'static { session: SessionId, ) -> FleetAdapterFuture<'a, Option>; + /// Loads every accepted receiver action for this immutable attempt/effect. + /// Implementations return a complete bounded set so a new controller can + /// find the latest durable route after an ambiguous response. The default + /// preserves older adapters by checking only the originally preferred + /// endpoint; production continuation adapters must override it. + fn load_movement_actions<'a>( + &'a self, + scope: FleetScope, + attempt: &'a MoveAttempt, + effect: MovementAction, + ) -> FleetAdapterFuture<'a, Vec> { + Box::pin(async move { + let spec = attempt.spec(); + let (node, session) = if effect.is_source_release() { + (spec.source_node, spec.source) + } else { + (spec.destination_node, spec.destination) + }; + Ok(self + .load_movement_action(scope, spec.id, effect, node, session) + .await? + .into_iter() + .collect()) + }) + } + /// Durably records the exact checked input before receiver acquisition. /// /// Bind it to the original acceptance atomically. Identical accepted/control @@ -133,4 +184,41 @@ pub trait FleetActionJournal: Send + Sync + 'static { &'a self, accepted: &'a AcceptedFleetAction, ) -> FleetAdapterFuture<'a, Option>; + + /// Confirms immutable failed-receiver input before the canonical takeover. + /// Bind atomically to the original routed activation; changed controls + /// conflict and identical writes return the original observation time. + fn record_receiver_recovery_basis<'a>( + &'a self, + _basis: &'a ReceiverRecoveryBasis, + ) -> FleetAdapterFuture<'a, ReceiverRecoveryBasis> { + Box::pin(async { + Err(std::io::Error::other("receiver recovery journal is not configured").into()) + }) + } + /// Returns retained input without inferring it from successor state. + fn load_receiver_recovery_basis<'a>( + &'a self, + _accepted: &'a AcceptedFleetAction, + ) -> FleetAdapterFuture<'a, Option> { + Box::pin(async { Ok(None) }) + } + /// Confirms the canonical recovered position before actor admission. + /// Require the exact retained basis; incompatible/ambiguous writes cannot + /// authorize activation. Preserve this record through later publication. + fn record_receiver_recovery_evidence<'a>( + &'a self, + _evidence: &'a ReceiverRecoveryEvidence, + ) -> FleetAdapterFuture<'a, ReceiverRecoveryEvidence> { + Box::pin(async { + Err(std::io::Error::other("receiver recovery journal is not configured").into()) + }) + } + /// Returns immutable pre-admission evidence for the accepted route. + fn load_receiver_recovery_evidence<'a>( + &'a self, + _accepted: &'a AcceptedFleetAction, + ) -> FleetAdapterFuture<'a, Option> { + Box::pin(async { Ok(None) }) + } } diff --git a/crates/cellule-host/src/fleet/maintenance.rs b/crates/cellule-host/src/fleet/maintenance.rs index 58f7cf32..4a022c4c 100644 --- a/crates/cellule-host/src/fleet/maintenance.rs +++ b/crates/cellule-host/src/fleet/maintenance.rs @@ -3,7 +3,7 @@ use super::actions::{ActionResult, FleetActionExecutor}; use cellule_runtime::Error; use cellule_runtime::fleet::operations::{ - AcceptedFleetAction, FleetActionKind, FleetOutcome, MaintenanceAction, + AcceptedFleetAction, FleetAction, FleetActionKind, FleetOutcome, MaintenanceAction, }; use cellule_runtime::node::NodeMode; @@ -11,22 +11,30 @@ impl FleetActionExecutor { pub(super) async fn perform_action( &self, accepted: &AcceptedFleetAction, + request: &FleetAction, + settlement: Option>, ) -> cellule_runtime::Result { match accepted.action().kind() { FleetActionKind::Movement { .. } => self.perform_movement(accepted).await, - FleetActionKind::Maintenance { .. } => self.perform_maintenance(accepted), + FleetActionKind::Maintenance { .. } => { + self.perform_maintenance(accepted, request, settlement.as_deref()) + } } } pub(super) fn perform_maintenance( &self, accepted: &AcceptedFleetAction, + request: &FleetAction, + settlement: Option<&crate::fleet::FleetRoleSettlement>, ) -> cellule_runtime::Result { - let FleetActionKind::Maintenance { - action: MaintenanceAction::Cordon, - operation, - } = accepted.action().kind() - else { + // Acceptance retains the original endpoint/key and execution identity. + // A read-only role refresh covers the current request's journal barrier; + // publication atomically checks that barrier before replacing its receipt. + accepted + .validate_replay(request, self.node, self.session) + .map_err(super::actions::operation)?; + let FleetActionKind::Maintenance { action, operation } = request.kind() else { return Err(Error::Control("unsupported fleet maintenance effect")); }; if accepted.action().scope() != self.scope @@ -35,14 +43,30 @@ impl FleetActionExecutor { { return Err(Error::Fenced); } - // The journal accepted the retained Draining intent. This closes the - // same gate used by writers, readers and followers, without releasing - // any existing responsibility or taking the node shutdown lane. - let gate = self.runtime.node_admission(); - gate.begin_drain()?; - if gate.mode()? != NodeMode::Draining { - return Err(Error::CellDraining); + match action { + MaintenanceAction::Cordon => { + // The journal accepted the retained Draining intent. This closes the + // same gate used by writers, readers and followers, without releasing + // any existing responsibility or taking the node shutdown lane. + let gate = self.runtime.node_admission(); + gate.begin_drain()?; + if gate.mode()? != NodeMode::Draining { + return Err(Error::CellDraining); + } + Ok(ActionResult::checked(FleetOutcome::Cordoned)) + } + MaintenanceAction::SettleRoles => { + let proof = settlement.ok_or(Error::Control( + "role settlement requires fresh complete host evidence", + ))?; + proof.validate_for(request)?; + Ok(ActionResult::checked(FleetOutcome::RolesSettledAt { + inventory: proof.inventory(), + head_revision: proof.head_revision(), + registry: proof.registry(), + })) + } + _ => Err(Error::Control("unsupported fleet maintenance effect")), } - Ok(ActionResult::checked(FleetOutcome::Cordoned)) } } diff --git a/crates/cellule-host/src/fleet/mod.rs b/crates/cellule-host/src/fleet/mod.rs index c0cab4d9..1b59b231 100644 --- a/crates/cellule-host/src/fleet/mod.rs +++ b/crates/cellule-host/src/fleet/mod.rs @@ -67,7 +67,7 @@ pub use inventory::{FleetNodeInventory, FleetNodeInventoryRecheck, FleetNodeInve pub use journal::{FleetActionAcceptance, FleetActionJournal, FleetAdapterFuture}; pub use reconciler::{ FleetAttemptFailure, FleetObservation, FleetObserver, FleetOwnedCell, FleetReconcileReport, - FleetReconciler, FleetTransport, + FleetReconciler, FleetRoleSettlement, FleetTransport, }; pub use recovered::{ FleetRecoveredFollowerClosure, FleetRecoveredFollowerMember, FleetRecoveredFollowerPublication, @@ -80,8 +80,8 @@ pub use snapshot::{ FleetSnapshotSubject, FleetSnapshotTransport, }; -pub(crate) use actions::FleetActionExecutor; pub(crate) use actions::operation; +pub(crate) use actions::{FinalizeActionAdmission, FleetActionExecutor, wall_time_ms}; /// Stable name of the node-owned finite fleet-action executor. pub const FLEET_ACTION_COMPONENT: &str = "fleet-actions"; diff --git a/crates/cellule-host/src/fleet/movement/activation.rs b/crates/cellule-host/src/fleet/movement/activation.rs index 251df0a4..623a90c6 100644 --- a/crates/cellule-host/src/fleet/movement/activation.rs +++ b/crates/cellule-host/src/fleet/movement/activation.rs @@ -6,19 +6,17 @@ impl FleetActionExecutor { accepted: &AcceptedFleetAction, attempt: &MoveAttempt, ) -> cellule_runtime::Result { + let closed_receiver_route = accepted.action().receiver_route().is_some(); let prepared = match self.runtime.prepared_receiver(attempt.spec().id)? { Some(_) => { let prepared = self.prepared(attempt)?; match prepared.state()? { ReceiverState::Prepared => Some(prepared), - ReceiverState::Cancelled - if self.confirmed_credit_settlement(attempt).await? => - { - None - } + ReceiverState::Cancelled => None, _ => return Ok(ActionResult::checked(FleetOutcome::Unknown)), } } + None if closed_receiver_route => None, None if self.confirmed_credit_settlement(attempt).await? => None, None => { return Err(Error::Peer( @@ -33,6 +31,28 @@ impl FleetActionExecutor { .await? .ok_or(Error::Control("fleet activation authority is absent"))?; self.check_contract(attempt, &inputs, &observed)?; + if closed_receiver_route + && prepared.is_none() + && !(observed.value().state == ControlState::Idle + && observed.value().owner.is_none() + && observed.value().root.is_some()) + && !(observed.value().state == ControlState::Serving + && observed + .value() + .owner + .as_ref() + .is_some_and(|owner| owner.session == self.session)) + { + if matches!( + observed.value().state, + ControlState::Serving | ControlState::Recovering + ) { + return self + .recover_receiver(accepted, attempt, &inputs, observed) + .await; + } + return Ok(ActionResult::checked(FleetOutcome::Unknown)); + } if prepared.is_none() && observed.value().state == ControlState::Serving && observed @@ -41,6 +61,18 @@ impl FleetActionExecutor { .as_ref() .is_some_and(|owner| owner.session == self.session) { + if closed_receiver_route + && self + .journal + .load_receiver_recovery_basis(accepted) + .await + .map_err(journal_error)? + .is_some() + { + return self + .receiver_recovered_serving(accepted, attempt, &inputs) + .await; + } // Ordinary acquisition can win after unused credit was joined. It // needs current serving proof, not a second ownership CAS. return self @@ -48,6 +80,18 @@ impl FleetActionExecutor { .await .map(ActionResult::checked); } + if closed_receiver_route + && self + .journal + .load_receiver_recovery_basis(accepted) + .await + .map_err(journal_error)? + .is_some() + { + return self + .resume_receiver_recovery(accepted, attempt, &inputs, observed) + .await; + } let basis = AcquisitionBasis::new(accepted.clone(), observed.value().clone(), wall_time_ms()?) .map_err(operation)?; diff --git a/crates/cellule-host/src/fleet/movement/inspection.rs b/crates/cellule-host/src/fleet/movement/inspection.rs index 8d1c0699..9ad64902 100644 --- a/crates/cellule-host/src/fleet/movement/inspection.rs +++ b/crates/cellule-host/src/fleet/movement/inspection.rs @@ -3,10 +3,15 @@ use super::*; impl FleetActionExecutor { pub(super) async fn cancel( &self, + accepted: &AcceptedFleetAction, attempt: &MoveAttempt, ) -> cellule_runtime::Result { if self.runtime.prepared_receiver(attempt.spec().id)?.is_none() { - return if self.confirmed_credit_settlement(attempt).await? { + let confirmed = self.confirmed_credit_settlement(attempt).await?; + let routed_empty = accepted.action().receiver_route().is_some() + && !confirmed + && self.routed_receiver_has_no_activation(attempt).await?; + return if confirmed || routed_empty { Ok(ActionResult::checked(FleetOutcome::ReceiverCleaned)) } else { Err(Error::Peer( @@ -29,6 +34,53 @@ impl FleetActionExecutor { Ok(ActionResult::checked(FleetOutcome::ReceiverCleaned)) } + async fn routed_receiver_has_no_activation( + &self, + attempt: &MoveAttempt, + ) -> cellule_runtime::Result { + for effect in [MovementAction::Activate, MovementAction::Recover] { + let retained = self + .journal + .load_movement_action( + self.scope, + attempt.spec().id, + effect, + self.node, + self.session, + ) + .await + .map_err(journal_error)?; + let Some(FleetActionAcceptance::Existing { accepted, result }) = retained else { + continue; + }; + if accepted.node() != self.node || accepted.session() != self.session { + return Err(Error::Fenced); + } + accepted + .validate_replay(accepted.action(), self.node, self.session) + .map_err(operation)?; + let FleetActionKind::Movement { + action: accepted_effect, + attempt: accepted_attempt, + } = accepted.action().kind() + else { + return Err(Error::Fenced); + }; + if *accepted_effect != effect || accepted_attempt.spec() != attempt.spec() { + return Err(Error::Fenced); + } + let Some(result) = result else { + return Ok(false); + }; + accepted.validate_result(&result).map_err(operation)?; + // A route with no local acquisition record is settled. Any retained + // record must already be recognized by confirmed_credit_settlement; + // unresolved or nonterminal evidence stays charged. + return Ok(false); + } + Ok(true) + } + pub(super) async fn confirmed_credit_settlement( &self, attempt: &MoveAttempt, diff --git a/crates/cellule-host/src/fleet/movement/mod.rs b/crates/cellule-host/src/fleet/movement/mod.rs index 5cd2caae..829f4983 100644 --- a/crates/cellule-host/src/fleet/movement/mod.rs +++ b/crates/cellule-host/src/fleet/movement/mod.rs @@ -17,11 +17,15 @@ mod activation; mod inspection; mod prefix; mod receiver; +mod receiver_recovery; +mod receiver_resume; mod recovery; +mod recovery_resume; pub(super) enum ServingPrefix<'a> { Released(&'a cellule_runtime::control::RootRef), Recovered(&'a cellule_runtime::fleet::operations::RecoveryEvidence), + ReceiverRecovered(&'a cellule_runtime::fleet::operations::ReceiverRecoveryEvidence), } impl FleetActionExecutor { @@ -101,7 +105,7 @@ impl FleetActionExecutor { MovementAction::Recover => self.recover(accepted, attempt).await, MovementAction::Prepare => self.prepare(attempt).await, MovementAction::Activate => self.activate(accepted, attempt).await, - MovementAction::Cancel => self.cancel(attempt).await, + MovementAction::Cancel => self.cancel(accepted, attempt).await, MovementAction::Inspect => Err(Error::Control( "fleet Inspect requires request-bound inspection", )), @@ -117,7 +121,7 @@ impl FleetActionExecutor { accepted.action().kind(), FleetActionKind::Maintenance { .. } ) { - return self.perform_maintenance(accepted); + return self.perform_maintenance(accepted, accepted.action(), None); } let FleetActionKind::Movement { action, attempt } = accepted.action().kind() else { return Err(Error::Control("accepted fleet effect is not movement")); @@ -128,6 +132,17 @@ impl FleetActionExecutor { match action { MovementAction::Recover => self.inspect_recovery(accepted, attempt).await, MovementAction::Prepare => self.inspect_preparation(attempt), + MovementAction::Activate if accepted.action().receiver_route().is_some() => { + if self.runtime.prepared_receiver(attempt.spec().id)?.is_some() { + return Err(Error::Peer( + "routed activation unexpectedly owns a prepared receiver credit", + )); + } + // Re-read canonical control and retry only through the exact + // routed acceptance. AcquisitionBasis checks the immutable + // release root before a new local acquisition can begin. + self.activate(accepted, attempt).await + } MovementAction::Activate => { if self.runtime.prepared_receiver(attempt.spec().id)?.is_none() && self.confirmed_credit_settlement(attempt).await? @@ -173,7 +188,7 @@ impl FleetActionExecutor { _ => Ok(ActionResult::checked(FleetOutcome::Unknown)), } } - MovementAction::Cancel => self.cancel(attempt).await, + MovementAction::Cancel => self.cancel(accepted, attempt).await, MovementAction::Inspect => Err(Error::Control( "fleet Inspect requires request-bound inspection", )), diff --git a/crates/cellule-host/src/fleet/movement/prefix.rs b/crates/cellule-host/src/fleet/movement/prefix.rs index a46f3090..4931be70 100644 --- a/crates/cellule-host/src/fleet/movement/prefix.rs +++ b/crates/cellule-host/src/fleet/movement/prefix.rs @@ -24,82 +24,178 @@ impl FleetActionExecutor { .await?; } ServingPrefix::Recovered(recovery) => { - // Charge canonical acquisition metadata, the 2-MiB manifest - // envelope and decoded row/vector growth before provider I/O. - // Hold the token through the complete suffix/origin proof. + // Charge canonical metadata and bounded manifest growth before + // provider I/O; hold it through the suffix/origin verification. let _recovery_memory = self.runtime.try_reserve_node_bytes(8 << 20)?; - let original = recovery.basis().control(); - let restored = recovery.restored(); - let canonical = inputs - .authority - .acquisition_record(original.cell, original.incarnation, restored.epoch) - .await? - .ok_or(Error::AcquisitionHistoryIncomplete { - cell: original.cell, - incarnation: original.incarnation, - epoch: restored.epoch, - })?; - // Journal shape alone cannot prove native materialization. Bind - // its entire original input/result to canonical acquisition. - if canonical.input() != original || canonical.materialized() != restored { - return Err(Error::Control( - "journal recovery differs from canonical acquisition", - )); - } - if let Some(overlay) = &original.recovery { - // The overlay can survive interrupted earlier claims. Its - // original Cell epoch comes from the digest-verified sealed - // manifest, never from this later acquisition's epoch. - let stores = self.cells.recovery_inputs(spec).await.map_err(|source| { + let stores = if recovery.basis().control().recovery.is_some() { + Some(self.cells.recovery_inputs(spec).await.map_err(|source| { Error::Facility { name: "fleet-recovery-provider", source, } - })?; - let inventory = stores - .manifests - .load_manifest( - overlay.leader_session, - overlay.log_epoch, - overlay.manifest_digest, - ) - .await?; - let mut rows = inventory.cells().iter().filter(|row| { - row.application == spec.target.application() - && row.cell == original.cell - && row.incarnation == original.incarnation - && &row.recovery == overlay - }); - let suffix = rows.next().ok_or(Error::Control( - "recovered suffix is absent from its manifest", - ))?; - if rows.next().is_some() { - return Err(Error::Control("recovered suffix manifest is ambiguous")); - } - self.runtime - .verify_recovered_prefix( - &inputs.catalog, - &inputs.authority, - inputs.replica.clone(), - suffix, - root, - 10_000, - ) - .await?; + })?) } else { - self.runtime - .verify_root_prefix( - &inputs.catalog, - &inputs.authority, - inputs.replica.clone(), - restored.ltx_root().ok_or(Error::Fenced)?, - root, - 10_000, - ) - .await?; - } + None + }; + self.verify_materialized_prefix( + attempt, + inputs, + RecoveryPrefixInput { + original: recovery.basis().control(), + restored: recovery.restored(), + stores: stores.as_ref(), + }, + root, + ) + .await?; } + ServingPrefix::ReceiverRecovered(recovery) => { + let _recovery_memory = self.runtime.try_reserve_node_bytes(8 << 20)?; + let basis = recovery.basis(); + let stores = if basis.control().recovery.is_some() { + Some( + self.cells + .receiver_recovery_inputs(basis.accepted(), basis.control()) + .await + .map_err(|source| Error::Facility { + name: "fleet-receiver-recovery-provider", + source, + })? + .ok_or(Error::Peer("receiver recovery prerequisites unavailable"))?, + ) + } else { + None + }; + self.verify_materialized_prefix( + attempt, + inputs, + RecoveryPrefixInput { + original: basis.control(), + restored: recovery.restored(), + stores: stores.as_ref(), + }, + root, + ) + .await?; + let released = attempt.released().ok_or(Error::Fenced)?; + self.runtime + .verify_root_prefix( + &inputs.catalog, + &inputs.authority, + inputs.replica.clone(), + released + .root + .to_ltx(spec.target.cell_id(), spec.incarnation), + root, + 10_000, + ) + .await?; + } + } + Ok(()) + } + + async fn verify_materialized_prefix( + &self, + attempt: &MoveAttempt, + inputs: &FleetCellInputs, + recovery: RecoveryPrefixInput<'_>, + root: cellule_runtime::ltx::RootRef, + ) -> cellule_runtime::Result<()> { + let spec = attempt.spec(); + let RecoveryPrefixInput { + original, + restored, + stores, + } = recovery; + let canonical = inputs + .authority + .acquisition_record(original.cell, original.incarnation, restored.epoch) + .await? + .ok_or(Error::AcquisitionHistoryIncomplete { + cell: original.cell, + incarnation: original.incarnation, + epoch: restored.epoch, + })?; + // Journal shape is historical input, not proof of native materialization. + if canonical.input() != original || canonical.materialized() != restored { + return Err(Error::Control( + "journal recovery differs from canonical acquisition", + )); + } + if let Some(overlay) = &original.recovery { + // An overlay may survive earlier interrupted claims. Its original + // Cell epoch comes from the sealed manifest, never the latest claim. + let stores = stores.ok_or(Error::Peer("recovered prefix lacks manifest store"))?; + let inventory = stores + .manifests + .load_manifest( + overlay.leader_session, + overlay.log_epoch, + overlay.manifest_digest, + ) + .await?; + let mut rows = inventory.cells().iter().filter(|row| { + row.application == spec.target.application() + && row.cell == original.cell + && row.incarnation == original.incarnation + && &row.recovery == overlay + }); + let suffix = rows.next().ok_or(Error::Control( + "recovered suffix is absent from its manifest", + ))?; + if rows.next().is_some() { + return Err(Error::Control("recovered suffix manifest is ambiguous")); + } + let observed = inputs + .authority + .load(original.cell) + .await? + .ok_or(Error::Fenced)?; + if observed.value().ltx_root() != Some(root) { + return Err(Error::Fenced); + } + if observed.value().state == ControlState::Idle { + self.runtime + .verify_recovered_idle_prefix( + &inputs.catalog, + &inputs.authority, + inputs.replica.clone(), + suffix, + &observed, + 10_000, + ) + .await?; + } else { + self.runtime + .verify_recovered_prefix( + &inputs.catalog, + &inputs.authority, + inputs.replica.clone(), + suffix, + root, + 10_000, + ) + .await?; + } + } else { + self.runtime + .verify_root_prefix( + &inputs.catalog, + &inputs.authority, + inputs.replica.clone(), + restored.ltx_root().ok_or(Error::Fenced)?, + root, + 10_000, + ) + .await?; } Ok(()) } } + +struct RecoveryPrefixInput<'a> { + original: &'a cellule_runtime::control::Control, + restored: &'a cellule_runtime::control::Control, + stores: Option<&'a crate::fleet::FleetRecoveryInputs>, +} diff --git a/crates/cellule-host/src/fleet/movement/receiver_recovery.rs b/crates/cellule-host/src/fleet/movement/receiver_recovery.rs new file mode 100644 index 00000000..71ce0da3 --- /dev/null +++ b/crates/cellule-host/src/fleet/movement/receiver_recovery.rs @@ -0,0 +1,289 @@ +//! Failed receiver takeover reuses the ordinary observed acquisition path. +use super::*; +use crate::fleet::FleetActionJournal; +use cellule_runtime::cell::actor::{AcquisitionObservation, AcquisitionObserver}; +use cellule_runtime::control::Control; +use cellule_runtime::fleet::operations::{ReceiverRecoveryBasis, ReceiverRecoveryEvidence}; +use cellule_runtime::node::NodeTakeoverProof; +use std::sync::Arc; + +struct ReceiverRecoveryRecorder { + accepted: AcceptedFleetAction, + takeover: NodeTakeoverProof, + journal: Arc, +} +impl AcquisitionObserver for ReceiverRecoveryRecorder { + fn before_claim<'a>(&'a self, input: &'a Control) -> AcquisitionObservation<'a> { + Box::pin(async move { + let basis = ReceiverRecoveryBasis::new( + self.accepted.clone(), + input.clone(), + self.takeover, + wall_time_ms()?, + ) + .map_err(operation)?; + let retained = self + .journal + .record_receiver_recovery_basis(&basis) + .await + .map_err(journal_error)?; + if retained.accepted() != &self.accepted + || retained.control() != input + || retained.observed_at_ms() > basis.observed_at_ms() + { + return Err(Error::Peer( + "journal changed checked receiver recovery basis", + )); + } + Ok(()) + }) + } + fn before_activation<'a>( + &'a self, + input: &'a Control, + restored: &'a Control, + ) -> AcquisitionObservation<'a> { + Box::pin(async move { + let basis = self + .journal + .load_receiver_recovery_basis(&self.accepted) + .await + .map_err(journal_error)? + .ok_or(Error::Peer("receiver recovery lacks retained input"))?; + if basis.accepted() != &self.accepted || basis.control() != input { + return Err(Error::Fenced); + } + let evidence = ReceiverRecoveryEvidence::new(basis, restored.clone(), wall_time_ms()?) + .map_err(operation)?; + let retained = self + .journal + .record_receiver_recovery_evidence(&evidence) + .await + .map_err(journal_error)?; + if retained.basis() != evidence.basis() + || retained.restored() != restored + || retained.recorded_at_ms() > evidence.recorded_at_ms() + { + return Err(Error::Peer( + "journal changed checked receiver recovery result", + )); + } + Ok(()) + }) + } +} + +impl FleetActionExecutor { + pub(super) async fn recover_receiver( + &self, + accepted: &AcceptedFleetAction, + attempt: &MoveAttempt, + inputs: &FleetCellInputs, + observed: VersionedControl, + ) -> cellule_runtime::Result { + if let Some(basis) = self + .journal + .load_receiver_recovery_basis(accepted) + .await + .map_err(journal_error)? + && observed + .value() + .owner + .as_ref() + .is_some_and(|owner| owner.session == self.session) + { + if basis.accepted() != accepted { + return Err(Error::Fenced); + } + let released = attempt.released().ok_or(Error::Fenced)?; + self.verify_serving_prefix( + attempt, + inputs, + ServingPrefix::Released(&released.root), + observed.value().ltx_root().ok_or(Error::Fenced)?, + ) + .await?; + let recovery = self + .cells + .receiver_recovery_inputs(accepted, basis.control()) + .await + .map_err(|source| Error::Facility { + name: "fleet-receiver-recovery-provider", + source, + })? + .ok_or(Error::Peer("receiver recovery prerequisites unavailable"))?; + let recorder: Arc = Arc::new(ReceiverRecoveryRecorder { + accepted: accepted.clone(), + takeover: recovery.takeover, + journal: self.journal.clone(), + }); + self.runtime + .resume_takeover_restored_observed( + inputs.catalog.clone(), + inputs.replica.clone(), + inputs.authority.clone(), + basis.control().clone(), + observed, + recovery.takeover, + recovery.manifests, + inputs.destination.clone(), + recorder, + ) + .await?; + return self + .receiver_recovered_serving(accepted, attempt, inputs) + .await; + } + let Some(recovery) = self + .cells + .receiver_recovery_inputs(accepted, observed.value()) + .await + .map_err(|source| Error::Facility { + name: "fleet-receiver-recovery-provider", + source, + })? + else { + return Ok(ActionResult::checked(FleetOutcome::Unknown)); + }; + ReceiverRecoveryBasis::new( + accepted.clone(), + observed.value().clone(), + recovery.takeover, + wall_time_ms()?, + ) + .map_err(operation)?; + if let Some(original) = self + .journal + .load_receiver_recovery_basis(accepted) + .await + .map_err(journal_error)? + && (original.accepted() != accepted || original.control() != observed.value()) + { + return Ok(ActionResult::checked(FleetOutcome::Unknown)); + } + // Prove the failed receiver still derives from the exact release before + // takeover. Its pinned overlay is materialized by native recovery. + let released = attempt.released().ok_or(Error::Fenced)?; + self.verify_serving_prefix( + attempt, + inputs, + ServingPrefix::Released(&released.root), + observed.value().ltx_root().ok_or(Error::Fenced)?, + ) + .await?; + let recorder: Arc = Arc::new(ReceiverRecoveryRecorder { + accepted: accepted.clone(), + takeover: recovery.takeover, + journal: self.journal.clone(), + }); + self.runtime + .takeover_restored_observed( + inputs.catalog.clone(), + inputs.replica.clone(), + inputs.authority.clone(), + observed, + recovery.takeover, + recovery.manifests, + inputs.destination.clone(), + inputs.owner.clone(), + Some(recorder), + ) + .await?; + self.receiver_recovered_serving(accepted, attempt, inputs) + .await + } + + pub(super) async fn receiver_recovered_serving( + &self, + accepted: &AcceptedFleetAction, + attempt: &MoveAttempt, + inputs: &FleetCellInputs, + ) -> cellule_runtime::Result { + let evidence = self + .confirm_receiver_recovery_evidence(accepted, inputs) + .await?; + let serving = self + .serving_evidence(attempt, inputs, ServingPrefix::ReceiverRecovered(&evidence)) + .await?; + let outcome = FleetOutcome::Activated(serving); + let envelope = cellule_runtime::fleet::operations::FleetActionOutcome { + scope: self.scope, + action_key: accepted.action().key().map_err(operation)?, + node: self.node, + session: self.session, + observed_at_ms: wall_time_ms()?, + outcome: outcome.clone(), + }; + evidence.validate_result(&envelope).map_err(operation)?; + Ok(ActionResult::checked(outcome)) + } + + // Effect replay may repair an interrupted evidence write. This lookup is + // never used by read-only inspection and starts no acquisition. Native + // history retained the exact input/materialization before actor admission; + // current owner, counters or root equality cannot replace that history. + pub(super) async fn confirm_receiver_recovery_evidence( + &self, + accepted: &AcceptedFleetAction, + inputs: &FleetCellInputs, + ) -> cellule_runtime::Result { + if let Some(evidence) = self + .journal + .load_receiver_recovery_evidence(accepted) + .await + .map_err(journal_error)? + { + if evidence.basis().accepted() != accepted { + return Err(Error::Fenced); + } + return Ok(evidence); + } + let basis = self + .journal + .load_receiver_recovery_basis(accepted) + .await + .map_err(journal_error)? + .ok_or(Error::Peer( + "serving receiver lacks retained recovery input", + ))?; + if basis.accepted() != accepted { + return Err(Error::Fenced); + } + let original = basis.control(); + let epoch = original + .epoch + .checked_add(1) + .ok_or(Error::Control("receiver recovery epoch overflow"))?; + let canonical = inputs + .authority + .acquisition_record(original.cell, original.incarnation, epoch) + .await? + .ok_or(Error::AcquisitionHistoryIncomplete { + cell: original.cell, + incarnation: original.incarnation, + epoch, + })?; + if canonical.input() != original { + return Err(Error::Control( + "receiver recovery input differs from canonical acquisition", + )); + } + let evidence = + ReceiverRecoveryEvidence::new(basis, canonical.materialized().clone(), wall_time_ms()?) + .map_err(operation)?; + let retained = self + .journal + .record_receiver_recovery_evidence(&evidence) + .await + .map_err(journal_error)?; + if retained.basis() != evidence.basis() + || retained.restored() != evidence.restored() + || retained.recorded_at_ms() > evidence.recorded_at_ms() + { + return Err(Error::Peer( + "journal changed reconstructed receiver recovery result", + )); + } + Ok(retained) + } +} diff --git a/crates/cellule-host/src/fleet/movement/receiver_resume.rs b/crates/cellule-host/src/fleet/movement/receiver_resume.rs new file mode 100644 index 00000000..70e8eec0 --- /dev/null +++ b/crates/cellule-host/src/fleet/movement/receiver_resume.rs @@ -0,0 +1,47 @@ +//! Resume a safely rolled-back receiver through ordinary admitted acquisition. +use super::*; + +impl FleetActionExecutor { + pub(super) async fn resume_receiver_recovery( + &self, + accepted: &AcceptedFleetAction, + attempt: &MoveAttempt, + inputs: &FleetCellInputs, + observed: VersionedControl, + ) -> cellule_runtime::Result { + let evidence = self + .confirm_receiver_recovery_evidence(accepted, inputs) + .await?; + if observed.value().state != ControlState::Idle + || observed.value().owner.is_some() + || observed.value().epoch < evidence.restored().epoch + { + return Ok(ActionResult::checked(FleetOutcome::Unknown)); + } + // The original takeover remains immutable. Prove its materialization, + // original suffix and exact release prefix before claiming the Idle + // root. A later root/epoch alone cannot authorize continuation. + self.verify_serving_prefix( + attempt, + inputs, + ServingPrefix::ReceiverRecovered(&evidence), + observed.value().ltx_root().ok_or(Error::Fenced)?, + ) + .await?; + // This exact accepted action already owns the journaled recovery basis. + // Native acquisition retains the new Idle input before admission; do + // not replace the original basis or mix it with AcquisitionBasis. + self.runtime + .acquire_idle_restored( + inputs.catalog.clone(), + inputs.replica.clone(), + inputs.authority.clone(), + observed, + inputs.destination.clone(), + inputs.owner.clone(), + ) + .await?; + self.receiver_recovered_serving(accepted, attempt, inputs) + .await + } +} diff --git a/crates/cellule-host/src/fleet/movement/recovery.rs b/crates/cellule-host/src/fleet/movement/recovery.rs index c1ce4b0a..7ba90885 100644 --- a/crates/cellule-host/src/fleet/movement/recovery.rs +++ b/crates/cellule-host/src/fleet/movement/recovery.rs @@ -171,6 +171,66 @@ impl FleetActionExecutor { return Err(Error::Fenced); } let inputs = self.inputs(attempt).await?; + if let Some(basis) = self + .journal + .load_recovery_basis(accepted) + .await + .map_err(journal_error)? + { + basis.validate_acceptance(accepted).map_err(operation)?; + let current = inputs + .authority + .load(attempt.spec().target.cell_id()) + .await? + .ok_or(Error::Fenced)?; + if current.value().state == ControlState::Recovering + && current + .value() + .owner + .as_ref() + .is_some_and(|owner| owner.session == self.session) + { + let recovery = + self.cells + .recovery_inputs(attempt.spec()) + .await + .map_err(|source| Error::Facility { + name: "fleet-recovery-provider", + source, + })?; + let recorder: Arc = Arc::new(RecoveryRecorder { + accepted: accepted.clone(), + takeover: recovery.takeover, + journal: self.journal.clone(), + }); + self.runtime + .resume_takeover_restored_observed( + inputs.catalog.clone(), + inputs.replica.clone(), + inputs.authority.clone(), + basis.control().clone(), + current, + recovery.takeover, + recovery.manifests, + inputs.destination.clone(), + recorder, + ) + .await?; + return self.recovered_serving(accepted, attempt, &inputs).await; + } + if current.value() != basis.control() + && (current.value().state == ControlState::Idle + || current + .value() + .owner + .as_ref() + .is_some_and(|owner| owner.session == self.session)) + { + return self + .resume_source_recovery(accepted, attempt, &inputs, current) + .await; + } + } if self .journal .load_recovery_evidence(accepted) diff --git a/crates/cellule-host/src/fleet/movement/recovery_resume.rs b/crates/cellule-host/src/fleet/movement/recovery_resume.rs new file mode 100644 index 00000000..c5673162 --- /dev/null +++ b/crates/cellule-host/src/fleet/movement/recovery_resume.rs @@ -0,0 +1,107 @@ +//! Preserve the original failed-source evidence across safe native rollback. +use super::*; +use cellule_runtime::fleet::operations::RecoveryEvidence; + +impl FleetActionExecutor { + pub(super) async fn resume_source_recovery( + &self, + accepted: &AcceptedFleetAction, + attempt: &MoveAttempt, + inputs: &FleetCellInputs, + observed: VersionedControl, + ) -> cellule_runtime::Result { + let evidence = self + .confirm_source_recovery_evidence(accepted, inputs) + .await?; + if observed.value().state == ControlState::Idle { + if observed.value().owner.is_some() + || observed.value().epoch < evidence.restored().epoch + { + return Err(Error::Fenced); + } + // Prove native materialization and the original pinned suffix before + // another ordinary claim. An Idle root or newer epoch alone cannot + // replace the original acquisition and acknowledged history. + self.verify_serving_prefix( + attempt, + inputs, + ServingPrefix::Recovered(&evidence), + observed.value().ltx_root().ok_or(Error::Fenced)?, + ) + .await?; + self.runtime + .acquire_idle_restored( + inputs.catalog.clone(), + inputs.replica.clone(), + inputs.authority.clone(), + observed, + inputs.destination.clone(), + inputs.owner.clone(), + ) + .await?; + } + self.recovered_serving(accepted, attempt, inputs).await + } + + // Only accepted effect replay reaches this repair. Read-only inspection + // keeps missing evidence blocking and never starts an acquisition. + async fn confirm_source_recovery_evidence( + &self, + accepted: &AcceptedFleetAction, + inputs: &FleetCellInputs, + ) -> cellule_runtime::Result { + if let Some(evidence) = self + .journal + .load_recovery_evidence(accepted) + .await + .map_err(journal_error)? + { + evidence + .basis() + .validate_acceptance(accepted) + .map_err(operation)?; + return Ok(evidence); + } + let basis = self + .journal + .load_recovery_basis(accepted) + .await + .map_err(journal_error)? + .ok_or(Error::Peer("recovery lacks retained original input"))?; + basis.validate_acceptance(accepted).map_err(operation)?; + let original = basis.control(); + let epoch = original + .epoch + .checked_add(1) + .ok_or(Error::Control("recovery epoch overflow"))?; + let canonical = inputs + .authority + .acquisition_record(original.cell, original.incarnation, epoch) + .await? + .ok_or(Error::AcquisitionHistoryIncomplete { + cell: original.cell, + incarnation: original.incarnation, + epoch, + })?; + if canonical.input() != original { + return Err(Error::Control( + "recovery input differs from canonical acquisition", + )); + } + let evidence = + RecoveryEvidence::new(basis, canonical.materialized().clone(), wall_time_ms()?) + .map_err(operation)?; + let retained = self + .journal + .record_recovery_evidence(accepted, &evidence) + .await + .map_err(journal_error)?; + if retained.basis() != evidence.basis() + || retained.restored() != evidence.restored() + || retained.recorded_at_ms() > evidence.recorded_at_ms() + { + return Err(Error::Peer("journal changed reconstructed recovery result")); + } + Ok(retained) + } +} diff --git a/crates/cellule-host/src/fleet/reconciler/maintenance.rs b/crates/cellule-host/src/fleet/reconciler/maintenance.rs index 81f6525f..43176a1a 100644 --- a/crates/cellule-host/src/fleet/reconciler/maintenance.rs +++ b/crates/cellule-host/src/fleet/reconciler/maintenance.rs @@ -1,11 +1,14 @@ use super::*; -use cellule_runtime::fleet::operations::{FleetOutcome, MaintenanceAction, MaintenanceEvent}; +use cellule_runtime::fleet::operations::{ + FleetOutcome, MaintenanceAction, MaintenanceEvent, MaintenanceOperation, +}; impl FleetReconciler { pub(super) async fn advance_maintenance( &self, clock: &PassClock<'_>, report: &mut FleetReconcileReport, + pass_observation: &mut Option, ) -> Result<()> { let Some(maintenance) = report.snapshot.head().maintenance().cloned() else { return Ok(()); @@ -71,10 +74,325 @@ impl FleetReconciler { ) .await?; } - MaintenancePhase::Evacuating | MaintenancePhase::Closing => { - // Full role inventory and joined facility/withdrawal evidence - // are still required. Empty movement permits cannot prove them. - report.blocked(DrainBlocker::IncompleteObservation); + MaintenancePhase::Evacuating => { + let roster = + FleetRoster::collect(self.journal.as_ref(), &report.snapshot, clock.deadline) + .await?; + let observation = call( + clock.deadline, + "fleet-observer", + self.observer.observe(&roster, clock.now()?, clock.deadline), + ) + .await?; + if observation.scope != self.scope + || observation.registry != report.snapshot.registry() + { + return Err(operation(OperationError::Conflict)); + } + roster + .confirm(self.journal.as_ref(), clock.deadline) + .await?; + let now = clock.now()?; + let boot_complete = roster.covers_advertisements(&observation.nodes, now)?; + let enrollment_settled = roster.enrollments().iter().all(|row| { + row.status() != cellule_runtime::fleet::operations::EnrollmentStatus::Pending + }); + *pass_observation = Some(observation.with_roster(roster)?); + let observation = pass_observation.as_ref().ok_or(Error::Control( + "fleet maintenance observation was not retained", + ))?; + report.maintenance_policy = observation + .maintenance_policy_coverage() + .map(|coverage| coverage.progress()); + let placements = observation.placements(now)?; + if !observation.complete + || !boot_complete + || !enrollment_settled + || !observation.counts_match(&placements) + { + report.blocked(DrainBlocker::IncompleteObservation); + return Ok(()); + } + if report.snapshot.head().attempts().iter().any(|attempt| { + attempt.spec().source_node == maintenance.node() + || attempt.spec().destination_node == maintenance.node() + }) { + report.blocked(DrainBlocker::PendingPublication); + return Ok(()); + } + let action = report + .snapshot + .head() + .maintenance_action(MaintenanceAction::SettleRoles, now) + .map_err(operation)?; + let settlement = match observation.role_settlement(&action, now) { + Ok(settlement) => settlement, + Err(_) => { + report.blocked(DrainBlocker::IncompleteObservation); + return Ok(()); + } + }; + report.dispatched += 1; + let completion = if let Some(closure_digest) = settlement.failed_boot_closure() { + let closure = observation + .failed_boot_closures() + .and_then(|closures| { + closures + .iter() + .find(|closure| closure.digest() == closure_digest) + }) + .ok_or(Error::Control( + "closed-boot settlement lost its process closure", + ))?; + settlement.validate_failed_boot_closure(&action, closure)?; + call( + clock.deadline, + "fleet-transport", + self.transport.settle_roles_after_process_closure( + &action, + &settlement, + closure, + clock.deadline, + ), + ) + .await? + } else { + call( + clock.deadline, + "fleet-transport", + self.transport + .settle_roles(&action, &settlement, clock.deadline), + ) + .await? + }; + completion + .accepted + .validate_replay(&action, maintenance.node(), maintenance.session()) + .map_err(operation)?; + completion + .accepted + .validate_result(&completion.outcome) + .map_err(operation)?; + report.maintenance_failure = completion.execution_error.clone(); + if !completion.committed { + if let Some(error) = &completion.journal_error { + report.maintenance_failure = Some(Arc::clone(error)); + } + report.blocked(DrainBlocker::PendingPublication); + return Ok(()); + } + match completion.outcome.outcome { + FleetOutcome::RolesSettledAt { + inventory, + head_revision, + registry, + } if inventory == settlement.inventory() + && head_revision == settlement.head_revision() + && registry == settlement.registry() => + { + self.commit( + clock, + report, + JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose( + cellule_runtime::fleet::operations::DrainEvidence { + node: maintenance.node(), + session: maintenance.session(), + remaining_cells: 0, + unresolved_attempts: 0, + relocated: true, + readers_settled: true, + followers_settled: true, + facilities_closed: false, + stopped: false, + withdrawn: false, + }, + )), + ) + .await?; + } + FleetOutcome::Rejected(blocker) | FleetOutcome::Blocked(blocker) => { + report.blocked(blocker); + } + FleetOutcome::Unknown => { + report.blocked(DrainBlocker::OutcomeUnknown); + } + _ => return Err(operation(OperationError::Conflict)), + } + } + MaintenancePhase::Closing => { + // The committed Closing record carries the complete relocation + // and role barrier. Finalize is a separately retained host task; + // its action result is committed only after drain and withdrawal. + let roster = + FleetRoster::collect(self.journal.as_ref(), &report.snapshot, clock.deadline) + .await?; + let target_boot = roster + .enrollments() + .iter() + .find(|row| { + row.spec().target.node == maintenance.node() + && row.spec().target.session == maintenance.session() + && matches!( + row.spec().role, + cellule_runtime::fleet::operations::EnrollmentRole::Node { .. } + ) + }) + .cloned(); + let Some(target_boot) = target_boot else { + report.blocked(DrainBlocker::IncompleteObservation); + return Ok(()); + }; + let target_retired = target_boot.status() + == cellule_runtime::fleet::operations::EnrollmentStatus::Retired; + let mut closed_observation = None; + if target_retired { + let observation = call( + clock.deadline, + "fleet-observer", + self.observer.observe(&roster, clock.now()?, clock.deadline), + ) + .await?; + if observation.scope != self.scope + || observation.registry != report.snapshot.registry() + { + return Err(operation(OperationError::Conflict)); + } + let observation = observation.with_roster(roster)?; + observation + .roster() + .ok_or(Error::Control("closed-boot roster was not retained"))? + .confirm(self.journal.as_ref(), clock.deadline) + .await?; + let closure = observation.failed_boot_closures().and_then(|closures| { + closures + .iter() + .find(|closure| closure.boot() == &target_boot) + }); + if closure.is_none_or(|closure| { + closure.snapshot() != &report.snapshot + || closure.boot().status() + != cellule_runtime::fleet::operations::EnrollmentStatus::Retired + || closure.canonical().node() != maintenance.node() + || closure.canonical().session() != maintenance.session() + }) { + report.blocked(DrainBlocker::IncompleteObservation); + return Ok(()); + } + closed_observation = Some(observation); + } + let action = report + .snapshot + .head() + .maintenance_action(MaintenanceAction::Finalize, clock.now()?) + .map_err(operation)?; + report.dispatched += 1; + let completion = if let Some(observation) = &closed_observation { + let closure = observation + .failed_boot_closures() + .and_then(|closures| { + closures.iter().find(|closure| { + closure.boot().spec().target.node == maintenance.node() + && closure.boot().spec().target.session == maintenance.session() + }) + }) + .ok_or(Error::Control( + "closed-boot finalization lost its process closure", + ))?; + call( + clock.deadline, + "fleet-transport", + self.transport.finalize_after_process_closure( + &action, + closure, + clock.deadline, + ), + ) + .await? + } else { + call( + clock.deadline, + "fleet-transport", + self.transport.dispatch(&action, clock.deadline), + ) + .await? + }; + completion + .accepted + .validate_replay(&action, maintenance.node(), maintenance.session()) + .map_err(operation)?; + completion + .accepted + .validate_result(&completion.outcome) + .map_err(operation)?; + report.maintenance_failure = completion.execution_error.clone(); + if !completion.committed { + if let Some(error) = &completion.journal_error { + report.maintenance_failure = Some(Arc::clone(error)); + } + report.blocked(DrainBlocker::PendingPublication); + return Ok(()); + } + match completion.outcome.outcome { + FleetOutcome::Stopped(evidence) => { + // Exact boot retirement advances the registry inside + // the enrollment journal. Refresh the complete snapshot + // before the head CAS, then revalidate that the same + // Closing operation and evidence still own this result. + let latest = call( + clock.deadline, + "fleet-journal", + self.journal.load_snapshot(self.scope), + ) + .await?; + let latest_epoch = self.controller_epoch(&latest, clock.now()?)?; + let dispatched_epoch = report + .snapshot + .head() + .controller() + .map(|lease| lease.epoch) + .ok_or_else(|| operation(OperationError::Fenced))?; + if latest_epoch != dispatched_epoch { + return Err(operation(OperationError::Fenced)); + } + let current = latest + .head() + .maintenance() + .ok_or_else(|| operation(OperationError::NotFound))?; + if !same_maintenance_identity(current, &maintenance) { + return Err(operation(OperationError::Conflict)); + } + match current.phase() { + MaintenancePhase::Closing + if current.drain_evidence() == maintenance.drain_evidence() => + { + report.snapshot = latest; + self.commit( + clock, + report, + JournalTransition::Maintenance(MaintenanceEvent::Stopped( + evidence, + )), + ) + .await?; + } + MaintenancePhase::Completed + if current.drain_evidence() == Some(evidence) => + { + // Another controller already published this + // exact terminal result while this waiter ran. + report.snapshot = latest; + } + _ => return Err(operation(OperationError::Conflict)), + } + } + FleetOutcome::Rejected(blocker) | FleetOutcome::Blocked(blocker) => { + report.blocked(blocker); + } + FleetOutcome::Unknown => { + report.blocked(DrainBlocker::OutcomeUnknown); + } + _ => return Err(operation(OperationError::Conflict)), + } } MaintenancePhase::Completed => {} } @@ -86,3 +404,14 @@ impl FleetReconciler { Ok(()) } } + +fn same_maintenance_identity( + current: &MaintenanceOperation, + dispatched: &MaintenanceOperation, +) -> bool { + current.id() == dispatched.id() + && current.request_digest() == dispatched.request_digest() + && current.node() == dispatched.node() + && current.session() == dispatched.session() + && current.intent_revision() == dispatched.intent_revision() +} diff --git a/crates/cellule-host/src/fleet/reconciler/mod.rs b/crates/cellule-host/src/fleet/reconciler/mod.rs index bf27dc4a..de158648 100644 --- a/crates/cellule-host/src/fleet/reconciler/mod.rs +++ b/crates/cellule-host/src/fleet/reconciler/mod.rs @@ -2,10 +2,9 @@ use std::{future::Future, sync::Arc, time::Duration}; -use cellule_runtime::fleet::operations::AttemptId; use cellule_runtime::fleet::operations::{ - DrainBlocker, FleetAction, FleetInspectionObservation, FleetInspectionRequest, FleetProfile, - FleetScope, JournalTransition, MaintenancePhase, OperationError, + AttemptId, DrainBlocker, FleetAction, FleetInspectionObservation, FleetInspectionRequest, + FleetProfile, FleetScope, JournalTransition, MaintenancePhase, OperationError, }; use cellule_runtime::identity::SessionId; use cellule_runtime::{Error, Result}; @@ -22,7 +21,7 @@ mod movement; mod observation; mod planning; mod successor; -pub use observation::{FleetObservation, FleetOwnedCell}; +pub use observation::{FleetObservation, FleetOwnedCell, FleetRoleSettlement}; /// Application-owned complete roster and authenticated paginated observation. /// @@ -54,6 +53,62 @@ pub trait FleetTransport: Send + Sync + 'static { deadline: Instant, ) -> FleetAdapterFuture<'a, Arc>; + /// Publishes a fresh complete role settlement at the exact action barrier. + /// Network adapters authenticate the action and carry the opaque proof to + /// the bound target node. The default keeps adapters that do not support + /// role settlement fail closed. + fn settle_roles<'a>( + &'a self, + _action: &'a FleetAction, + _settlement: &'a super::FleetRoleSettlement, + _deadline: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + Box::pin(async { + Err(Box::new(Error::Control( + "fleet role settlement transport is not configured", + )) as Box) + }) + } + + /// Publishes SettleRoles after the exact maintenance boot has already + /// stopped. Implementations must validate the opaque host settlement and + /// matching retained process-closure proof, then durably accept and publish + /// the exact RolesSettledAt result through the fleet journal. They must not + /// dispatch to, or infer completion from, the stopped endpoint. The default + /// refuses closed-boot settlement. + fn settle_roles_after_process_closure<'a>( + &'a self, + _action: &'a FleetAction, + _settlement: &'a super::FleetRoleSettlement, + _closure: &'a crate::fleet::FleetFailedBootClosure, + _deadline: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + Box::pin(async { + Err(Box::new(Error::Control( + "closed-boot role settlement transport is not configured", + )) as Box) + }) + } + + /// Publishes terminal Finalize evidence after the operation's exact boot + /// has already been retired. Implementations must verify that the fresh + /// closure matches the Closing action and exact current journal snapshot, + /// then durably accept and publish `Stopped` using the committed ready-to- + /// close evidence. They must not contact the stopped endpoint. The default + /// refuses this path. + fn finalize_after_process_closure<'a>( + &'a self, + _action: &'a FleetAction, + _closure: &'a crate::fleet::FleetFailedBootClosure, + _deadline: Instant, + ) -> FleetAdapterFuture<'a, Arc> { + Box::pin(async { + Err(Box::new(Error::Control( + "closed-boot finalization transport is not configured", + )) as Box) + }) + } + /// Captures current authority and actor evidence for the entire request. /// Cached effect receipts cannot satisfy this boundary. fn inspect<'a>( @@ -110,6 +165,21 @@ pub struct FleetReconcileReport { } impl FleetReconcileReport { + fn failed(&mut self, attempt: AttemptId, error: Error) { + // A fallback may progress after an endpoint fails. Preserve the first + // source for that attempt without exceeding the per-attempt bound. + if !self + .failures + .iter() + .any(|failure| failure.attempt == attempt) + { + self.failures.push(FleetAttemptFailure { + attempt, + error: Arc::new(error), + }); + } + } + fn blocked(&mut self, reason: DrainBlocker) { if !self.blockers.contains(&reason) { self.blockers.push(reason); @@ -264,14 +334,12 @@ impl FleetReconciler { report.snapshot = snapshot; } report.blocked(DrainBlocker::OutcomeUnknown); - report.failures.push(FleetAttemptFailure { - attempt: id, - error: Arc::new(error), - }); + report.failed(id, error); } } + let mut pass_observation = None; if let Err(error) = self - .advance_maintenance(&clock.partition(2), &mut report) + .advance_maintenance(&clock.partition(2), &mut report, &mut pass_observation) .await { let timed_out = matches!( @@ -308,7 +376,8 @@ impl FleetReconciler { report.maintenance_failure = Some(Arc::new(error)); } if report.snapshot.registry().scheduling_enabled() { - self.plan(&clock, &mut report).await?; + self.plan(&clock, &mut report, &mut pass_observation) + .await?; } // Proven cancellation keeps the periodic retry interval, avoiding an // immediate allocation/refusal loop when a receiver cannot admit work. diff --git a/crates/cellule-host/src/fleet/reconciler/movement.rs b/crates/cellule-host/src/fleet/reconciler/movement.rs index 2cfcee57..0c8aa26e 100644 --- a/crates/cellule-host/src/fleet/reconciler/movement.rs +++ b/crates/cellule-host/src/fleet/reconciler/movement.rs @@ -1,9 +1,23 @@ use super::*; use cellule_runtime::fleet::operations::{ - AttemptEvent, AttemptId, AttemptPhase, FleetActionOutcome, FleetOutcome, MoveAttempt, - MovementAction, + AttemptEvent, AttemptId, AttemptPhase, EnrollmentRole, EnrollmentStatus, FleetAction, + FleetActionOutcome, FleetOutcome, MoveAttempt, MovementAction, ReceiverRoute, }; +use cellule_runtime::fleet::placement::{PlacementEligibility, PlacementPlanner}; use cellule_runtime::identity::{NodeId, SessionId}; +use cellule_runtime::node::NodeMode; +use std::collections::HashSet; + +struct RetainedMovementAcceptance { + hops: usize, + acceptance: super::super::FleetActionAcceptance, +} + +enum ReceiverDispatchChoice { + Direct, + Routed(FleetAction), + Blocked, +} impl FleetReconciler { pub(super) async fn advance( @@ -87,7 +101,43 @@ impl FleetReconciler { if let Some(event) = transition { self.commit_event(id, event, clock, report).await?; } else if effect == MovementAction::Inspect { - let outcome = self.inspect_attempt(&attempt, clock, report).await?; + let outcome = match self.inspect_attempt(&attempt, clock, report).await { + Ok(outcome) => outcome, + Err(error) + if attempt.phase() == AttemptPhase::Activating + && attempt.released().is_some() + && matches!( + &error, + Error::Facility { + name: "fleet-transport", + .. + } + ) => + { + // A stopped receiver cannot answer inspection. Only a + // fresh typed closure and committed route can authorize a + // continuation. Canonical claims still require the + // receiver executor's ordinary ownership checks. + let choice = match self + .receiver_dispatch_action(&attempt, MovementAction::Activate, clock, report) + .await + { + Ok(choice) => choice, + Err(route_error) => { + report.failed(id, error); + return Err(route_error); + } + }; + let ReceiverDispatchChoice::Routed(action) = choice else { + return Err(error); + }; + report.failed(id, error); + return self + .dispatch_action(&attempt, MovementAction::Activate, &action, clock, report) + .await; + } + Err(error) => return Err(error), + }; if attempt.phase() == AttemptPhase::Preparing && clock.now()? >= attempt.spec().deadline_ms { @@ -190,7 +240,33 @@ impl FleetReconciler { .await; } } - let action = if attempt.blocker() == Some(DrainBlocker::OutcomeUnknown) { + let action = if matches!(effect, MovementAction::Activate | MovementAction::Cancel) { + match self + .receiver_dispatch_action(&attempt, effect, clock, report) + .await? + { + ReceiverDispatchChoice::Routed(action) => action, + ReceiverDispatchChoice::Direct => { + if attempt.blocker() == Some(DrainBlocker::OutcomeUnknown) { + let Some(super::super::FleetActionAcceptance::Existing { + accepted, .. + }) = self.retained_action(&attempt, effect, clock).await? + else { + report.blocked(DrainBlocker::OutcomeUnknown); + return Ok(()); + }; + accepted.action().clone() + } else { + report + .snapshot + .head() + .movement_action(id, effect, clock.now()?) + .map_err(operation)? + } + } + ReceiverDispatchChoice::Blocked => return Ok(()), + } + } else if attempt.blocker() == Some(DrainBlocker::OutcomeUnknown) { // Unknown does not authorize a new effect. Replay only the exact // original acceptance; the node may inspect/resume its owned work. let Some(super::super::FleetActionAcceptance::Existing { accepted, .. }) = @@ -207,17 +283,36 @@ impl FleetReconciler { .movement_action(id, effect, clock.now()?) .map_err(operation)? }; + self.dispatch_action(&attempt, effect, &action, clock, report) + .await + } + + async fn dispatch_action( + &self, + attempt: &MoveAttempt, + effect: MovementAction, + action: &FleetAction, + clock: &PassClock<'_>, + report: &mut FleetReconcileReport, + ) -> Result<()> { + let id = attempt.spec().id; report.dispatched += 1; let completion = call( clock.deadline, "fleet-transport", - self.transport.dispatch(&action, clock.deadline), + self.transport.dispatch(action, clock.deadline), ) .await?; - let (node, session) = endpoint(&attempt, effect); + let (node, session) = if effect.is_source_release() { + endpoint(attempt, effect) + } else { + action + .receiver_endpoint() + .ok_or(Error::Control("fleet receiver action has no endpoint"))? + }; completion .accepted - .validate_replay(&action, node, session) + .validate_replay(action, node, session) .map_err(operation)?; completion .accepted @@ -234,7 +329,7 @@ impl FleetReconciler { completion.outcome.outcome, FleetOutcome::Activated(_) | FleetOutcome::Recovered(_) ) { - let fresh = self.inspect_attempt(&attempt, clock, report).await?; + let fresh = self.inspect_attempt(attempt, clock, report).await?; self.consume(id, &fresh, true, clock, report).await } else { self.consume(id, &completion.outcome, false, clock, report) @@ -242,6 +337,321 @@ impl FleetReconciler { } } + async fn receiver_dispatch_action( + &self, + attempt: &MoveAttempt, + effect: MovementAction, + clock: &PassClock<'_>, + report: &mut FleetReconcileReport, + ) -> Result { + let retained = self.retained_action(attempt, effect, clock).await?; + let accepted = match &retained { + Some(super::super::FleetActionAcceptance::Existing { accepted, .. }) => { + Some(accepted.action().clone()) + } + _ => None, + }; + + // A durable terminal result always wins over a new route. In particular, + // process closure alone cannot prove that a receiver did not publish a + // Serving control immediately before it stopped. + if matches!( + &retained, + Some(super::super::FleetActionAcceptance::Existing { + result: Some(result), + .. + }) if matches!( + &result.outcome, + FleetOutcome::Activated(_) + | FleetOutcome::Recovered(_) + | FleetOutcome::ReceiverCleaned + ) + ) { + return Ok(accepted.map_or( + ReceiverDispatchChoice::Direct, + ReceiverDispatchChoice::Routed, + )); + } + + let route_source = if effect == MovementAction::Cancel { + match self + .retained_action(attempt, MovementAction::Activate, clock) + .await? + { + Some(super::super::FleetActionAcceptance::Existing { accepted, .. }) => { + accepted.action().receiver_route().cloned() + } + _ => None, + } + } else { + None + }; + let current_route = accepted + .as_ref() + .and_then(|action| action.receiver_route().cloned()) + .or(route_source); + let previous = current_route.as_ref().map_or( + (attempt.spec().destination_node, attempt.spec().destination), + |route| route.target(attempt.spec()), + ); + + let roster = crate::fleet::FleetRoster::collect( + self.journal.as_ref(), + &report.snapshot, + clock.deadline, + ) + .await?; + let observation = call( + clock.deadline, + "fleet-observer", + self.observer.observe(&roster, clock.now()?, clock.deadline), + ) + .await?; + if observation.scope != self.scope || observation.registry != report.snapshot.registry() { + return Err(operation(OperationError::Conflict)); + } + roster + .confirm(self.journal.as_ref(), clock.deadline) + .await?; + let observation = observation.with_roster(roster)?; + let now = clock.now()?; + let placements = observation.placements(now)?; + let Some(closures) = observation.failed_boot_closures() else { + if effect == MovementAction::Cancel && current_route.is_some() && accepted.is_none() { + report.blocked(DrainBlocker::OutcomeUnknown); + return Ok(ReceiverDispatchChoice::Blocked); + } + return Ok(accepted.map_or( + ReceiverDispatchChoice::Direct, + ReceiverDispatchChoice::Routed, + )); + }; + let closure_for = |endpoint| { + closures.iter().find(|closure| { + let target = closure.boot().spec().target; + (target.node, target.session) == endpoint + && closure.boot().status() == EnrollmentStatus::Retired + && closure.snapshot() == &report.snapshot + }) + }; + + // An existing route continues to replay at its current endpoint while + // that boot remains live. A routed Cancel still needs a fresh typed + // closure for the original handoff in the same head/registry barrier. + if current_route + .as_ref() + .is_some_and(|route| route.target(attempt.spec()) == previous) + && closure_for(previous).is_none() + { + if effect == MovementAction::Activate { + return Ok(accepted.map_or( + ReceiverDispatchChoice::Blocked, + ReceiverDispatchChoice::Routed, + )); + } + if accepted + .as_ref() + .is_some_and(|action| action.receiver_route().is_some()) + { + return Ok(ReceiverDispatchChoice::Routed(accepted.ok_or( + Error::Control("routed cancellation acceptance is absent"), + )?)); + } + let Some(route) = current_route else { + return Ok(accepted.map_or( + ReceiverDispatchChoice::Direct, + ReceiverDispatchChoice::Routed, + )); + }; + let Some(handoff) = route.latest_handoff() else { + return Err(operation(OperationError::Conflict)); + }; + let Some(closure) = closure_for(handoff.previous()) else { + report.blocked(DrainBlocker::OutcomeUnknown); + return Ok(ReceiverDispatchChoice::Blocked); + }; + if route.registry() != Some(report.snapshot.registry()) { + report.blocked(DrainBlocker::StaleObservation); + return Ok(ReceiverDispatchChoice::Blocked); + } + return self + .accept_routed_receiver_action(attempt, effect, route, closure, clock, report) + .await; + } + + let Some(closure) = closure_for(previous) else { + return Ok(accepted.map_or( + ReceiverDispatchChoice::Direct, + ReceiverDispatchChoice::Routed, + )); + }; + let roster = observation + .roster() + .ok_or(Error::Control("fleet roster was not retained"))?; + if !roster.covers_advertisements(&observation.nodes, now)? { + report.blocked(DrainBlocker::IncompleteObservation); + return Ok(ReceiverDispatchChoice::Blocked); + } + + let mut projected = placements; + for other in report.snapshot.head().attempts() { + if other.spec().id == attempt.spec().id { + continue; + } + let mut endpoint = (other.spec().destination_node, other.spec().destination); + if let Some(super::super::FleetActionAcceptance::Existing { accepted, .. }) = self + .retained_action(other, MovementAction::Activate, clock) + .await? + && let Some(routed) = accepted.action().receiver_endpoint() + { + endpoint = routed; + } + if let Some(receiver) = projected + .iter_mut() + .find(|node| (node.node, node.session) == endpoint) + { + receiver.free_memory_bytes = receiver + .free_memory_bytes + .saturating_sub(other.spec().cost.memory_bytes); + receiver.free_disk_bytes = receiver + .free_disk_bytes + .saturating_sub(other.spec().cost.disk_bytes); + receiver.active_cells = receiver.active_cells.saturating_add(1); + receiver.running_jobs = receiver + .running_jobs + .saturating_add(other.spec().cost.job_credits); + } + } + + let planner = PlacementPlanner::default(); + let ranked = planner.rank(attempt.spec().target.cell_id(), now, &projected)?; + let mut visited = + HashSet::from([attempt.spec().source_node, attempt.spec().destination_node]); + if let Some(route) = ¤t_route + && let Some(handoff) = route.latest_handoff() + { + visited.insert(handoff.target().0); + } + let needs_capacity = effect == MovementAction::Activate; + let roster = observation + .roster() + .ok_or(Error::Control("fleet roster was not retained"))?; + let candidate = ranked.into_iter().find(|score| { + let usable = if needs_capacity { + score.eligibility == PlacementEligibility::Eligible + } else { + !matches!( + score.eligibility, + PlacementEligibility::Unauthenticated | PlacementEligibility::Stale + ) + }; + if !usable || visited.contains(&score.node) { + return false; + } + let Some(intent) = roster + .intents() + .iter() + .find(|intent| intent.node() == score.node) + else { + return false; + }; + if intent.session() != score.session || intent.mode() != NodeMode::Active { + return false; + } + let enrolled = roster.enrollments().iter().any(|row| { + row.status() == EnrollmentStatus::Established + && row.spec().target.node == score.node + && row.spec().target.session == score.session + && matches!(row.spec().role, EnrollmentRole::Node { .. }) + }); + if !enrolled { + return false; + } + !needs_capacity + || projected + .iter() + .find(|node| node.node == score.node && node.session == score.session) + .is_some_and(|node| { + node.free_memory_bytes >= attempt.spec().cost.memory_bytes + && node.free_disk_bytes >= attempt.spec().cost.disk_bytes + && node.max_active_cells.saturating_sub(node.active_cells) >= 1 + && node.job_capacity.saturating_sub(node.running_jobs) + >= attempt.spec().cost.job_credits + }) + }); + let Some(candidate) = candidate else { + report.blocked(if needs_capacity { + DrainBlocker::ReceiverCapacity + } else { + DrainBlocker::IncompleteObservation + }); + return Ok(ReceiverDispatchChoice::Blocked); + }; + + let route = match ¤t_route { + Some(route) => route + .extend( + self.scope, + attempt.spec(), + candidate.node, + candidate.session, + closure.digest(), + report.snapshot.registry(), + ) + .map_err(operation)?, + None => ReceiverRoute::begin( + self.scope, + attempt.spec(), + candidate.node, + candidate.session, + closure.digest(), + report.snapshot.registry(), + ) + .map_err(operation)?, + }; + self.accept_routed_receiver_action(attempt, effect, route, closure, clock, report) + .await + } + + async fn accept_routed_receiver_action( + &self, + attempt: &MoveAttempt, + effect: MovementAction, + route: ReceiverRoute, + closure: &crate::fleet::FleetFailedBootClosure, + clock: &PassClock<'_>, + report: &FleetReconcileReport, + ) -> Result { + let action = report + .snapshot + .head() + .movement_action_with_receiver_route(attempt.spec().id, effect, route, clock.now()?) + .map_err(operation)?; + let (node, session) = action + .receiver_endpoint() + .ok_or(Error::Control("routed receiver action has no endpoint"))?; + let acceptance = call( + clock.deadline, + "fleet-journal", + self.journal.accept_closed_receiver_action( + &action, + node, + session, + clock.now()?, + closure, + ), + ) + .await?; + let accepted = match acceptance { + super::super::FleetActionAcceptance::New(accepted) + | super::super::FleetActionAcceptance::Existing { accepted, .. } => accepted, + }; + accepted + .validate_replay(&action, node, session) + .map_err(operation)?; + Ok(ReceiverDispatchChoice::Routed(accepted.action().clone())) + } + async fn inspect_attempt( &self, attempt: &MoveAttempt, @@ -305,10 +715,7 @@ impl FleetReconciler { if let Err(error) = original { // Preserve the failed original endpoint even when another // boot proves serving. No failure can free receiver credit. - report.failures.push(super::FleetAttemptFailure { - attempt: attempt.spec().id, - error: Arc::new(error), - }); + report.failed(attempt.spec().id, error); } Ok(fresh) } @@ -372,42 +779,57 @@ impl FleetReconciler { effect: MovementAction, clock: &PassClock<'_>, ) -> Result> { - let (node, session) = endpoint(attempt, effect); - let original = call( + let actions = call( clock.deadline, "fleet-journal", self.journal - .load_movement_action(self.scope, attempt.spec().id, effect, node, session), + .load_movement_actions(self.scope, attempt, effect), ) .await?; - match &original { - None => {} - Some(super::super::FleetActionAcceptance::New(_)) => { + let mut selected: Option = None; + for candidate in actions { + let super::super::FleetActionAcceptance::Existing { accepted, result } = candidate + else { + return Err(operation(OperationError::Conflict)); + }; + accepted + .validate_replay(accepted.action(), accepted.node(), accepted.session()) + .map_err(operation)?; + let cellule_runtime::fleet::operations::FleetActionKind::Movement { + action, + attempt: input, + } = accepted.action().kind() + else { + return Err(operation(OperationError::Conflict)); + }; + if accepted.action().scope() != self.scope + || *action != effect + || input.spec() != attempt.spec() + { return Err(operation(OperationError::Conflict)); } - Some(super::super::FleetActionAcceptance::Existing { accepted, result }) => { - accepted - .validate_replay(accepted.action(), node, session) - .map_err(operation)?; - let cellule_runtime::fleet::operations::FleetActionKind::Movement { - action, - attempt: input, - } = accepted.action().kind() - else { - return Err(operation(OperationError::Conflict)); - }; - if accepted.action().scope() != self.scope - || *action != effect - || input.spec() != attempt.spec() - { - return Err(operation(OperationError::Conflict)); - } - if let Some(result) = result { - accepted.validate_result(result).map_err(operation)?; + if let Some(result) = &result { + accepted.validate_result(result).map_err(operation)?; + } + let route = accepted.action().receiver_route().cloned(); + let hops = route.as_ref().map_or( + 0, + cellule_runtime::fleet::operations::ReceiverRoute::hop_count, + ); + let candidate = RetainedMovementAcceptance { + hops, + acceptance: super::super::FleetActionAcceptance::Existing { accepted, result }, + }; + match &selected { + None => selected = Some(candidate), + Some(previous) if hops > previous.hops => { + selected = Some(candidate); } + Some(previous) if hops < previous.hops => {} + Some(_) => return Err(operation(OperationError::Conflict)), } } - Ok(original) + Ok(selected.map(|selected| selected.acceptance)) } async fn commit_event( diff --git a/crates/cellule-host/src/fleet/reconciler/observation/mod.rs b/crates/cellule-host/src/fleet/reconciler/observation/mod.rs index 383a0277..1fdef4e4 100644 --- a/crates/cellule-host/src/fleet/reconciler/observation/mod.rs +++ b/crates/cellule-host/src/fleet/reconciler/observation/mod.rs @@ -208,6 +208,12 @@ impl FleetObservation { .iter() .any(|node| node.node == owned.node && node.session == owned.session) || row.generation == 0 + || row.owner_fence.incarnation != row.incarnation + || row.owner_fence.epoch == 0 + || row.position.as_ref().is_some_and(|position| { + position.incarnation != row.owner_fence.incarnation + || position.epoch != row.owner_fence.epoch + }) || row.resident_since_ms < 0 || row.resident_since_ms > now_ms { @@ -223,7 +229,7 @@ impl FleetObservation { pub(super) fn digest(&self, now_ms: i64) -> Result { let nodes = self.placements(now_ms)?; let mut hash = blake3::Hasher::new(); - hash.update(b"cellule.fleet-planner-inputs.v14\0"); + hash.update(b"cellule.fleet-planner-inputs.v15\0"); hash.update(self.scope.fleet.as_bytes()); hash.update(self.scope.application.as_bytes()); hash.update(&self.registry.to_bytes().map_err(super::operation)?); @@ -323,6 +329,8 @@ impl FleetObservation { hash.update(owned.session.as_bytes()); hash.update(row.target.cell_id().as_bytes()); hash.update(row.incarnation.as_bytes()); + hash.update(row.owner_fence.incarnation.as_bytes()); + hash.update(&row.owner_fence.epoch.to_be_bytes()); hash.update(row.code.as_bytes()); hash.update(&row.schema.to_be_bytes()); // Role now controls busy-maintenance eligibility. Use explicit tags @@ -393,6 +401,8 @@ mod maintenance_policies; mod nonexecution; mod original_writers; mod recovered_followers; +mod role_settlement; mod source_readers; +pub use role_settlement::FleetRoleSettlement; #[cfg(test)] mod tests; diff --git a/crates/cellule-host/src/fleet/reconciler/observation/role_settlement/mod.rs b/crates/cellule-host/src/fleet/reconciler/observation/role_settlement/mod.rs new file mode 100644 index 00000000..47881de8 --- /dev/null +++ b/crates/cellule-host/src/fleet/reconciler/observation/role_settlement/mod.rs @@ -0,0 +1,218 @@ +use super::*; +use cellule_runtime::fleet::operations::{ + FleetAction, FleetActionKind, MaintenanceAction, MaintenanceOperation, MaintenancePhase, +}; + +#[cfg(test)] +mod tests; + +/// Opaque proof that complete role coverage and replacement policy matched one +/// fresh full journal barrier before a SettleRoles action was issued. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct FleetRoleSettlement { + scope: FleetScope, + action_key: Digest, + operation: MaintenanceOperation, + snapshot: crate::fleet::FleetJournalSnapshot, + inventory: Digest, + failed_boot_closure: Option, + capture_interval: (i64, i64), +} + +impl FleetRoleSettlement { + /// Exact inventory digest retained by the durable SettleRoles receipt. + #[must_use] + pub const fn inventory(&self) -> Digest { + self.inventory + } + + /// Exact process-closure proof for a maintenance node that has already + /// stopped. `None` means settlement must be performed by its live session. + #[must_use] + pub const fn failed_boot_closure(&self) -> Option { + self.failed_boot_closure + } + + /// Exact current journal head revision covered by this proof. + #[must_use] + pub fn head_revision(&self) -> u64 { + self.snapshot.head().revision() + } + + /// Exact current registry version covered by this proof. + #[must_use] + pub const fn registry(&self) -> RegistryVersion { + self.snapshot.registry() + } + + /// Original full native/foreign collection interval. + #[must_use] + pub const fn capture_interval(&self) -> (i64, i64) { + self.capture_interval + } + + /// Checks that the certificate belongs to this exact SettleRoles envelope. + pub fn validate_for(&self, action: &FleetAction) -> Result<()> { + let operation = match action.kind() { + FleetActionKind::Maintenance { + action: MaintenanceAction::SettleRoles, + operation, + } => operation, + _ => return Err(Error::Fenced), + }; + if action.scope() != self.scope + || action.key().map_err(crate::fleet::operation)? != self.action_key + || action.journal_revision() != self.snapshot.head().revision() + || self.snapshot.head().scope() != self.scope + || self.snapshot.head().maintenance() != Some(operation.as_ref()) + || &self.operation != operation.as_ref() + || operation.phase() != MaintenancePhase::Evacuating + || self.inventory.as_bytes().iter().all(|byte| *byte == 0) + || self + .failed_boot_closure + .is_some_and(|digest| digest.as_bytes().iter().all(|byte| *byte == 0)) + || self.capture_interval.0 < 0 + || self.capture_interval.1 < self.capture_interval.0 + || action.issued_at_ms() < self.capture_interval.1 + || action.issued_at_ms() - self.capture_interval.0 > 30_000 + { + return Err(Error::Fenced); + } + Ok(()) + } + + /// Checks that this settlement and a fresh closed-boot proof cover the same + /// exact node, session, full journal barrier, and process-closure digest. + pub fn validate_failed_boot_closure( + &self, + action: &FleetAction, + closure: &FleetFailedBootClosure, + ) -> Result<()> { + self.validate_for(action)?; + let operation = &self.operation; + let target = closure.boot().spec().target; + if self.failed_boot_closure != Some(closure.digest()) + || closure.snapshot() != &self.snapshot + || closure.boot().status() + != cellule_runtime::fleet::operations::EnrollmentStatus::Retired + || target.node != operation.node() + || target.session != operation.session() + || closure.canonical().node() != operation.node() + || closure.canonical().session() != operation.session() + { + return Err(Error::Fenced); + } + Ok(()) + } +} + +impl FleetObservation { + /// Creates the exact host proof accepted by SettleRoles only when the + /// complete inventory, original obligations, native role graph, and fresh + /// reader/follower replacement policies all cover the same snapshot. + pub fn role_settlement( + &self, + action: &FleetAction, + now_ms: i64, + ) -> Result { + let operation = match action.kind() { + FleetActionKind::Maintenance { + action: MaintenanceAction::SettleRoles, + operation, + } => operation, + _ => return Err(Error::Fenced), + }; + let roster = self.roster.as_ref().ok_or(Error::Control( + "role settlement requires the complete roster", + ))?; + let role_coverage = self.role_coverage.as_ref().ok_or(Error::Control( + "role settlement requires full role coverage", + ))?; + let original = self.maintenance_enrollments.as_ref().ok_or(Error::Control( + "role settlement requires original maintenance enrollments", + ))?; + let policies = self.maintenance_policies.as_ref().ok_or(Error::Control( + "role settlement requires complete replacement policy coverage", + ))?; + let placements = self.placements(now_ms)?; + let snapshot = roster.snapshot(); + let roster_digest = roster.digest()?; + let original_operation = original.original().operation(); + let covered_interval = |interval: (i64, i64)| { + interval.0 >= self.capture_started_at_ms + && interval.1 <= self.capture_finished_at_ms + && interval.1 >= interval.0 + }; + if !self.complete + || !self.counts_match(&placements) + || !roster.covers_advertisements(&self.nodes, now_ms)? + || roster.enrollments().iter().any(|row| { + row.status() == cellule_runtime::fleet::operations::EnrollmentStatus::Pending + }) + || self + .cells + .iter() + .any(|owned| owned.node == operation.node()) + || role_coverage.snapshot() != snapshot + || role_coverage.roster_digest() != roster_digest + || role_coverage.pending_enrollments() != 0 + || !covered_interval(role_coverage.interval()) + || original.snapshot() != snapshot + || original.roster_digest() != roster_digest + || !covered_interval(original.interval()) + || original_operation.id() != operation.id() + || original_operation.request_digest() != operation.request_digest() + || original_operation.node() != operation.node() + || original_operation.created_at_ms() != operation.created_at_ms() + || policies.snapshot() != snapshot + || policies.roster_digest() != roster_digest + || !policies.is_complete() + || !covered_interval(policies.interval()) + || operation.phase() != MaintenancePhase::Evacuating + || action.scope() != self.scope + || action.journal_revision() != snapshot.head().revision() + || self.registry != snapshot.registry() + { + return Err(Error::Control( + "role settlement evidence does not cover the current maintenance barrier", + )); + } + action + .authorize_against(snapshot.head(), now_ms) + .map_err(crate::fleet::operation)?; + let inventory = self.digest(now_ms)?; + let failed_boot_closure = self.failed_boot_closures.as_ref().and_then(|closures| { + closures.iter().find(|closure| { + let target = closure.boot().spec().target; + target.node == operation.node() && target.session == operation.session() + }) + }); + let target_is_live = self + .nodes + .iter() + .any(|node| node.node() == operation.node() && node.session() == operation.session()); + if target_is_live == failed_boot_closure.is_some() { + return Err(Error::Control( + "role settlement target must have exactly one live boot or process closure", + )); + } + if let Some(closure) = failed_boot_closure + && (closure.snapshot() != snapshot + || closure.boot().status() + != cellule_runtime::fleet::operations::EnrollmentStatus::Retired) + { + return Err(Error::Fenced); + } + let proof = FleetRoleSettlement { + scope: self.scope, + action_key: action.key().map_err(crate::fleet::operation)?, + operation: operation.as_ref().clone(), + snapshot: snapshot.clone(), + inventory, + failed_boot_closure: failed_boot_closure.map(FleetFailedBootClosure::digest), + capture_interval: (self.capture_started_at_ms, self.capture_finished_at_ms), + }; + proof.validate_for(action)?; + Ok(proof) + } +} diff --git a/crates/cellule-host/src/fleet/reconciler/observation/role_settlement/tests.rs b/crates/cellule-host/src/fleet/reconciler/observation/role_settlement/tests.rs new file mode 100644 index 00000000..fc2603e3 --- /dev/null +++ b/crates/cellule-host/src/fleet/reconciler/observation/role_settlement/tests.rs @@ -0,0 +1,107 @@ +use super::*; +use cellule_runtime::fleet::operations::{ + FleetHead, FleetProfile, FleetScope, JournalTransition, MaintenanceEvent, OperationId, +}; +use cellule_runtime::identity::{ApplicationId, Digest, NodeId, SessionId}; + +fn scope() -> FleetScope { + FleetScope { + fleet: Digest::from_bytes([1; 32]), + application: ApplicationId::from_bytes([2; 16]), + } +} + +fn next(head: &FleetHead, transition: JournalTransition, now_ms: i64) -> FleetHead { + head.transition( + FleetProfile::default(), + head.revision(), + head.controller().unwrap().epoch, + now_ms, + transition, + ) + .unwrap() +} + +fn proof() -> (FleetRoleSettlement, FleetAction, FleetHead) { + let scope = scope(); + let head = FleetHead::new(scope, 0) + .unwrap() + .claim( + FleetProfile::default(), + 0, + SessionId::from_bytes([3; 16]), + 0, + ) + .unwrap(); + let operation = MaintenanceOperation::new( + OperationId::from_bytes([4; 16]).unwrap(), + Digest::from_bytes([5; 32]), + NodeId::from_bytes([6; 16]), + SessionId::from_bytes([7; 16]), + 1, + 0, + 10_000, + ) + .unwrap(); + let head = next(&head, JournalTransition::BeginMaintenance(operation), 0); + let head = next( + &head, + JournalTransition::Maintenance(MaintenanceEvent::Cordoned), + 0, + ); + let head = next( + &head, + JournalTransition::Maintenance(MaintenanceEvent::BeginEvacuation), + 0, + ); + let action = head + .maintenance_action(MaintenanceAction::SettleRoles, 10) + .unwrap(); + let registry = RegistryVersion::new(scope).unwrap().bootstrap(0).unwrap(); + let snapshot = crate::fleet::FleetJournalSnapshot::new(head.clone(), registry).unwrap(); + let operation = snapshot.head().maintenance().unwrap().clone(); + let settlement = FleetRoleSettlement { + scope, + action_key: action.key().unwrap(), + operation, + snapshot, + inventory: Digest::from_bytes([8; 32]), + failed_boot_closure: None, + capture_interval: (0, 10), + }; + (settlement, action, head) +} + +#[test] +fn settlement_binds_exact_settle_action_and_snapshot_head() { + let (settlement, action, _) = proof(); + settlement.validate_for(&action).unwrap(); + assert_eq!(settlement.failed_boot_closure(), None); + assert_eq!(settlement.head_revision(), action.journal_revision()); +} + +#[test] +fn settlement_rejects_an_invalid_closed_boot_marker() { + let (mut settlement, action, _) = proof(); + settlement.failed_boot_closure = Some(Digest::from_bytes([0; 32])); + assert!(settlement.validate_for(&action).is_err()); +} + +#[test] +fn settlement_cannot_be_replayed_after_head_changes_with_same_registry() { + let (settlement, action, head) = proof(); + let next_head = head + .claim( + FleetProfile::default(), + head.revision(), + SessionId::from_bytes([3; 16]), + 11, + ) + .unwrap(); + let replay = next_head + .maintenance_action(MaintenanceAction::SettleRoles, 11) + .unwrap(); + assert_eq!(action.key().unwrap(), replay.key().unwrap()); + assert_ne!(action.journal_revision(), replay.journal_revision()); + assert!(settlement.validate_for(&replay).is_err()); +} diff --git a/crates/cellule-host/src/fleet/reconciler/observation/tests.rs b/crates/cellule-host/src/fleet/reconciler/observation/tests.rs index 15c34bbe..b5b60a88 100644 --- a/crates/cellule-host/src/fleet/reconciler/observation/tests.rs +++ b/crates/cellule-host/src/fleet/reconciler/observation/tests.rs @@ -77,6 +77,10 @@ fn observation() -> FleetObservation { .unwrap(), generation: 1, incarnation: IncarnationId::from_bytes([11; 16]), + owner_fence: cellule_runtime::control::OwnerFence { + incarnation: IncarnationId::from_bytes([11; 16]), + epoch: 1, + }, code: Digest::from_bytes([12; 32]), schema: 1, role: CatalogRole::Sql, @@ -174,3 +178,21 @@ fn busy_envelope_cannot_refresh_an_expired_collection_barrier() { assert!(observation.digest(30_101).is_err()); assert!(observation.placements(30_101).is_err()); } + +#[test] +fn owner_fence_is_bound_and_invalid_or_conflicting_identity_refuses_planning() { + let baseline = observation().digest(100).unwrap(); + let mut changed = observation(); + changed.cells[0].observation.owner_fence.epoch += 1; + assert_ne!(changed.digest(100).unwrap(), baseline); + for incarnation in [true, false] { + let mut invalid = observation(); + let fence = &mut invalid.cells[0].observation.owner_fence; + if incarnation { + fence.incarnation = IncarnationId::from_bytes([99; 16]); + } else { + fence.epoch = 0; + } + assert!(invalid.digest(100).is_err()); + } +} diff --git a/crates/cellule-host/src/fleet/reconciler/planning.rs b/crates/cellule-host/src/fleet/reconciler/planning.rs index d23858ad..97b019dd 100644 --- a/crates/cellule-host/src/fleet/reconciler/planning.rs +++ b/crates/cellule-host/src/fleet/reconciler/planning.rs @@ -9,6 +9,7 @@ impl FleetReconciler { &self, clock: &PassClock<'_>, report: &mut FleetReconcileReport, + pass_observation: &mut Option, ) -> Result<()> { if report.snapshot.head().attempts().len() >= self.profile.max_inflight { report.blocked(DrainBlocker::MovementBudget); @@ -18,30 +19,46 @@ impl FleetReconciler { report.blocked(DrainBlocker::IncompleteObservation); return Ok(()); } - let roster = crate::fleet::FleetRoster::collect( - self.journal.as_ref(), - &report.snapshot, - clock.deadline, - ) - .await?; - let observation = call( - clock.deadline, - "fleet-observer", - self.observer.observe(&roster, clock.now()?, clock.deadline), - ) - .await?; - if observation.scope != self.scope || observation.registry != report.snapshot.registry() { - return Err(operation(OperationError::Conflict)); - } - roster - .confirm(self.journal.as_ref(), clock.deadline) + let cached = pass_observation.take().filter(|observation| { + observation.scope == self.scope + && observation.registry == report.snapshot.registry() + && observation.roster().is_some_and(|roster| { + roster.snapshot().head().revision() == report.snapshot.head().revision() + && roster.snapshot().registry() == report.snapshot.registry() + }) + }); + let observation = if let Some(observation) = cached { + observation + } else { + let roster = crate::fleet::FleetRoster::collect( + self.journal.as_ref(), + &report.snapshot, + clock.deadline, + ) .await?; + let observation = call( + clock.deadline, + "fleet-observer", + self.observer.observe(&roster, clock.now()?, clock.deadline), + ) + .await?; + if observation.scope != self.scope || observation.registry != report.snapshot.registry() + { + return Err(operation(OperationError::Conflict)); + } + roster + .confirm(self.journal.as_ref(), clock.deadline) + .await?; + observation.with_roster(roster)? + }; + let roster = observation + .roster() + .ok_or(Error::Control("fleet roster was not retained"))?; let now = clock.now()?; let boot_complete = roster.covers_advertisements(&observation.nodes, now)?; let enrollment_settled = roster.enrollments().iter().all(|record| { record.status() != cellule_runtime::fleet::operations::EnrollmentStatus::Pending }); - let observation = observation.with_roster(roster)?; let mut placements = observation.placements(now)?; report.maintenance_policy = observation .maintenance_policy_coverage() @@ -70,11 +87,7 @@ impl FleetReconciler { report.blocked(DrainBlocker::StaleObservation); } // Retained intents take precedence over cached signed advertisements. - for intent in observation - .roster() - .ok_or(Error::Control("fleet roster was not retained"))? - .intents() - { + for intent in roster.intents() { if let Some(node) = placements .iter_mut() .find(|node| node.node == intent.node()) @@ -141,9 +154,11 @@ impl FleetReconciler { report.blocked(*blocker); continue; } - if row.position.as_ref().is_none_or(|position| { - position.incarnation != row.incarnation || position.epoch == 0 - }) || (!maintenance && !settled) + if !maintenance + && (!settled + || row.position.as_ref().is_none_or(|position| { + position.incarnation != row.incarnation || position.epoch == 0 + })) { continue; } @@ -227,10 +242,6 @@ impl FleetReconciler { .iter() .find(|node| node.session == proposal.destination) .ok_or(Error::Node("fleet proposal receiver absent"))?; - let position = row - .position - .as_ref() - .ok_or(Error::Node("fleet proposal position absent"))?; let maintenance = maintenance_for(head, owned, now); let cost = if maintenance.is_some() { row.maintenance_cost @@ -249,7 +260,7 @@ impl FleetReconciler { source_node: owned.node, source: owned.session, generation: row.generation, - source_epoch: position.epoch, + source_epoch: row.owner_fence.epoch, destination_node: destination.node, destination: destination.session, cost, diff --git a/crates/cellule-host/src/fleet/references/batch.rs b/crates/cellule-host/src/fleet/references/batch.rs new file mode 100644 index 00000000..73e42980 --- /dev/null +++ b/crates/cellule-host/src/fleet/references/batch.rs @@ -0,0 +1,138 @@ +//! Shared directory reads retain each physical follower's full traversal proof. +use super::*; +use std::collections::HashSet; +use tokio::time::timeout_at; + +const PAGE_ENTRIES: usize = 128; + +impl FleetFollowerReferences { + /// Collects up to 128 distinct physical followers in request order. + /// Every continuation traverses fresh authenticated canonical records; + /// total page rows across the active windows stay bounded by 128. Complete + /// per-member inventories retain at most 10,000 rows each, accounted by the + /// application exactly as for individual collection. Any failure returns no + /// partial set. This does not confirm the journal or native role barrier. + pub async fn collect_all( + directory: &NodeDirectory, + roster: &FleetRoster, + members: &[NodeId], + page_limit: usize, + deadline: Instant, + mut clock: impl FnMut() -> Result, + ) -> Result> { + let mut unique = HashSet::new(); + if members.is_empty() + || members.len() > PAGE_ENTRIES + || !(1..=PAGE_ENTRIES).contains(&page_limit) + || roster.snapshot().registry().bootstrap_revision().is_none() + || directory.fleet() != roster.snapshot().head().scope().fleet + || members.iter().any(|member| { + !unique.insert(*member) + || !roster + .intents() + .iter() + .any(|intent| intent.node() == *member) + }) + { + return Err(Error::Fenced); + } + let mut scans = (0..members.len()) + .map(|_| traversal::Scan::default()) + .collect::>(); + let mut pending = (0..members.len()).collect::>(); + while !pending.is_empty() { + if Instant::now() >= deadline { + return Err(Error::Deadline); + } + let requests = pending + .iter() + .map(|index| (members[*index], scans[*index].next)) + .collect::>(); + let limit = page_limit.min(PAGE_ENTRIES / pending.len()); + let started = clock()?; + let pages = timeout_at( + deadline, + directory.follower_logs_pages(&requests, limit, started), + ) + .await + .map_err(|source| Error::Facility { + name: "fleet-log-inventory-deadline", + source: Box::new(source), + })??; + let finished = clock()?; + if pages.len() != pending.len() { + return Err(Error::Node("authoritative follower batch differs")); + } + let mut next = Vec::new(); + for (index, page) in pending.into_iter().zip(pages) { + if !scans[index].accept(members[index], page, started, finished)? { + next.push(index); + } + } + pending = next; + } + let digest = roster.digest()?; + scans + .into_iter() + .zip(members) + .map(|(scan, member)| { + Ok(Self { + member: *member, + roster: digest, + snapshot: roster.snapshot().clone(), + topology: scan.topology.ok_or(Error::Fenced)?, + started_at_ms: scan.started_at_ms.ok_or(Error::Fenced)?, + finished_at_ms: scan.finished_at_ms, + collected_at_ms: scan.finished_at_ms, + rechecked: None, + entries: scan.entries, + }) + }) + .collect() + } + + /// Rechecks the whole set after native collection using fresh shared reads. + /// Starting this call invalidates every previous recheck immediately. All + /// members must match their original exact rows/topology, full roster and + /// capture interval before any confirmation advances. A changed member, + /// deadline, cancellation or source error leaves all members unconfirmed + /// with their original rows/times intact; retry performs a fresh traversal. + pub async fn recheck_all( + original: &mut [Self], + directory: &NodeDirectory, + roster: &FleetRoster, + page_limit: usize, + deadline: Instant, + clock: impl FnMut() -> Result, + ) -> Result<()> { + for references in original.iter_mut() { + references.rechecked = None; + } + let digest = roster.digest()?; + if original.iter().any(|references| { + roster.snapshot() != &references.snapshot || digest != references.roster + }) { + return Err(Error::Fenced); + } + let members = original + .iter() + .map(|references| references.member) + .collect::>(); + let fresh = + Self::collect_all(directory, roster, &members, page_limit, deadline, clock).await?; + for (original, fresh) in original.iter().zip(&fresh) { + if fresh.started_at_ms < original.finished_at_ms + || fresh.finished_at_ms - original.started_at_ms > 30_000 + || fresh.topology != original.topology + || fresh.entries != original.entries + { + return Err(Error::Node("authoritative follower inventory changed")); + } + } + for (original, fresh) in original.iter_mut().zip(fresh) { + original.finished_at_ms = fresh.finished_at_ms; + original.rechecked = Some((fresh.started_at_ms, fresh.finished_at_ms)); + } + Ok(()) + } +} diff --git a/crates/cellule-host/src/fleet/references/mod.rs b/crates/cellule-host/src/fleet/references/mod.rs index 3934f055..7086a4f0 100644 --- a/crates/cellule-host/src/fleet/references/mod.rs +++ b/crates/cellule-host/src/fleet/references/mod.rs @@ -6,8 +6,9 @@ use cellule_runtime::{ identity::{Digest, NodeId}, node::{FollowerLogObservation, NodeDirectory, log_state::NodeLogPhase}, }; -use tokio::time::{Instant, timeout_at}; +use tokio::time::Instant; +mod batch; mod traversal; /// All authoritative log references to one physical follower, including failed @@ -43,47 +44,18 @@ impl FleetFollowerReferences { deadline: Instant, mut clock: impl FnMut() -> Result, ) -> Result { - if roster.snapshot().registry().bootstrap_revision().is_none() - || directory.fleet() != roster.snapshot().head().scope().fleet - || !roster - .intents() - .iter() - .any(|intent| intent.node() == member) - || !(1..=128).contains(&page_limit) - { - return Err(Error::Fenced); - } - let mut scan = traversal::Scan::default(); - loop { - if Instant::now() >= deadline { - return Err(Error::Deadline); - } - let now = clock()?; - let page = timeout_at( - deadline, - directory.follower_logs_page(member, scan.next, page_limit, now), - ) - .await - .map_err(|source| Error::Facility { - name: "fleet-log-inventory-deadline", - source: Box::new(source), - })??; - let done = scan.accept(member, page, now, clock()?)?; - if done { - break; - } - } - Ok(Self { - member, - roster: roster.digest()?, - snapshot: roster.snapshot().clone(), - topology: scan.topology.ok_or(Error::Fenced)?, - started_at_ms: scan.started_at_ms.ok_or(Error::Fenced)?, - finished_at_ms: scan.finished_at_ms, - collected_at_ms: scan.finished_at_ms, - rechecked: None, - entries: scan.entries, - }) + Self::collect_all( + directory, + roster, + &[member], + page_limit, + deadline, + &mut clock, + ) + .await? + .into_iter() + .next() + .ok_or(Error::Fenced) } /// Physical follower node; no current boot substitutes an earlier lane. @@ -180,22 +152,15 @@ impl FleetFollowerReferences { deadline: Instant, clock: impl FnMut() -> Result, ) -> Result<()> { - self.rechecked = None; - if roster.snapshot() != &self.snapshot || roster.digest()? != self.roster { - return Err(Error::Fenced); - } - let fresh = - Self::collect(directory, roster, self.member, page_limit, deadline, clock).await?; - if fresh.started_at_ms < self.finished_at_ms - || fresh.finished_at_ms - self.started_at_ms > 30_000 - || fresh.topology != self.topology - || fresh.entries != self.entries - { - return Err(Error::Node("authoritative follower inventory changed")); - } - self.finished_at_ms = fresh.finished_at_ms; - self.rechecked = Some((fresh.started_at_ms, fresh.finished_at_ms)); - Ok(()) + Self::recheck_all( + std::slice::from_mut(self), + directory, + roster, + page_limit, + deadline, + clock, + ) + .await } } diff --git a/crates/cellule-host/src/lib.rs b/crates/cellule-host/src/lib.rs index 8e0abdb0..96ee5889 100644 --- a/crates/cellule-host/src/lib.rs +++ b/crates/cellule-host/src/lib.rs @@ -75,6 +75,8 @@ const MAX_NODE_TASKS: usize = 256; /// Stable host-owned component name for the follower store. pub const FOLLOWER_STORE_COMPONENT: &str = "follower-store"; +/// Stable host-owned component name for the Blob artifact store. +pub const BLOB_ARTIFACT_STORE_COMPONENT: &str = "blob-artifact-store"; /// Stable host-owned component name for the node-log enrollment provider. pub const NODE_DURABILITY_PROVIDER_COMPONENT: &str = "node-durability-provider"; pub(crate) const NODE_DURABILITY_SUPERVISOR_COMPONENT: &str = "node-durability-supervisor"; diff --git a/crates/cellule-host/src/node/blob_artifacts.rs b/crates/cellule-host/src/node/blob_artifacts.rs new file mode 100644 index 00000000..b7958844 --- /dev/null +++ b/crates/cellule-host/src/node/blob_artifacts.rs @@ -0,0 +1,36 @@ +//! Original Blob I/O ownership through the existing ordered node drain. + +use super::*; +use cellule_runtime::primitives::blob::BlobArtifactStore; + +impl CellNode { + /// Retains a configured Blob store before readiness and returns its shared + /// capability for `CellClient::with_blob_artifact_store`. + /// + /// The existing drain lane closes every clone and joins original accepted + /// operations before runtime shutdown. A deadline or cancelled drain waiter + /// cannot cancel those operations. This local barrier does not establish + /// Cell-scoped upload/pin coverage, remote success or global GC authority. + pub fn install_blob_artifact_store( + &self, + store: BlobArtifactStore, + ) -> cellule_runtime::Result { + self.require_task_group()?; + if store.lifecycle_observation()?.admission_closed() { + return Err(Error::CellDraining); + } + let retained = Arc::new(store.clone()); + self.install_owned_component_with_drain(BLOB_ARTIFACT_STORE_COMPONENT, retained, { + let store = store.clone(); + move || { + let store = store.clone(); + async move { + store.close_and_join().await.map(|_| ()).map_err(|source| { + Box::new(source) as Box + }) + } + } + })?; + Ok(store) + } +} diff --git a/crates/cellule-host/src/node/drain/owner.rs b/crates/cellule-host/src/node/drain/owner.rs index af3fa385..58db4c10 100644 --- a/crates/cellule-host/src/node/drain/owner.rs +++ b/crates/cellule-host/src/node/drain/owner.rs @@ -81,6 +81,42 @@ impl DrainOwner { Ok(()) } + pub(crate) fn has_boot_withdrawal(&self) -> cellule_runtime::Result { + self.resources + .boot_withdrawal + .lock() + .map(|binding| binding.is_some()) + .map_err(|_| Error::Control("CellNode boot withdrawal lock poisoned")) + } + + /// Confirms that the canonical retained host attempt joined successfully, + /// reached Stopped, and had an exact managed-boot withdrawal to execute. + pub(crate) fn confirms_fleet_terminal(&self) -> cellule_runtime::Result { + if *self + .resources + .state + .lock() + .map_err(|_| Error::Control("CellNode lifecycle lock poisoned"))? + != NodeState::Stopped + || !self.has_boot_withdrawal()? + { + return Ok(false); + } + let current = self + .bank + .lock() + .map_err(|_| Error::Control("CellNode drain bank poisoned"))? + .current + .clone(); + let Some(current) = current else { + return Ok(false); + }; + if !current.joined.load(Ordering::Acquire) { + return Ok(false); + } + Ok(matches!(current.returned.borrow().clone(), Some(Ok(())))) + } + pub(crate) async fn drain( &self, shutdown: OwnedMutexGuard<()>, diff --git a/crates/cellule-host/src/node/finalize.rs b/crates/cellule-host/src/node/finalize.rs new file mode 100644 index 00000000..923fc7dc --- /dev/null +++ b/crates/cellule-host/src/node/finalize.rs @@ -0,0 +1,542 @@ +//! Retained fleet terminal action ownership outside the finite action bank. + +use std::sync::Arc; + +use cellule_runtime::Error; +use cellule_runtime::cell::actor::NodeByteReservation; +use cellule_runtime::fleet::operations::{ + AcceptedFleetAction, DrainEvidence, FleetAction, FleetActionKind, FleetActionOutcome, + FleetOutcome, MaintenanceAction, MaintenancePhase, OperationError, +}; +use cellule_runtime::identity::Digest; +use tokio::sync::{Mutex, watch}; +use tokio::task::JoinHandle; + +use crate::fleet::{ + FinalizeActionAdmission, FleetActionCompletion, FleetActionExecutor, operation, wall_time_ms, +}; + +use super::drain::DrainOwner; + +/// One canonical Finalize action retained per node boot. +pub(crate) struct FleetFinalizeOwner { + admission: Mutex<()>, + current: Mutex>, +} + +#[derive(Clone)] +enum FinalizeSlot { + Running(Arc), + /// Durable results are reloaded from the journal, avoiding a second retained + /// copy of the terminal result after publication succeeds. An Unknown result + /// keeps the original bounded reservation so a retry remains admissible even + /// if the runtime has already closed. + Settled { + key: Digest, + executor: Arc, + retained: Arc>>, + }, +} + +struct FinalizeAttempt { + key: Digest, + accepted: AcceptedFleetAction, + executor: Arc, + completion: watch::Receiver>>, + completed: watch::Sender>>, + task: Mutex>>, + resolution: Mutex<()>, + retained: Arc>>, +} + +enum StartOutcome { + Attempt(Arc), + Completed(Arc), +} + +impl FleetFinalizeOwner { + pub(crate) fn new() -> Self { + Self { + admission: Mutex::new(()), + current: Mutex::new(None), + } + } + + pub(crate) async fn apply( + self: &Arc, + action: FleetAction, + now_ms: i64, + executor: Option>, + drain_owner: Arc, + shutdown_lock: Arc>, + allow_new: bool, + ) -> Result, Arc> { + let key = action.key().map_err(operation).map_err(Arc::new)?; + let _admission = self.admission.lock().await; + let mut current = self.current.lock().await; + let existing = current.clone(); + + let started = match existing { + Some(FinalizeSlot::Running(attempt)) => { + if attempt.key != key { + return Err(Arc::new(operation(OperationError::Conflict))); + } + attempt + .accepted + .validate_replay(&action, attempt.accepted.node(), attempt.accepted.session()) + .map_err(operation) + .map_err(Arc::new)?; + StartOutcome::Attempt(attempt) + } + Some(FinalizeSlot::Settled { + key: settled_key, + executor, + retained, + }) => { + if settled_key != key { + return Err(Arc::new(operation(OperationError::Conflict))); + } + match executor.accept_finalize(&action, now_ms).await? { + FinalizeActionAdmission::Completed(completion) => { + retained.lock().await.take(); + StartOutcome::Completed(completion) + } + FinalizeActionAdmission::Accepted(accepted) => { + if retained.lock().await.is_none() { + return Err(Arc::new(Error::Control( + "fleet Finalize retry lost its retained reservation", + ))); + } + self.start_accepted( + &mut current, + key, + accepted, + executor, + Arc::clone(&drain_owner), + Arc::clone(&shutdown_lock), + retained, + )? + } + } + } + None => { + if !allow_new { + return Err(Arc::new(Error::CellDraining)); + } + let executor = executor.ok_or_else(|| { + Arc::new(Error::Control("fleet action executor is not installed")) + })?; + self.accept_or_start( + &mut current, + key, + action, + now_ms, + executor, + Arc::clone(&drain_owner), + Arc::clone(&shutdown_lock), + ) + .await? + } + }; + + drop(current); + drop(_admission); + match started { + StartOutcome::Completed(completion) => { + verify_stopped(&completion, &drain_owner)?; + Ok(completion) + } + StartOutcome::Attempt(attempt) => { + let completion = attempt.resolve().await?; + if completion.committed { + self.mark_settled(&attempt).await; + verify_stopped(&completion, &drain_owner)?; + } + Ok(completion) + } + } + } + + async fn accept_or_start( + self: &Arc, + current: &mut Option, + key: Digest, + action: FleetAction, + now_ms: i64, + executor: Arc, + drain_owner: Arc, + shutdown_lock: Arc>, + ) -> Result> { + if !drain_owner.has_boot_withdrawal().map_err(Arc::new)? { + return Err(Arc::new(Error::Control( + "fleet Finalize requires a bound managed-boot withdrawal", + ))); + } + // Reserve the retained accepted/result envelope before the journal can + // commit acceptance. The fixed slot shares this charge across retries. + let retained = Arc::new(Mutex::new(Some( + executor.reserve_finalize_retention().map_err(Arc::new)?, + ))); + match executor.accept_finalize(&action, now_ms).await? { + FinalizeActionAdmission::Completed(completion) => { + retained.lock().await.take(); + *current = Some(FinalizeSlot::Settled { + key, + executor, + retained, + }); + Ok(StartOutcome::Completed(completion)) + } + FinalizeActionAdmission::Accepted(accepted) => self.start_accepted( + current, + key, + accepted, + executor, + drain_owner, + shutdown_lock, + retained, + ), + } + } + + fn start_accepted( + self: &Arc, + current: &mut Option, + key: Digest, + accepted: AcceptedFleetAction, + executor: Arc, + drain_owner: Arc, + shutdown_lock: Arc>, + retained: Arc>>, + ) -> Result> { + let attempt = Self::spawn( + self, + key, + accepted, + executor, + drain_owner, + shutdown_lock, + retained, + )?; + *current = Some(FinalizeSlot::Running(Arc::clone(&attempt))); + Ok(StartOutcome::Attempt(attempt)) + } + + fn spawn( + owner: &Arc, + key: Digest, + accepted: AcceptedFleetAction, + executor: Arc, + drain_owner: Arc, + shutdown_lock: Arc>, + retained: Arc>>, + ) -> Result, Arc> { + let (completed, completion) = watch::channel(None); + let attempt = Arc::new(FinalizeAttempt { + key, + accepted: accepted.clone(), + executor: Arc::clone(&executor), + completion, + completed: completed.clone(), + task: Mutex::new(None), + resolution: Mutex::new(()), + retained: Arc::clone(&retained), + }); + let owner = Arc::downgrade(owner); + let weak_attempt = Arc::downgrade(&attempt); + let settled_executor = Arc::clone(&executor); + let task = tokio::spawn(async move { + let completion = + run_finalize(key, accepted, executor, drain_owner, shutdown_lock).await; + completed.send_replace(Some(Arc::clone(&completion))); + if completion.committed { + if !matches!(completion.outcome.outcome, FleetOutcome::Unknown) { + retained.lock().await.take(); + } + if let (Some(owner), Some(attempt)) = (owner.upgrade(), weak_attempt.upgrade()) { + owner + .mark_settled_with_executor( + &attempt, + settled_executor, + Arc::clone(&retained), + ) + .await; + } + } + }); + match attempt.task.try_lock() { + Ok(mut slot) => *slot = Some(task), + Err(_) => { + task.abort(); + return Err(Arc::new(Error::Control( + "fleet Finalize task slot was unexpectedly busy", + ))); + } + } + Ok(attempt) + } + + async fn mark_settled(&self, attempt: &Arc) { + self.mark_settled_with_executor( + attempt, + Arc::clone(&attempt.executor), + Arc::clone(&attempt.retained), + ) + .await; + } + + async fn mark_settled_with_executor( + &self, + attempt: &Arc, + executor: Arc, + retained: Arc>>, + ) { + let mut current = self.current.lock().await; + if let Some(FinalizeSlot::Running(active)) = current.as_ref() + && Arc::ptr_eq(active, attempt) + { + *current = Some(FinalizeSlot::Settled { + key: attempt.key, + executor, + retained, + }); + } + } +} + +impl FinalizeAttempt { + async fn resolve(self: &Arc) -> Result, Arc> { + let _resolution = self.resolution.lock().await; + let mut completion = self.wait().await?; + if !completion.committed { + completion = Arc::new( + self.executor + .publish_terminal( + completion.accepted.clone(), + completion.outcome.clone(), + completion.execution_error.clone(), + ) + .await, + ); + self.completed.send_replace(Some(Arc::clone(&completion))); + } + if completion.committed && !matches!(completion.outcome.outcome, FleetOutcome::Unknown) { + self.retained.lock().await.take(); + } + Ok(completion) + } + + async fn wait(self: &Arc) -> Result, Arc> { + let mut completion = self.completion.clone(); + loop { + let returned = { completion.borrow().clone() }; + if let Some(result) = returned { + // The result was constructed and published before the task + // signalled. A task epilogue failure cannot invalidate it. + let _ = self.join().await; + return Ok(result); + } + if completion.changed().await.is_err() { + let error = match self.join().await { + Err(error) => error, + Ok(()) => Arc::new(Error::Control( + "fleet Finalize task ended without publishing its result", + )), + }; + let returned = { completion.borrow().clone() }; + if let Some(result) = returned { + return Ok(result); + } + let result = + unknown_completion(self.key, &self.accepted, Arc::clone(&self.executor), error) + .await; + self.completed.send_replace(Some(Arc::clone(&result))); + return Ok(result); + } + } + } + + async fn join(&self) -> Result<(), Arc> { + let mut slot = self.task.lock().await; + if let Some(task) = slot.as_mut() { + let result = task.await; + *slot = None; + if let Err(source) = result { + return Err(Arc::new(Error::Facility { + name: "fleet-finalize-task", + source: Box::new(source), + })); + } + } + Ok(()) + } +} + +async fn run_finalize( + key: Digest, + accepted: AcceptedFleetAction, + executor: Arc, + drain_owner: Arc, + shutdown_lock: Arc>, +) -> Arc { + let base = match ready_evidence(&accepted) { + Ok(evidence) => evidence, + Err(error) => { + return unknown_completion(key, &accepted, executor, Arc::new(error)).await; + } + }; + let shutdown = Arc::clone(&shutdown_lock).lock_owned().await; + let drain = drain_owner.drain(shutdown, None).await; + let terminal = match drain { + Ok(()) => drain_owner + .confirms_fleet_terminal() + .map_err(Arc::new) + .and_then(|confirmed| { + confirmed.then_some(()).ok_or_else(|| { + Arc::new(Error::Control( + "canonical node drain returned without withdrawal proof", + )) + }) + }), + Err(error) => Err(Arc::new(error)), + }; + match terminal { + Ok(()) => { + let mut evidence = base; + evidence.facilities_closed = true; + evidence.stopped = true; + evidence.withdrawn = true; + let (result, time_error) = outcome(&accepted, key, FleetOutcome::Stopped(evidence)); + match accepted.validate_result(&result) { + Ok(()) => { + let execution_error = time_error.map(Arc::new); + Arc::new( + executor + .publish_terminal(accepted, result, execution_error) + .await, + ) + } + Err(error) => { + unknown_completion(key, &accepted, executor, Arc::new(operation(error))).await + } + } + } + Err(error) => unknown_completion(key, &accepted, executor, error).await, + } +} + +async fn unknown_completion( + key: Digest, + accepted: &AcceptedFleetAction, + executor: Arc, + error: Arc, +) -> Arc { + let (result, time_error) = outcome(accepted, key, FleetOutcome::Unknown); + let execution_error = Some(match time_error { + None => error, + Some(clock) => Arc::new(Error::Facility { + name: "fleet-finalize-clock", + source: Box::new(FinalizeFailure { + primary: error, + clock, + }), + }), + }); + Arc::new( + executor + .publish_terminal(accepted.clone(), result, execution_error) + .await, + ) +} + +fn outcome( + accepted: &AcceptedFleetAction, + key: Digest, + outcome: FleetOutcome, +) -> (FleetActionOutcome, Option) { + let (observed_at_ms, time_error) = match wall_time_ms() { + Ok(now) if now >= accepted.accepted_at_ms() => (now, None), + Ok(_) => ( + accepted.accepted_at_ms(), + Some(Error::Control("fleet Finalize clock regressed")), + ), + Err(error) => (accepted.accepted_at_ms(), Some(error)), + }; + ( + FleetActionOutcome { + scope: accepted.action().scope(), + action_key: key, + node: accepted.node(), + session: accepted.session(), + observed_at_ms, + outcome, + }, + time_error, + ) +} + +fn ready_evidence(accepted: &AcceptedFleetAction) -> cellule_runtime::Result { + let FleetActionKind::Maintenance { + action: MaintenanceAction::Finalize, + operation, + } = accepted.action().kind() + else { + return Err(Error::Control("accepted action is not fleet Finalize")); + }; + let evidence = operation + .drain_evidence() + .ok_or(Error::Control("fleet Finalize lacks drain evidence"))?; + if operation.phase() != MaintenancePhase::Closing + || evidence.node != accepted.node() + || evidence.session != accepted.session() + || evidence.remaining_cells != 0 + || evidence.unresolved_attempts != 0 + || !evidence.relocated + || !evidence.readers_settled + || !evidence.followers_settled + || evidence.facilities_closed + || evidence.stopped + || evidence.withdrawn + { + return Err(Error::Control( + "fleet Finalize evidence is not ready to close", + )); + } + Ok(evidence) +} + +fn verify_stopped( + completion: &FleetActionCompletion, + drain_owner: &DrainOwner, +) -> Result<(), Arc> { + if matches!(completion.outcome.outcome, FleetOutcome::Stopped(_)) + && !drain_owner.confirms_fleet_terminal().map_err(Arc::new)? + { + return Err(Arc::new(Error::Control( + "committed fleet Finalize result lacks local terminal drain proof", + ))); + } + Ok(()) +} + +#[derive(Debug)] +struct FinalizeFailure { + primary: Arc, + clock: Error, +} + +impl std::fmt::Display for FinalizeFailure { + fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + write!( + formatter, + "{}; result clock failed: {}", + self.primary, self.clock + ) + } +} + +impl std::error::Error for FinalizeFailure { + fn source(&self) -> Option<&(dyn std::error::Error + 'static)> { + Some(self.primary.as_ref()) + } +} diff --git a/crates/cellule-host/src/node/fleet.rs b/crates/cellule-host/src/node/fleet.rs index 2fbae00e..fb250186 100644 --- a/crates/cellule-host/src/node/fleet.rs +++ b/crates/cellule-host/src/node/fleet.rs @@ -3,7 +3,9 @@ use crate::fleet::{ FLEET_ACTION_COMPONENT, FleetActionCompletion, FleetActionExecutor, FleetActionJournal, FleetCellProvider, FleetEnrollmentJournal, }; -use cellule_runtime::fleet::operations::{FleetAction, FleetScope}; +use cellule_runtime::fleet::operations::{ + FleetAction, FleetActionKind, FleetScope, MaintenanceAction, +}; use cellule_runtime::identity::NodeId; impl CellNode { @@ -446,9 +448,10 @@ impl CellNode { /// Applications authenticate the caller before invoking this local boundary. /// The executor supports settled movement, explicit busy maintenance /// release, receiver inspection, and cordon through the shared role gate. - /// Role settlement and - /// finalization require their host barriers and are refused. Count a - /// result only when `committed` is true, then inspect current serving evidence. + /// Finalize is accepted separately and runs through the retained host drain + /// owner after ready-to-close evidence and managed boot withdrawal are bound. + /// Role settlement still requires its host barriers. Count an effect only + /// when `committed` is true, then inspect current serving evidence. /// Raw Inspect actions are refused: use `inspect_fleet_action` so a durable /// historical acknowledgement cannot masquerade as a current observation. pub async fn apply_fleet_action( @@ -456,6 +459,26 @@ impl CellNode { action: FleetAction, now_ms: i64, ) -> Result, Arc> { + if matches!( + action.kind(), + FleetActionKind::Maintenance { + action: MaintenanceAction::Finalize, + .. + } + ) { + let executor = self.owned_component::(FLEET_ACTION_COMPONENT); + return self + .fleet_finalize + .apply( + action, + now_ms, + executor, + Arc::clone(&self.drain_owner), + Arc::clone(&self.shutdown_lock), + self.is_management_ready(), + ) + .await; + } if !self.is_management_ready() { return Err(Arc::new(Error::CellDraining)); } @@ -464,4 +487,23 @@ impl CellNode { .ok_or_else(|| Arc::new(Error::Control("fleet action executor is not installed")))?; executor.apply(action, now_ms).await } + + /// Applies SettleRoles only with an opaque complete inventory/policy proof + /// produced by the host reconciler at the action's exact journal barrier. + pub async fn apply_fleet_role_settlement( + &self, + action: FleetAction, + settlement: crate::fleet::FleetRoleSettlement, + now_ms: i64, + ) -> Result, Arc> { + if !self.is_management_ready() { + return Err(Arc::new(Error::CellDraining)); + } + let executor = self + .owned_component::(FLEET_ACTION_COMPONENT) + .ok_or_else(|| Arc::new(Error::Control("fleet action executor is not installed")))?; + executor + .apply_role_settlement(action, settlement, now_ms) + .await + } } diff --git a/crates/cellule-host/src/node/mod.rs b/crates/cellule-host/src/node/mod.rs index d7df7a19..a97d2444 100644 --- a/crates/cellule-host/src/node/mod.rs +++ b/crates/cellule-host/src/node/mod.rs @@ -4,6 +4,8 @@ use super::*; use crate::builder::append_required_components; use crate::durability::DurabilitySupervisor; +mod finalize; + pub(crate) struct FleetStartup { pub(crate) intent: cellule_runtime::fleet::operations::NodeIntent, pub(crate) boot: Option, @@ -18,6 +20,7 @@ pub struct CellNode { pub(super) lease_installed: AtomicBool, pub(super) shutdown_lock: Arc>, pub(super) drain_owner: Arc, + pub(super) fleet_finalize: Arc, pub(super) facilities: Arc>>, pub(super) required_components: Arc>>, pub(super) task_group: Arc>>>, @@ -40,6 +43,7 @@ impl CellNode { Arc::clone(&facilities), Arc::clone(&task_group), )); + let fleet_finalize = Arc::new(finalize::FleetFinalizeOwner::new()); Self { application, runtime, @@ -48,6 +52,7 @@ impl CellNode { lease_installed: AtomicBool::new(false), shutdown_lock: Arc::new(tokio::sync::Mutex::new(())), drain_owner, + fleet_finalize, facilities, required_components: Arc::new(Mutex::new(required_components)), task_group, @@ -132,6 +137,7 @@ impl CellNode { } } +mod blob_artifacts; mod components; mod drain; mod fleet; diff --git a/crates/cellule-host/tests/node.rs b/crates/cellule-host/tests/node.rs index 28f81398..b2adf184 100644 --- a/crates/cellule-host/tests/node.rs +++ b/crates/cellule-host/tests/node.rs @@ -101,6 +101,7 @@ mod node { Arc::new(builder.finish().unwrap()) } + mod blob_artifacts; pub mod builder; pub mod components; pub mod durability; diff --git a/crates/cellule-host/tests/node/blob_artifacts.rs b/crates/cellule-host/tests/node/blob_artifacts.rs new file mode 100644 index 00000000..f19da418 --- /dev/null +++ b/crates/cellule-host/tests/node/blob_artifacts.rs @@ -0,0 +1,238 @@ +//! Real Blob provider work retained through the existing host drain owner. +use super::*; +use cellule_runtime::primitives::blob::BlobArtifactStore; +use cellule_store::Store; +use futures_util::{StreamExt, stream::BoxStream}; +use object_store::{ObjectStore, ObjectStoreExt, memory::InMemory, path::Path}; +use std::{collections::BTreeSet, fmt, sync::atomic::AtomicUsize}; +use tokio::sync::Notify; + +#[derive(Default, Debug)] +struct Gate { + entered: AtomicUsize, + released: AtomicBool, + changed: Notify, + resume: Notify, +} +impl Gate { + async fn wait(&self) { + tokio::time::timeout(Duration::from_secs(5), async { + loop { + let changed = self.changed.notified(); + tokio::pin!(changed); + changed.as_mut().enable(); + if self.entered.load(Ordering::Acquire) != 0 { + return; + } + changed.await; + } + }) + .await + .unwrap(); + } + fn release(&self) { + self.released.store(true, Ordering::Release); + self.resume.notify_waiters(); + } +} +#[derive(Debug)] +struct Provider { + inner: InMemory, + gate: Arc, +} +impl fmt::Display for Provider { + fn fmt(&self, out: &mut fmt::Formatter<'_>) -> fmt::Result { + out.write_str("host-blob-drain") + } +} +#[async_trait::async_trait] +impl ObjectStore for Provider { + async fn put_opts( + &self, + path: &Path, + payload: object_store::PutPayload, + options: object_store::PutOptions, + ) -> object_store::Result { + self.inner.put_opts(path, payload, options).await + } + async fn put_multipart_opts( + &self, + path: &Path, + options: object_store::PutMultipartOptions, + ) -> object_store::Result> { + self.inner.put_multipart_opts(path, options).await + } + async fn get_opts( + &self, + path: &Path, + options: object_store::GetOptions, + ) -> object_store::Result { + self.inner.get_opts(path, options).await + } + fn delete_stream( + &self, + paths: BoxStream<'static, object_store::Result>, + ) -> BoxStream<'static, object_store::Result> { + self.inner.delete_stream(paths) + } + fn list( + &self, + prefix: Option<&Path>, + ) -> BoxStream<'static, object_store::Result> { + let gate = self.gate.clone(); + Box::pin(self.inner.list(prefix).then(move |item| { + let gate = gate.clone(); + async move { + gate.entered.fetch_add(1, Ordering::AcqRel); + gate.changed.notify_waiters(); + loop { + let resume = gate.resume.notified(); + tokio::pin!(resume); + resume.as_mut().enable(); + if gate.released.load(Ordering::Acquire) { + return item; + } + resume.await; + } + } + })) + } + async fn list_with_delimiter( + &self, + prefix: Option<&Path>, + ) -> object_store::Result { + self.inner.list_with_delimiter(prefix).await + } + async fn copy_opts( + &self, + from: &Path, + to: &Path, + options: object_store::CopyOptions, + ) -> object_store::Result<()> { + self.inner.copy_opts(from, to, options).await + } +} +fn node() -> CellNode { + CellNodeBuilder::new(application()) + .with_runtime(SqlWorkerPool::new(1, 1).unwrap(), 16 << 20) + .with_replica_host(ReplicaHost::default().with_local_disk_budget(DiskBudget::new(1 << 30))) + .with_session(SessionId::from_bytes([252; 16])) + .with_required_owned_components([BLOB_ARTIFACT_STORE_COMPONENT]) + .unwrap() + .build() + .unwrap() +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn original_blob_sweep_survives_caller_and_shutdown_waiter_cancellation() { + let node = Arc::new(node()); + node.install_task_group(CancellationToken::new(), CancellationToken::new()) + .unwrap(); + let provider = Arc::new(Provider { + inner: InMemory::new(), + gate: Arc::new(Gate::default()), + }); + let path = cellule_store::global_content_path( + cellule_store::GLOBAL_PREFIX, + "blob-parts", + &"ab".repeat(32), + ); + provider + .inner + .put(&path, b"orphan".to_vec().into()) + .await + .unwrap(); + let artifacts = node + .install_blob_artifact_store(BlobArtifactStore::new(Store::new(provider.clone()))) + .unwrap(); + assert!(node.install_blob_artifact_store(artifacts.clone()).is_err()); + node.install_node_lease(NodeLeaseGuard::new(0, 60_000).unwrap()) + .unwrap(); + let unrelated = BlobArtifactStore::new(Store::new(Arc::new(InMemory::new()))); + assert!(node.install_blob_artifact_store(unrelated.clone()).is_err()); + let references = Arc::new(BTreeSet::new()); + let weak = Arc::downgrade(&references); + let caller = { + let store = artifacts.clone(); + tokio::spawn(async move { store.sweep_unreferenced(references, i64::MAX).await }) + }; + provider.gate.wait().await; + caller.abort(); + assert!(caller.await.unwrap_err().is_cancelled()); + let shutdown = { + let node = node.clone(); + tokio::spawn(async move { node.shutdown().await }) + }; + let closing = tokio::time::timeout(Duration::from_secs(5), async { + while !artifacts + .lifecycle_observation() + .unwrap() + .admission_closed() + { + tokio::task::yield_now().await; + } + }) + .await; + let before = artifacts.lifecycle_observation().unwrap(); + let stopped_early = node.state() == NodeState::Stopped; + let retained_references = weak.upgrade().is_some(); + shutdown.abort(); + let cancelled = shutdown.await.is_err_and(|error| error.is_cancelled()); + // Always release the real provider and join native drain before assertions. + provider.gate.release(); + let drained = node.shutdown().await; + assert!(closing.is_ok() && cancelled && !stopped_early && retained_references); + assert_eq!(before.accepted_jobs(), 1); + drained.unwrap(); + assert_eq!(node.state(), NodeState::Stopped); + let joined = artifacts.close_and_join().await.unwrap(); + assert!(joined.locally_joined() && joined.first_failure().is_none()); + assert!( + !unrelated + .lifecycle_observation() + .unwrap() + .admission_closed() + ); + assert!(weak.upgrade().is_none()); + assert_eq!(provider.gate.entered.load(Ordering::Acquire), 1); + assert!(matches!( + provider.inner.get(&path).await, + Err(object_store::Error::NotFound { .. }) + )); + assert!( + node.try_owned_component::(BLOB_ARTIFACT_STORE_COMPONENT) + .unwrap() + .is_none() + ); + assert_eq!(node.stats().active_cells(), 0); + assert_eq!(node.stats().worker_jobs(), 0); + assert_eq!(node.stats().retained_bytes(), 0); + assert_eq!(node.stats().local_disk_reserved_bytes(), 0); +} + +#[tokio::test] +async fn blob_installation_requires_owned_tasks_and_an_open_original_store() { + let node = node(); + let artifacts = BlobArtifactStore::new(Store::new(Arc::new(InMemory::new()))); + assert!(node.install_blob_artifact_store(artifacts.clone()).is_err()); + node.install_task_group(CancellationToken::new(), CancellationToken::new()) + .unwrap(); + artifacts.close(); + assert!(matches!( + node.install_blob_artifact_store(artifacts), + Err(Error::CellDraining) + )); + assert!( + node.try_owned_component::(BLOB_ARTIFACT_STORE_COMPONENT) + .unwrap() + .is_none() + ); + // Complete the required owner set before normal startup/cleanup. + let installed = node + .install_blob_artifact_store(BlobArtifactStore::new(Store::new( + Arc::new(InMemory::new()), + ))) + .unwrap(); + node.shutdown().await.unwrap(); + assert!(installed.close_and_join().await.unwrap().locally_joined()); +} diff --git a/crates/cellule-host/tests/node/fleet_actions.rs b/crates/cellule-host/tests/node/fleet_actions.rs index 9fae0f66..c9d6637f 100644 --- a/crates/cellule-host/tests/node/fleet_actions.rs +++ b/crates/cellule-host/tests/node/fleet_actions.rs @@ -58,6 +58,10 @@ pub(super) struct Journal { pub(super) basis_writes: AtomicUsize, pub(super) panic_basis: AtomicBool, pub(super) lose_recovery_evidence_reply: AtomicBool, + pub(super) fail_recovery_evidence_writes: AtomicUsize, + pub(super) block_recovery_evidence: AtomicBool, + pub(super) recovery_evidence_entered: tokio::sync::Notify, + pub(super) recovery_evidence_resume: tokio::sync::Semaphore, pub(super) block_inspections: AtomicBool, pub(super) inspection_entered: tokio::sync::Notify, pub(super) inspection_resume: tokio::sync::Semaphore, @@ -123,6 +127,10 @@ impl Journal { basis_writes: AtomicUsize::new(0), panic_basis: AtomicBool::new(false), lose_recovery_evidence_reply: AtomicBool::new(false), + fail_recovery_evidence_writes: AtomicUsize::new(0), + block_recovery_evidence: AtomicBool::new(false), + recovery_evidence_entered: tokio::sync::Notify::new(), + recovery_evidence_resume: tokio::sync::Semaphore::new(0), block_inspections: AtomicBool::new(false), inspection_entered: tokio::sync::Notify::new(), inspection_resume: tokio::sync::Semaphore::new(0), @@ -532,6 +540,26 @@ impl FleetActionJournal for Journal { evidence: &'a RecoveryEvidence, ) -> FleetAdapterFuture<'a, RecoveryEvidence> { Box::pin(async move { + if self.block_recovery_evidence.swap(false, Ordering::SeqCst) { + self.recovery_evidence_entered.notify_one(); + self.recovery_evidence_resume + .acquire() + .await + .unwrap() + .forget(); + } + if self + .fail_recovery_evidence_writes + .try_update(Ordering::SeqCst, Ordering::SeqCst, |remaining| { + remaining.checked_sub(1) + }) + .is_ok() + { + return Err(Box::new(std::io::Error::other( + "injected recovery evidence write failure", + )) + as Box); + } evidence.basis().validate_acceptance(accepted)?; let retained = { let mut state = self.state.lock().unwrap(); diff --git a/crates/cellule-host/tests/node/fleet_maintenance.rs b/crates/cellule-host/tests/node/fleet_maintenance.rs index 116c3ad7..689617f5 100644 --- a/crates/cellule-host/tests/node/fleet_maintenance.rs +++ b/crates/cellule-host/tests/node/fleet_maintenance.rs @@ -455,6 +455,30 @@ async fn wrong_boot_stale_intent_and_unsupported_roles_do_not_accept_effects() { .await .is_err() ); + fixture.journal.transition(JournalTransition::Maintenance( + MaintenanceEvent::ReadyToClose(DrainEvidence { + node: NodeId::from_bytes([201; 16]), + session: SessionId::from_bytes([201; 16]), + remaining_cells: 0, + unresolved_attempts: 0, + relocated: true, + readers_settled: true, + followers_settled: true, + facilities_closed: false, + stopped: false, + withdrawn: false, + }), + )); + let finalize = fixture + .journal + .maintenance_action(MaintenanceAction::Finalize); + assert!( + fixture + .node + .apply_fleet_action(finalize, clock()) + .await + .is_err() + ); let inspection = fixture .node .inspect_fleet_action(maintenance_inspection(&fixture, 213)) diff --git a/crates/cellule-host/tests/node/fleet_receivers/mod.rs b/crates/cellule-host/tests/node/fleet_receivers/mod.rs index 641f2656..f6e82df7 100644 --- a/crates/cellule-host/tests/node/fleet_receivers/mod.rs +++ b/crates/cellule-host/tests/node/fleet_receivers/mod.rs @@ -14,8 +14,10 @@ use cellule_runtime::identity::RequestId; use cellule_runtime::ltx::CellReplica; mod prefix; +mod recovery_resume; mod successor; mod suffix; +mod suffix_resume; struct Cells(FleetCellInputs, Arc>>); @@ -1726,9 +1728,16 @@ async fn lost_recovery_evidence_reply_cannot_admit_actor_or_claim_completion() { assert_eq!(current.value().epoch, evidence.restored().epoch); assert_eq!(movement.receiver.stats().active_cells(), 0); assert_eq!(movement.receiver.stats().worker_jobs(), 0); - let repeated = apply(&movement.receiver, action.clone()).await; - assert!(matches!(repeated.outcome.outcome, FleetOutcome::Unknown)); - assert!(repeated.execution_error.is_some()); + // Inspection cannot restart a safely rolled-back acquisition. Keep the + // ordinary winner independent of the accepted effect's automatic replay. + assert!( + movement + .receiver + .inspect_fleet_action(movement.inspection(248)) + .await + .is_err() + ); + assert_eq!(movement.receiver.stats().active_cells(), 0); assert_eq!( movement.source.journal.current_attempt().phase(), AttemptPhase::Recovering diff --git a/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/continuation.rs b/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/continuation.rs new file mode 100644 index 00000000..2a2bef4b --- /dev/null +++ b/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/continuation.rs @@ -0,0 +1,204 @@ +use super::*; + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn lost_failed_source_evidence_reply_resumes_with_unchanged_durable_evidence() { + let (movement, action, idle, receipt) = failed(0, 77, false).await; + let evidence = movement + .source + .journal + .recovery_evidence(action.key().unwrap()) + .unwrap(); + let basis = movement + .source + .journal + .recovery_basis(action.key().unwrap()) + .unwrap(); + assert_eq!(evidence.basis(), &basis); + assert!( + movement + .receiver + .inspect_fleet_action(movement.inspection(249)) + .await + .is_err() + ); + assert_eq!( + movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap() + .value(), + &idle + ); + // The helper compares the returned evidence to the original canonical + // materialization; capture time must also retain this committed value. + let completion = apply(&movement.receiver, action.clone()).await; + assert!( + completion.committed && completion.execution_error.is_none(), + "{completion:?}" + ); + let FleetOutcome::Recovered(result) = &completion.outcome.outcome else { + panic!("{completion:?}"); + }; + assert_eq!(result.recovery, evidence); + finish(movement, action, &idle, 77, true, receipt).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn cancelled_failed_source_repair_waiter_leaves_the_original_work_owned() { + let (movement, action, idle, receipt) = failed(1, 42, false).await; + let basis = movement + .source + .journal + .recovery_basis(action.key().unwrap()) + .unwrap(); + movement + .source + .journal + .block_recovery_evidence + .store(true, Ordering::SeqCst); + let node = movement.receiver.clone(); + let issued = action.clone(); + let waiter = tokio::spawn(async move { apply(&node, issued).await }); + tokio::time::timeout( + Duration::from_secs(5), + movement.source.journal.recovery_evidence_entered.notified(), + ) + .await + .unwrap(); + waiter.abort(); + assert!(waiter.await.unwrap_err().is_cancelled()); + assert_eq!( + movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap() + .value(), + &idle + ); + assert_eq!(movement.receiver.stats().active_cells(), 0); + let inspection = movement.inspect(250).await; + assert!(matches!( + inspection.outcome().outcome, + FleetOutcome::Unknown + )); + assert!( + movement + .source + .journal + .recovery_evidence(action.key().unwrap()) + .is_none() + ); + movement + .source + .journal + .recovery_evidence_resume + .add_permits(1); + let completion = apply(&movement.receiver, action.clone()).await; + assert!( + completion.committed && completion.execution_error.is_none(), + "{completion:?}" + ); + let FleetOutcome::Recovered(result) = &completion.outcome.outcome else { + panic!("{completion:?}"); + }; + assert_eq!(result.recovery.basis(), &basis); + assert_eq!(result.serving.position.epoch, idle.epoch + 1); + finish(movement, action, &idle, 42, true, receipt).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn failed_source_evidence_repair_resumes_idle_without_replacing_original_basis() { + let (movement, action, idle, receipt) = failed(1, 77, false).await; + // Fresh inspection cannot reconstruct metadata or start native acquisition. + let observation = movement.inspect(247).await; + assert!(matches!( + observation.outcome().outcome, + FleetOutcome::Unknown + )); + assert!( + movement + .source + .journal + .recovery_evidence(action.key().unwrap()) + .is_none() + ); + assert_eq!( + movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap() + .value(), + &idle + ); + finish(movement, action, &idle, 77, false, receipt).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn repeated_failed_source_evidence_write_failure_cannot_start_another_claim() { + let (movement, action, idle, receipt) = failed(2, 42, false).await; + let failure = apply(&movement.receiver, action.clone()).await; + assert!(failure.committed && matches!(failure.outcome.outcome, FleetOutcome::Unknown)); + let Error::Facility { source, .. } = failure.execution_error.as_ref().unwrap().as_ref() else { + panic!("{failure:?}"); + }; + assert_eq!( + source.downcast_ref::().unwrap().to_string(), + "injected recovery evidence write failure" + ); + assert_eq!( + movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap() + .value(), + &idle + ); + assert!( + movement + .source + .journal + .recovery_evidence(action.key().unwrap()) + .is_none() + ); + assert_eq!(movement.receiver.stats().active_cells(), 0); + finish(movement, action, &idle, 42, false, receipt).await; +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn failed_source_evidence_repair_accepts_an_ordinary_acquisition_winner() { + let (movement, action, idle, receipt) = failed(1, 42, false).await; + let observed = movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap(); + let handle = movement + .receiver + .runtime() + .acquire_idle_restored( + movement.inputs.catalog.clone(), + movement.inputs.replica.clone(), + movement.inputs.authority.clone(), + observed, + movement.inputs.destination.clone(), + movement.inputs.owner.clone(), + ) + .await + .unwrap(); + assert_eq!(counter(&handle).await, 42); + finish(movement, action, &idle, 42, true, receipt).await; +} diff --git a/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/history.rs b/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/history.rs new file mode 100644 index 00000000..5de945e1 --- /dev/null +++ b/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/history.rs @@ -0,0 +1,176 @@ +use super::*; + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn failed_source_missing_or_corrupt_history_blocks_idle_and_ordinary_winner() { + for (corrupt, ordinary) in [(false, false), (true, false), (false, true), (true, true)] { + let (movement, action, idle, receipt) = failed(1, 42, false).await; + let observed = movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap(); + if ordinary { + let handle = movement + .receiver + .runtime() + .acquire_idle_restored( + movement.inputs.catalog.clone(), + movement.inputs.replica.clone(), + movement.inputs.authority.clone(), + observed, + movement.inputs.destination.clone(), + movement.inputs.owner.clone(), + ) + .await + .unwrap(); + assert_eq!(counter(&handle).await, 42); + } + let selected = movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap(); + let layout = movement.inputs.authority.layout(); + let path = layout.acquisition_record_path( + idle.cell.as_bytes(), + idle.incarnation.as_bytes(), + idle.epoch, + ); + let (body, _) = layout.store().get_with_etag(&path).await.unwrap(); + layout.store().delete(&path).await.unwrap(); + if corrupt { + layout + .store() + .create_strict(&path, Bytes::from_static(b"corrupt-source-acquisition")) + .await + .unwrap(); + } + let refused = apply(&movement.receiver, action.clone()).await; + assert!(refused.committed && matches!(refused.outcome.outcome, FleetOutcome::Unknown)); + let error = refused.execution_error.as_ref().unwrap().as_ref(); + if corrupt { + assert!(matches!(error, Error::Control(_)), "{error:?}"); + } else { + assert!( + matches!(error, Error::AcquisitionHistoryIncomplete { epoch, .. } if *epoch == idle.epoch), + "{error:?}" + ); + } + assert_eq!( + movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap() + .value(), + selected.value() + ); + assert_eq!( + movement.receiver.stats().active_cells(), + usize::from(ordinary) + ); + assert!( + movement + .source + .journal + .recovery_evidence(action.key().unwrap()) + .is_none() + ); + if corrupt { + layout.store().delete(&path).await.unwrap(); + } + layout.store().create_strict(&path, body).await.unwrap(); + finish(movement, action, &idle, 42, ordinary, receipt).await; + } +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn failed_source_valid_substituted_history_cannot_reconstruct_evidence() { + let (foreign, _, foreign_idle, _) = failed(1, 99, false).await; + let layout = foreign.inputs.authority.layout(); + let path = layout.acquisition_record_path( + foreign_idle.cell.as_bytes(), + foreign_idle.incarnation.as_bytes(), + foreign_idle.epoch, + ); + let (foreign_body, _) = layout.store().get_with_etag(&path).await.unwrap(); + let foreign_input = foreign + .inputs + .authority + .acquisition_record( + foreign_idle.cell, + foreign_idle.incarnation, + foreign_idle.epoch, + ) + .await + .unwrap() + .unwrap() + .input() + .clone(); + foreign.shutdown().await; + let (movement, action, idle, receipt) = failed(1, 42, false).await; + let basis = movement + .source + .journal + .recovery_basis(action.key().unwrap()) + .unwrap(); + assert_ne!(basis.control(), &foreign_input); + let layout = movement.inputs.authority.layout(); + let path = layout.acquisition_record_path( + idle.cell.as_bytes(), + idle.incarnation.as_bytes(), + idle.epoch, + ); + let (body, _) = layout.store().get_with_etag(&path).await.unwrap(); + layout.store().delete(&path).await.unwrap(); + layout + .store() + .create_strict(&path, foreign_body) + .await + .unwrap(); + // It decodes as a valid native record in the same Cell/epoch scope. + assert_eq!( + movement + .inputs + .authority + .acquisition_record(idle.cell, idle.incarnation, idle.epoch) + .await + .unwrap() + .unwrap() + .input(), + &foreign_input + ); + let refused = apply(&movement.receiver, action.clone()).await; + assert!(refused.committed && matches!(refused.outcome.outcome, FleetOutcome::Unknown)); + assert!(matches!( + refused.execution_error.as_ref().unwrap().as_ref(), + Error::Control("recovery input differs from canonical acquisition") + )); + assert_eq!( + movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap() + .value(), + &idle + ); + assert!( + movement + .source + .journal + .recovery_evidence(action.key().unwrap()) + .is_none() + ); + layout.store().delete(&path).await.unwrap(); + layout.store().create_strict(&path, body).await.unwrap(); + finish(movement, action, &idle, 42, false, receipt).await; +} diff --git a/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/mod.rs b/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/mod.rs new file mode 100644 index 00000000..70d97e09 --- /dev/null +++ b/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/mod.rs @@ -0,0 +1,210 @@ +//! Accepted failed-source recovery must finish a proven safe rollback itself. +use super::*; +use bytes::Bytes; +use cellule_runtime::cell::executor::StoredOutcome; +use cellule_runtime::control::Control; + +struct Acknowledged { + identity: MutationIdentity, + outcome: StoredOutcome, +} + +mod continuation; +mod history; +mod origin; +mod suffix; + +async fn failed( + writes: usize, + value: i64, + overlay: bool, +) -> (Movement, FleetAction, Control, Acknowledged) { + let movement = Movement::new(128 << 20).await; + let now = clock(); + let identity = MutationIdentity { + request_id: RequestId::from_bytes([247; 16]), + issued_at_ms: now, + expires_at_ms: now + 60_000, + }; + let outcome = movement + .source + .handle + .execute( + identity, + Digest::from_bytes([247; 32]), + now, + 64, + 64, + move |tx| { + tx.execute("UPDATE counter SET value = ?1", [value])?; + Ok(HandlerOutcome::Success(value.to_be_bytes().to_vec())) + }, + ) + .await + .unwrap(); + if overlay { + super::suffix::start_suffix_recovery(&movement).await; + } else { + movement.start_recovery(false).await; + } + if writes == 0 { + movement + .source + .journal + .lose_recovery_evidence_reply + .store(true, Ordering::SeqCst); + } + movement + .source + .journal + .fail_recovery_evidence_writes + .store(writes, Ordering::SeqCst); + let action = movement.action(MovementAction::Recover); + let failure = apply(&movement.receiver, action.clone()).await; + assert!(failure.committed && matches!(failure.outcome.outcome, FleetOutcome::Unknown)); + let Error::Facility { source, .. } = failure.execution_error.as_ref().unwrap().as_ref() else { + panic!("original journal error lost: {failure:?}"); + }; + assert_eq!( + source.downcast_ref::().unwrap().to_string(), + if writes == 0 { + "injected lost recovery evidence reply" + } else { + "injected recovery evidence write failure" + } + ); + assert_eq!( + movement + .source + .journal + .recovery_evidence(action.key().unwrap()) + .is_some(), + writes == 0 + ); + let current = movement + .inputs + .authority + .load(movement.spec.target.cell_id()) + .await + .unwrap() + .unwrap(); + assert_eq!(current.value().state, ControlState::Idle); + assert!(current.value().owner.is_none()); + assert_eq!(movement.receiver.stats().active_cells(), 0); + assert_eq!(movement.receiver.stats().worker_jobs(), 0); + assert_eq!(movement.receiver.stats().file_descriptors(), 0); + assert_eq!(movement.receiver.stats().local_disk_reserved_bytes(), 0); + ( + movement, + action, + current.value().clone(), + Acknowledged { identity, outcome }, + ) +} + +async fn finish( + movement: Movement, + action: FleetAction, + idle: &Control, + value: i64, + ordinary: bool, + receipt: Acknowledged, +) { + let basis = movement + .source + .journal + .recovery_basis(action.key().unwrap()) + .unwrap(); + let canonical = movement + .inputs + .authority + .acquisition_record(idle.cell, idle.incarnation, idle.epoch) + .await + .unwrap() + .unwrap(); + assert_eq!(canonical.input(), basis.control()); + let before = movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap(); + assert_eq!( + before.value().state, + if ordinary { + ControlState::Serving + } else { + ControlState::Idle + } + ); + let completion = apply(&movement.receiver, action.clone()).await; + assert!( + completion.committed && completion.execution_error.is_none(), + "{completion:?}" + ); + let FleetOutcome::Recovered(result) = &completion.outcome.outcome else { + panic!("{completion:?}"); + }; + assert_eq!(result.recovery.basis(), &basis); + assert_eq!(result.recovery.restored(), canonical.materialized()); + assert_eq!(result.serving.position.epoch, idle.epoch + 1); + let current = movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap(); + let handle = movement + .receiver + .runtime() + .local_handle(movement.inputs.catalog.clone(), ¤t) + .await + .unwrap() + .unwrap(); + assert_eq!(counter(&handle).await, value); + assert_eq!( + handle + .resolve(receipt.identity, Digest::from_bytes([247; 32]), clock(), 64) + .await + .unwrap(), + Resolution::Committed(receipt.outcome) + ); + if ordinary { + assert_eq!(before.value(), current.value()); + } + assert_eq!(movement.receiver.stats().active_cells(), 1); + assert_eq!( + movement.source.journal.basis_writes.load(Ordering::SeqCst), + 1 + ); + assert_eq!( + movement.source.journal.current_attempt().phase(), + AttemptPhase::Recovering + ); + assert_eq!( + movement.source.journal.current_attempt().spec().cost, + movement.spec.cost + ); + let repeated = apply(&movement.receiver, action).await; + assert!(repeated.committed && repeated.execution_error.is_none()); + let FleetOutcome::Recovered(repeated) = &repeated.outcome.outcome else { + panic!("{repeated:?}"); + }; + assert_eq!(repeated.recovery, result.recovery); + assert_eq!(repeated.serving.position.epoch, idle.epoch + 1); + assert_eq!( + movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap() + .value(), + current.value() + ); + assert_eq!(current.value().epoch, canonical.materialized().epoch + 1); + movement.shutdown().await; +} diff --git a/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/origin.rs b/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/origin.rs new file mode 100644 index 00000000..7022dd5e --- /dev/null +++ b/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/origin.rs @@ -0,0 +1,48 @@ +//! Recovery metadata cannot replace the actual origin bytes before reacquisition. +use super::*; + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn failed_source_idle_reacquisition_requires_complete_materialized_origin() { + let (movement, action, idle, receipt) = failed(1, 42, false).await; + let root = idle.ltx_root().unwrap(); + let layout = movement.inputs.authority.layout(); + let path = layout.incarnation_object_path( + &root.cell, + &root.incarnation, + &root.digest, + cellule_runtime::ltx::CellObjectKind::Root, + ); + let (body, _) = layout.store().get_with_etag(&path).await.unwrap(); + layout.store().delete(&path).await.unwrap(); + let refused = apply(&movement.receiver, action.clone()).await; + assert!(refused.committed && matches!(refused.outcome.outcome, FleetOutcome::Unknown)); + assert!(matches!( + refused.execution_error.as_ref().unwrap().as_ref(), + Error::Ltx(_) + )); + assert_eq!( + movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap() + .value(), + &idle + ); + assert_eq!(movement.receiver.stats().active_cells(), 0); + assert_eq!(movement.receiver.stats().worker_jobs(), 0); + assert_eq!(movement.receiver.stats().file_descriptors(), 0); + assert_eq!(movement.receiver.stats().local_disk_reserved_bytes(), 0); + // Recording the actual native materialization is allowed; origin absence + // still blocks the next ownership CAS and any serving/settlement result. + let evidence = movement + .source + .journal + .recovery_evidence(action.key().unwrap()) + .unwrap(); + assert_eq!(evidence.restored().ltx_root(), Some(root)); + layout.store().create_strict(&path, body).await.unwrap(); + finish(movement, action, &idle, 42, false, receipt).await; +} diff --git a/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/suffix.rs b/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/suffix.rs new file mode 100644 index 00000000..11d90cca --- /dev/null +++ b/crates/cellule-host/tests/node/fleet_receivers/recovery_resume/suffix.rs @@ -0,0 +1,109 @@ +use super::*; + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn failed_source_idle_continuation_requires_its_original_sealed_suffix() { + for corrupt in [false, true] { + let (movement, action, idle, receipt) = failed(1, 42, true).await; + let basis = movement + .source + .journal + .recovery_basis(action.key().unwrap()) + .unwrap(); + let overlay = basis.control().recovery.as_ref().unwrap(); + assert!(idle.recovery.is_none()); + assert_ne!(idle.root, basis.control().root); + let layout = movement.inputs.authority.layout(); + let path = layout.node_log_recovery_path( + overlay.leader_session.as_bytes(), + overlay.log_epoch, + overlay.manifest_digest.as_bytes(), + ); + let (body, _) = layout.store().get_with_etag(&path).await.unwrap(); + layout.store().delete(&path).await.unwrap(); + if corrupt { + layout + .store() + .create_strict(&path, Bytes::from_static(b"corrupt-original-suffix")) + .await + .unwrap(); + } + let refused = apply(&movement.receiver, action.clone()).await; + assert!(refused.committed && matches!(refused.outcome.outcome, FleetOutcome::Unknown)); + assert!(refused.execution_error.is_some()); + assert_eq!( + movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap() + .value(), + &idle + ); + assert_eq!(movement.receiver.stats().active_cells(), 0); + assert_eq!(movement.receiver.stats().worker_jobs(), 0); + assert_eq!(movement.receiver.stats().file_descriptors(), 0); + // Metadata records the actual native materialization, but cannot admit + // a writer without the original suffix and complete prefix proof. + let evidence = movement + .source + .journal + .recovery_evidence(action.key().unwrap()) + .unwrap(); + assert_eq!(evidence.basis(), &basis); + assert_eq!(evidence.restored().root, idle.root); + assert!( + movement + .receiver + .inspect_fleet_action(movement.inspection(251)) + .await + .is_err() + ); + assert_eq!( + movement.source.journal.current_attempt().spec().cost, + movement.spec.cost + ); + if corrupt { + layout.store().delete(&path).await.unwrap(); + } + layout.store().create_strict(&path, body).await.unwrap(); + let completed = apply(&movement.receiver, action.clone()).await; + assert!( + completed.committed && completed.execution_error.is_none(), + "{completed:?}" + ); + let FleetOutcome::Recovered(result) = &completed.outcome.outcome else { + panic!("{completed:?}"); + }; + assert_eq!(result.recovery, evidence); + finish(movement, action, &idle, 43, true, receipt).await; + } +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn failed_source_suffix_evidence_repair_preserves_an_ordinary_writer_and_receipt() { + let (movement, action, idle, receipt) = failed(1, 42, true).await; + let current = movement + .inputs + .authority + .load(idle.cell) + .await + .unwrap() + .unwrap(); + let handle = movement + .receiver + .runtime() + .acquire_idle_restored( + movement.inputs.catalog.clone(), + movement.inputs.replica.clone(), + movement.inputs.authority.clone(), + current, + movement.inputs.destination.clone(), + movement.inputs.owner.clone(), + ) + .await + .unwrap(); + assert_eq!(counter(&handle).await, 43); + finish(movement, action, &idle, 43, true, receipt).await; +} diff --git a/crates/cellule-host/tests/node/fleet_receivers/suffix.rs b/crates/cellule-host/tests/node/fleet_receivers/suffix.rs index 9b0ff078..25c0165b 100644 --- a/crates/cellule-host/tests/node/fleet_receivers/suffix.rs +++ b/crates/cellule-host/tests/node/fleet_receivers/suffix.rs @@ -114,7 +114,7 @@ async fn inspect_suffix(corrupt: bool) { movement.shutdown().await; } -async fn start_suffix_recovery(movement: &Movement) { +pub(super) async fn start_suffix_recovery(movement: &Movement) { let prepared = movement.prepare().await; let FleetOutcome::Reserved(reservation) = prepared.outcome.outcome else { panic!("not prepared") diff --git a/crates/cellule-host/tests/node/fleet_receivers/suffix_resume/interrupted.rs b/crates/cellule-host/tests/node/fleet_receivers/suffix_resume/interrupted.rs new file mode 100644 index 00000000..43fd9b3d --- /dev/null +++ b/crates/cellule-host/tests/node/fleet_receivers/suffix_resume/interrupted.rs @@ -0,0 +1,197 @@ +use super::*; + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn interrupted_overlay_materialization_resumes_exact_claim_after_waiter_cancellation() { + let movement = Movement::new(128 << 20).await; + let now = clock(); + let identity = MutationIdentity { + request_id: RequestId::from_bytes([246; 16]), + issued_at_ms: now, + expires_at_ms: now + 60_000, + }; + let digest = Digest::from_bytes([246; 32]); + let receipt = movement + .source + .handle + .execute(identity, digest, now, 64, 64, |tx| { + tx.execute("UPDATE counter SET value=value", [])?; + Ok(HandlerOutcome::Success(vec![42])) + }) + .await + .unwrap(); + suffix::start_suffix_recovery(&movement).await; + let original = movement + .inputs + .authority + .load(movement.spec.target.cell_id()) + .await + .unwrap() + .unwrap() + .value() + .clone(); + let recovery = original.recovery.as_ref().unwrap(); + let layout = movement.inputs.authority.layout(); + let path = layout.node_log_recovery_path( + recovery.leader_session.as_bytes(), + recovery.log_epoch, + recovery.manifest_digest.as_bytes(), + ); + let (body, _) = layout.store().get_with_etag(&path).await.unwrap(); + layout.store().delete(&path).await.unwrap(); + for _ in 0..2 { + let failure = apply(&movement.receiver, movement.action(MovementAction::Recover)).await; + assert!(failure.committed && matches!(failure.outcome.outcome, FleetOutcome::Unknown)); + assert!(matches!( + failure.execution_error.as_ref().unwrap().as_ref(), + Error::Storage(cellule_store::StorageError::NotFound { .. }) + )); + let claimed = movement + .inputs + .authority + .load(original.cell) + .await + .unwrap() + .unwrap(); + assert_eq!(claimed.value().epoch, original.epoch + 1); + assert_eq!(claimed.value().state, ControlState::Recovering); + assert_eq!( + claimed.value().owner.as_ref().unwrap().session, + movement.spec.destination + ); + assert_eq!(claimed.value().recovery, original.recovery); + assert_eq!(claimed.value().root, original.root); + assert_eq!(movement.receiver.stats().active_cells(), 0); + assert_eq!(movement.receiver.stats().worker_jobs(), 0); + assert_eq!(movement.receiver.stats().file_descriptors(), 0); + assert!( + movement + .inputs + .authority + .acquisition_record(original.cell, original.incarnation, claimed.value().epoch) + .await + .unwrap() + .is_none() + ); + } + let FleetActionAcceptance::Existing { accepted, .. } = movement + .source + .journal + .load_movement_action( + scope(), + movement.spec.id, + MovementAction::Recover, + movement.spec.destination_node, + movement.spec.destination, + ) + .await + .unwrap() + .unwrap() + else { + panic!("accepted recovery missing") + }; + let basis = movement + .source + .journal + .load_recovery_basis(&accepted) + .await + .unwrap() + .unwrap(); + assert_eq!(basis.control(), &original); + assert!( + movement + .source + .journal + .load_recovery_evidence(&accepted) + .await + .unwrap() + .is_none() + ); + layout.store().create_strict(&path, body).await.unwrap(); + movement + .source + .journal + .block_basis + .store(true, std::sync::atomic::Ordering::SeqCst); + let node = movement.receiver.clone(); + let action = movement.action(MovementAction::Recover); + let waiter = tokio::spawn(async move { apply(&node, action).await }); + tokio::time::timeout( + std::time::Duration::from_secs(5), + movement.source.journal.basis_entered.notified(), + ) + .await + .unwrap(); + waiter.abort(); + assert!(waiter.await.unwrap_err().is_cancelled()); + movement.source.journal.basis_resume.add_permits(1); + let recovered = apply(&movement.receiver, movement.action(MovementAction::Recover)).await; + assert!( + recovered.committed && recovered.execution_error.is_none(), + "{recovered:?}" + ); + let FleetOutcome::Recovered(evidence) = &recovered.outcome.outcome else { + panic!("not recovered: {recovered:?}") + }; + assert_eq!(evidence.recovery.basis(), &basis); + assert_eq!(evidence.recovery.restored().epoch, original.epoch + 1); + assert!(evidence.recovery.restored().recovery.is_none()); + assert_eq!(evidence.serving.position.epoch, original.epoch + 1); + movement.event(AttemptEvent::Recovered(evidence.clone())); + let current = movement + .inputs + .authority + .load(original.cell) + .await + .unwrap() + .unwrap(); + let canonical = movement + .inputs + .authority + .acquisition_record(original.cell, original.incarnation, current.value().epoch) + .await + .unwrap() + .unwrap(); + assert_eq!(canonical.input(), &original); + assert_eq!(canonical.materialized(), evidence.recovery.restored()); + let handle = movement + .receiver + .runtime() + .local_handle(movement.inputs.catalog.clone(), ¤t) + .await + .unwrap() + .unwrap(); + assert_eq!( + handle.resolve(identity, digest, clock(), 64).await.unwrap(), + Resolution::Committed(receipt) + ); + assert_eq!(counter(&handle).await, 43); + assert_eq!(movement.receiver.stats().active_cells(), 1); + let inputs = movement.recovery_inputs.lock().unwrap().clone().unwrap(); + let manifest = inputs + .manifests + .load_manifest( + recovery.leader_session, + recovery.log_epoch, + recovery.manifest_digest, + ) + .await + .unwrap(); + assert_eq!(manifest.cells()[0].cell_epoch, original.epoch); + assert_eq!(manifest.cells()[0].recovery, *recovery); + movement.inspect(247).await; + layout.store().delete(&path).await.unwrap(); + layout + .store() + .create_strict(&path, Bytes::from_static(b"substituted historical suffix")) + .await + .unwrap(); + assert!( + movement + .receiver + .inspect_fleet_action(movement.inspection(248)) + .await + .is_err() + ); + assert_eq!(counter(&handle).await, 43); + movement.shutdown().await; +} diff --git a/crates/cellule-host/tests/node/fleet_receivers/suffix_resume/materialized.rs b/crates/cellule-host/tests/node/fleet_receivers/suffix_resume/materialized.rs new file mode 100644 index 00000000..2ccccf3b --- /dev/null +++ b/crates/cellule-host/tests/node/fleet_receivers/suffix_resume/materialized.rs @@ -0,0 +1,231 @@ +use super::*; + +struct PausedRecovery { + journal: Arc, + accepted: AcceptedFleetAction, + takeover: cellule_runtime::node::NodeTakeoverProof, + entered: Mutex>>, + resume: tokio::sync::Mutex>, +} +impl AcquisitionObserver for PausedRecovery { + fn before_claim<'a>(&'a self, input: &'a Control) -> AcquisitionObservation<'a> { + Box::pin(async move { + let basis = + RecoveryBasis::new(&self.accepted, input.clone(), self.takeover, clock()).unwrap(); + let original = self + .journal + .record_recovery_basis(&self.accepted, &basis) + .await + .unwrap(); + assert_eq!(original.control(), input); + Ok(()) + }) + } + fn before_activation<'a>( + &'a self, + input: &'a Control, + restored: &'a Control, + ) -> AcquisitionObservation<'a> { + Box::pin(async move { + let basis = self + .journal + .load_recovery_basis(&self.accepted) + .await + .unwrap() + .unwrap(); + assert_eq!(basis.control(), input); + let evidence = RecoveryEvidence::new(basis, restored.clone(), clock()).unwrap(); + assert_eq!( + self.journal + .record_recovery_evidence(&self.accepted, &evidence) + .await + .unwrap(), + evidence + ); + self.entered + .lock() + .unwrap() + .take() + .unwrap() + .send(()) + .unwrap(); + let _ = (&mut *self.resume.lock().await).await; + Ok(()) + }) + } +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn committed_overlay_materialization_resumes_only_the_original_checked_control() { + let movement = Movement::new(128 << 20).await; + let now = clock(); + let identity = MutationIdentity { + request_id: RequestId::from_bytes([250; 16]), + issued_at_ms: now, + expires_at_ms: now + 60_000, + }; + let digest = Digest::from_bytes([250; 32]); + let receipt = movement + .source + .handle + .execute(identity, digest, now, 64, 64, |tx| { + tx.execute("UPDATE counter SET value=value", [])?; + Ok(HandlerOutcome::Success(vec![42])) + }) + .await + .unwrap(); + suffix::start_suffix_recovery(&movement).await; + let original = movement + .inputs + .authority + .load(movement.spec.target.cell_id()) + .await + .unwrap() + .unwrap(); + let recovery = movement.recovery_inputs.lock().unwrap().clone().unwrap(); + let FleetActionAcceptance::New(accepted) = movement + .source + .journal + .accept_action( + &movement.action(MovementAction::Recover), + movement.spec.destination_node, + movement.spec.destination, + clock(), + ) + .await + .unwrap() + else { + panic!("expected new recovery") + }; + let prepared = movement + .receiver + .runtime() + .prepared_receiver(movement.spec.id) + .unwrap() + .unwrap(); + movement + .receiver + .runtime() + .cancel_prepared_receiver(&prepared) + .unwrap(); + let (entered, captured) = tokio::sync::oneshot::channel(); + let (_resume, paused) = tokio::sync::oneshot::channel(); + let recorder = Arc::new(PausedRecovery { + journal: movement.source.journal.clone(), + accepted: accepted.clone(), + takeover: recovery.takeover, + entered: Mutex::new(Some(entered)), + resume: tokio::sync::Mutex::new(paused), + }); + let node = movement.receiver.clone(); + let inputs = movement.inputs.clone(); + let expected = original.clone(); + let takeover = recovery.clone(); + let future_recorder = recorder.clone(); + let owner = tokio::spawn(async move { + node.runtime() + .takeover_restored_observed( + inputs.catalog, + inputs.replica, + inputs.authority, + expected, + takeover.takeover, + takeover.manifests, + inputs.destination, + inputs.owner, + Some(future_recorder), + ) + .await + }); + tokio::time::timeout(std::time::Duration::from_secs(5), captured) + .await + .unwrap() + .unwrap(); + // Simulate an embedding owner's interruption after native materialization + // and durable recording, before actor admission. Transport cancellation is + // separately owned by the host action tests and cannot abort that owner. + owner.abort(); + assert!(matches!(owner.await, Err(error) if error.is_cancelled())); + assert_eq!(movement.receiver.stats().active_cells(), 0); + let materialized = movement + .inputs + .authority + .load(original.value().cell) + .await + .unwrap() + .unwrap(); + assert_eq!(materialized.value().state, ControlState::Recovering); + assert!(materialized.value().recovery.is_none()); + let evidence = movement + .source + .journal + .load_recovery_evidence(&accepted) + .await + .unwrap() + .unwrap(); + assert_eq!(evidence.basis().control(), original.value()); + assert_eq!(evidence.restored(), materialized.value()); + let mut changed = original.value().clone(); + changed.epoch += 1; + let refused = movement + .receiver + .runtime() + .resume_takeover_restored_observed( + movement.inputs.catalog.clone(), + movement.inputs.replica.clone(), + movement.inputs.authority.clone(), + changed, + materialized.clone(), + recovery.takeover, + recovery.manifests.clone(), + movement.inputs.destination.clone(), + recorder, + ) + .await; + assert!(matches!(refused, Err(Error::Fenced))); + assert_eq!( + movement + .inputs + .authority + .load(original.value().cell) + .await + .unwrap() + .unwrap() + .value(), + materialized.value() + ); + assert_eq!(movement.receiver.stats().active_cells(), 0); + let completion = apply(&movement.receiver, movement.action(MovementAction::Recover)).await; + assert!( + completion.committed && completion.execution_error.is_none(), + "{completion:?}" + ); + let FleetOutcome::Recovered(recovered) = &completion.outcome.outcome else { + panic!("not recovered") + }; + assert_eq!(recovered.recovery, evidence); + assert_eq!(recovered.serving.position.epoch, original.value().epoch + 1); + movement.event(AttemptEvent::Recovered(recovered.clone())); + let current = movement + .inputs + .authority + .load(original.value().cell) + .await + .unwrap() + .unwrap(); + let handle = movement + .receiver + .runtime() + .local_handle(movement.inputs.catalog.clone(), ¤t) + .await + .unwrap() + .unwrap(); + assert_eq!(counter(&handle).await, 43); + assert_eq!( + handle.resolve(identity, digest, clock(), 64).await.unwrap(), + Resolution::Committed(receipt) + ); + assert_eq!(movement.receiver.stats().active_cells(), 1); + movement.inspect(249).await; + movement.shutdown().await; +} diff --git a/crates/cellule-host/tests/node/fleet_receivers/suffix_resume/mod.rs b/crates/cellule-host/tests/node/fleet_receivers/suffix_resume/mod.rs new file mode 100644 index 00000000..7070d1db --- /dev/null +++ b/crates/cellule-host/tests/node/fleet_receivers/suffix_resume/mod.rs @@ -0,0 +1,9 @@ +//! Retained original input resumes an inherited suffix without another claim. +use super::*; +use bytes::Bytes; +use cellule_host::fleet::FleetActionAcceptance; +use cellule_runtime::cell::actor::{AcquisitionObservation, AcquisitionObserver}; +use cellule_runtime::control::Control; + +mod interrupted; +mod materialized; diff --git a/crates/cellule-runtime/docs/deployment.md b/crates/cellule-runtime/docs/deployment.md index 28895fae..eddb6c57 100644 --- a/crates/cellule-runtime/docs/deployment.md +++ b/crates/cellule-runtime/docs/deployment.md @@ -829,6 +829,17 @@ trusted callbacks grant no authority and authenticate no caller. Their caller must own acquisition independently of transport cancellation; the host fleet executor provides that finite-task ownership. +`resume_takeover_restored_observed` resumes this boot's original interrupted +claim without another ownership CAS. It requires the exact original fenced +control, current Recovering control and takeover proof. The current control must +be that original takeover or its exact pinned-overlay materialization; a changed +generation, scope, owner or publication is refused. A committed materialization +is derived again from the original manifest and compared in full before it can +be retained as native acquisition history. The recorder reconfirms the original +input, then records the actual materialized position before actor admission. +Fresh takeover and resumption share the same materialization, activation, +admission credits and rollback owner. Inspection does not call this method. + `RecoveryBasis` binds an accepted recovery action to its exact failed source, canonical control, and original capture time. Construction requires existing `NodeTakeoverProof`. `RecoveryEvidence` checks the exact takeover and optional @@ -1007,7 +1018,10 @@ can call `CellRuntime::quiesce_cell_at` with the exact Cell ID, runtime boot, local generation, incarnation and ownership epoch. A mismatch fails before closing admission. The transition is sticky for that activation and returns when the actor installs the gate; it does not wait for accepted work to finish. -An unavailable publisher conservatively refuses the request, requiring a retry. +The activation's immutable admission fence remains available while accepted +publication owns its publisher, so closing admission does not require a quiet +publication gap. A wrong epoch or generation remains fenced. Publisher return, +accepted work and exact-root publication still gate final release. | Work | During Cell quiescence | | --- | --- | @@ -1065,6 +1079,12 @@ Cell refuses before closing admission because external stream/upload/pin owners are not covered. `OwnedCellObservation::maintenance_cost` is a separate peak receiver envelope derived from validated per-Cell LTX limits. It stays available while mutations invalidate measured worker samples; it establishes no readiness. +`OwnedCellObservation::owner_fence` retains the activation incarnation/epoch even +when `position` is unavailable during publication. An observer can retain that +exact writer's maintenance demand after independently rechecking its native +generation, canonical owner, boot, executable contract and configured envelope. +Root advancement still invalidates complete counts and ordinary movement demand; +this advisory identity supplies neither final release nor role-settlement proof. The driver can plan with it only for the exact boot/node of a retained Evacuating maintenance operation and a fresh authenticated collection barrier. Receiver preparation precedes source quiescence and rechecks the actual cost. Ordinary diff --git a/crates/cellule-runtime/docs/failover-and-followers.md b/crates/cellule-runtime/docs/failover-and-followers.md index b175374e..1ae46107 100644 --- a/crates/cellule-runtime/docs/failover-and-followers.md +++ b/crates/cellule-runtime/docs/failover-and-followers.md @@ -588,6 +588,13 @@ CAS-protected mutable authority: - A recoverer may change only recovery-owned fields after expiry. - Every transition validates all unchanged fields before conditional overwrite. +Canonical directory reads authenticate the identity and understood placement +signatures on each freshly decoded body. Exact loads and membership/follower +scans then apply their time, path, fleet and release policies to that same +immutable value. A later read authenticates its bytes again; an earlier +signature decision does not carry across versions or scans. Expired signed logs +remain visible obligations, and forged or noncanonical records fail closed. + | Session state | May route application work? | May append to its follower log? | May a peer recover it? | | --- | --- | --- | --- | | `live` before published expiry | Yes | Yes | No | diff --git a/crates/cellule-runtime/docs/primitives.md b/crates/cellule-runtime/docs/primitives.md index e7be6a67..46229b97 100644 --- a/crates/cellule-runtime/docs/primitives.md +++ b/crates/cellule-runtime/docs/primitives.md @@ -318,6 +318,50 @@ A product-level collector must: The helper is not wired to a product collector yet. +### Own accepted Blob operations during shutdown + +All clones of one `BlobArtifactStore` share one irreversible admission word. +The store admits at most 64 original operations. A public namespace mutation +retains staging through its command response; a range read retains metadata +lookup and every part read. Closing between parts cannot interrupt that accepted +read. Cancellation removes the caller's waiter while the original operation +continues. GC retains an `Arc>` with the complete supplied +reference set through original listing/deletion, even after caller loss. + +| API | Local guarantee | +| --- | --- | +| `close()` | Refuse new namespace operations and GC through every clone. | +| `close_and_join()` | Close admission and join known original operations; return an error if any original native join was lost. Cancelled join waiters do not cancel work or reopen admission. | +| `lifecycle_observation()` | Capture admission, accepted-operation count, unjoined original work and the first source-bearing failure. Local joining requires closed admission, zero accepted operations and zero unjoined work. | + +A joined operation can have failed or returned an uncertain command result. +Forced Tokio runtime teardown can discard a supervisor while an original +provider worker still runs. This irreversibly closes admission and retains an +unproven join even after that worker finishes. Zero accepted operations cannot +clear it; `locally_joined()` remains false and repeated close/join returns an +error. Native joining cannot be reconstructed from a later object-store read. +Retained diagnostics preserve the original source; they do not establish remote +absence or success. Preparing a mutation still returns a caller-owned +`PreparedCommand`: after return it holds no running store job. Execution through +that configured client acquires the same store admission and retains dispatch +through its response, including after caller cancellation. Closed admission +refuses new execution from clones and reconstructed snapshots. Convenience +mutation already holds its original operation and does not reacquire admission +between staging and dispatch. Request digests, snapshot bytes, normal Cell +admission and durability are unchanged. Retain evidence and resolve uncertain +earlier execution; local closure cannot prove its outcome or that writes through +other client/provider capabilities in the object-store scope are quiesced. + +Install the same store with `CellNode::install_blob_artifact_store` before +readiness, after the task group, and pass its returned clone to +`CellClient::with_blob_artifact_store`. The existing host drain closes and joins +it before runtime shutdown. This local lifetime boundary supplies no complete +Cell-scoped upload, stream, pin, migration or cross-Cell retention proof. +`BlobInventory` therefore still blocks maintenance release. The global collector +must protect outstanding read/pin obligations as well as authoritative manifest +references and quiesced writes; the store's local count alone cannot authorize +deletion or fleet finalization. + ## Use Cron for failover-safe recurring triggers diff --git a/crates/cellule-runtime/docs/storage.md b/crates/cellule-runtime/docs/storage.md index 05fe69e9..2f78e7b3 100644 --- a/crates/cellule-runtime/docs/storage.md +++ b/crates/cellule-runtime/docs/storage.md @@ -358,6 +358,7 @@ exactly. A matching endpoint without the original overlay is insufficient. | Bounds | The same caller limit caps both acquisition epochs and lineage traversal, at most 10,000. Excess refuses; storage/corrupt-record errors remain errors. | | Missing metadata | An absent original owner yields `OwnerHistoryIncomplete`. Without a matching materialization, missing acquisition records yield `AcquisitionHistoryIncomplete`; different recovery inputs refuse. No legacy rows are fabricated. | | Selected authority | Require Serving at the selected root and recheck owner, incarnation, epoch, state and exact root after origin verification. Lease renewal may continue. | +| Idle rollback | `verify_recovered_idle_prefix` accepts a complete observed unowned Idle control with cleared overlay, rechecking its entire value before and after the same bounded suffix/origin walk. It proves materialization for pre-acquisition validation; it neither acquires nor certifies serving. The Serving verifier still refuses Idle. | | Resource ownership | Reuse the root verifier's shared memory reservation and configured LTX I/O host before metadata/origin I/O. The caller supplies the finite deadline; cancellation drops the reservation. | | Proof scope | `VerifiedRecoveryPrefix` retains the exact required row, materialization epoch and opaque root-prefix proof. It grants no native serving, root pin, authenticated physical boot/backend scope or aggregate settlement. | @@ -389,6 +390,10 @@ before I/O and holds it through verification. After complete origin verification the same native actor through ordinary FIFO admission, selected root/owner/epoch and native ownership inventory before returning fresh serving evidence. Durable historical results remain historical; their replay does not refresh this proof. +Accepted recovery effect replay can reconstruct missing journal evidence from +the exact original native acquisition after safe Idle rollback. It uses the +explicit Idle verifier before ordinary admitted acquisition, then rechecks +current native serving. Read-only fleet inspection performs neither effect. Complete original-writer/suffix aggregation, physical boot/process scope, reader and follower replacement policy, and terminal action joining remain separate requirements before role settlement or finalization. diff --git a/crates/cellule-runtime/src/cell/actor/acquire.rs b/crates/cellule-runtime/src/cell/actor/acquire.rs index 901dc592..a7aaae26 100644 --- a/crates/cellule-runtime/src/cell/actor/acquire.rs +++ b/crates/cellule-runtime/src/cell/actor/acquire.rs @@ -4,6 +4,7 @@ //! takeover proof into a running local handle, and refuses when the //! authority, capacity ledger, or node lease does not agree. +use super::acquire_resume::TakeoverActivation; use super::*; impl CellRuntime { @@ -485,8 +486,6 @@ impl CellRuntime { .replica_with_directory_cache(replica, &destination) .await?; let rollback_node_lease = self.inner.node_lease.guard()?; - let rollback_authority = authority.clone(); - let rollback_replica = replica.clone(); let cell = self.claiming_cell(&catalog, &observed, &owner)?; if owner.session != takeover.claimant() { return Err(Error::Fenced); @@ -540,68 +539,24 @@ impl CellRuntime { } } }; - let recovery_rollback_claim = claimed.clone(); - let claimed = match self - .publish_attached_recovery(&replica, &authority, claimed, &recovery_store) - .await - { - Ok(claimed) => claimed, - Err(error) => { - match rollback_failed_acquisition( - &rollback_authority, - &recovery_rollback_claim, - &rollback_replica, - rollback_node_lease.clone(), - ) - .await - { - Ok(()) => return Err(error), - Err(cleanup) => return Err(cleanup), - } - } - }; - let rollback_claim = claimed.clone(); - let activation = async { - authority - .retain_acquisition(current.value(), claimed.value()) - .await?; - if let Some(observer) = &observer { - observer - .before_activation(current.value(), claimed.value()) - .await?; - } - self.activate_restored_reserved( + return self + .finish_takeover_restored(TakeoverActivation { catalog, replica, authority, + input: current.value().clone(), claimed, + recovery_store, destination, reservation, - None, - ) - .await - } - .await; - return match activation { - Ok(handle) => Ok(handle), - Err(error) => { - match rollback_failed_acquisition( - &rollback_authority, - &rollback_claim, - &rollback_replica, - rollback_node_lease, - ) - .await - { - Ok(()) => Err(error), - Err(cleanup) => Err(cleanup), - } - } - }; + node_lease: rollback_node_lease, + observer, + }) + .await; } } - async fn publish_attached_recovery( + pub(super) async fn publish_attached_recovery( &self, replica: &cellule_ltx::CellReplica, authority: &CellAuthority, @@ -645,7 +600,7 @@ impl CellRuntime { } } - async fn activate_restored_reserved( + pub(super) async fn activate_restored_reserved( &self, catalog: CatalogProof, replica: cellule_ltx::CellReplica, @@ -799,7 +754,7 @@ impl CellRuntime { Ok(()) } - fn activation_cell( + pub(super) fn activation_cell( &self, catalog: &CatalogProof, observed: &VersionedControl, @@ -880,7 +835,7 @@ impl CellRuntime { }) } - async fn replica_with_directory_cache( + pub(super) async fn replica_with_directory_cache( &self, replica: cellule_ltx::CellReplica, destination: &Path, diff --git a/crates/cellule-runtime/src/cell/actor/acquire_resume.rs b/crates/cellule-runtime/src/cell/actor/acquire_resume.rs new file mode 100644 index 00000000..5dd7f5b5 --- /dev/null +++ b/crates/cellule-runtime/src/cell/actor/acquire_resume.rs @@ -0,0 +1,178 @@ +//! Original takeover input survives interrupted materialization and activation. +use super::*; +use crate::control::{Control, ControlState}; +use crate::recovery::manifest::RecoveryManifestStore; + +pub(super) struct TakeoverActivation { + pub catalog: CatalogProof, + pub replica: cellule_ltx::CellReplica, + pub authority: CellAuthority, + pub input: Control, + pub claimed: VersionedControl, + pub recovery_store: RecoveryManifestStore, + pub destination: PathBuf, + pub reservation: CellReservation, + pub node_lease: Option, + pub observer: Option>, +} + +impl CellRuntime { + /// Resumes this boot's exact interrupted takeover without another ownership + /// CAS. The original fenced input must match the current claim or its exact + /// canonical overlay materialization. Reconfirms the durable original input + /// and result through the recorder before actor admission. + #[expect( + clippy::too_many_arguments, + reason = "resumption binds the original fenced input and exact current claim" + )] + pub async fn resume_takeover_restored_observed( + &self, + catalog: CatalogProof, + replica: cellule_ltx::CellReplica, + authority: CellAuthority, + original: Control, + observed: VersionedControl, + takeover: crate::node::NodeTakeoverProof, + recovery_store: RecoveryManifestStore, + destination: PathBuf, + observer: Arc, + ) -> crate::Result { + self.ensure_acquiring()?; + self.check_application_limits(&catalog, replica.limits())?; + self.check_application_limits(&catalog, recovery_store.limits())?; + self.activation_cell(&catalog, &observed)?; + let owner = observed.value().owner.clone().ok_or(Error::Fenced)?; + if observed.value().state != ControlState::Recovering + || original + .owner + .as_ref() + .is_none_or(|owner| owner.session != takeover.session()) + || owner.session != takeover.claimant() + { + return Err(Error::Fenced); + } + let expected = original.takeover(owner)?; + if observed.value() != &expected + && expected + .validate_transition(observed.value(), Transition::PublishRecovery) + .is_err() + { + return Err(Error::Fenced); + } + let current = authority.load(original.cell).await?.ok_or(Error::Fenced)?; + if current.value() != observed.value() { + return Err(Error::Fenced); + } + let replica = self + .replica_with_directory_cache(replica, &destination) + .await?; + let node_lease = self.inner.node_lease.guard()?; + let reservation = self.inner.pool.reserve_activation()?; + observer.before_claim(&original).await?; + if observed.value() != &expected { + // PublishRecovery may have committed before its reply or metadata + // write failed. Derive the exact canonical bytes again; a cleared + // overlay or matching counters alone cannot establish that result. + let recovery = original.recovery.as_ref().ok_or(Error::Fenced)?; + let overlay = recovery_store + .load_overlay(original.cell, original.incarnation, recovery) + .await?; + let prepared = replica + .prepare_recovered_overlay(&overlay, original.schema) + .await?; + let materialized = expected.publish_recovery(&prepared, expected.next_due_ms)?; + if observed.value() != &materialized { + return Err(Error::Fenced); + } + } + self.finish_takeover_restored(TakeoverActivation { + catalog, + replica, + authority, + input: original, + claimed: observed, + recovery_store, + destination, + reservation, + node_lease, + observer: Some(observer), + }) + .await + } + + pub(super) async fn finish_takeover_restored( + &self, + activation: TakeoverActivation, + ) -> crate::Result { + let TakeoverActivation { + catalog, + replica, + authority, + input, + claimed, + recovery_store, + destination, + reservation, + node_lease, + observer, + } = activation; + let rollback_claim = claimed.clone(); + let materialized = match self + .publish_attached_recovery(&replica, &authority, claimed, &recovery_store) + .await + { + Ok(materialized) => materialized, + Err(error) => { + return match rollback_failed_acquisition( + &authority, + &rollback_claim, + &replica, + node_lease, + ) + .await + { + Ok(()) => Err(error), + Err(cleanup) => Err(cleanup), + }; + } + }; + let rollback_claim = materialized.clone(); + let rollback_authority = authority.clone(); + let rollback_replica = replica.clone(); + let activation = async { + authority + .retain_acquisition(&input, materialized.value()) + .await?; + if let Some(observer) = &observer { + observer + .before_activation(&input, materialized.value()) + .await?; + } + self.activate_restored_reserved( + catalog, + replica, + authority, + materialized, + destination, + reservation, + None, + ) + .await + } + .await; + match activation { + Ok(handle) => Ok(handle), + Err(error) => match rollback_failed_acquisition( + &rollback_authority, + &rollback_claim, + &rollback_replica, + node_lease, + ) + .await + { + Ok(()) => Err(error), + Err(cleanup) => Err(cleanup), + }, + } + } +} diff --git a/crates/cellule-runtime/src/cell/actor/acquisition_observer.rs b/crates/cellule-runtime/src/cell/actor/acquisition_observer.rs index f33e7d19..b57c9a62 100644 --- a/crates/cellule-runtime/src/cell/actor/acquisition_observer.rs +++ b/crates/cellule-runtime/src/cell/actor/acquisition_observer.rs @@ -16,6 +16,8 @@ pub trait AcquisitionObserver: Send + Sync + 'static { /// Records the exact canonical input before its ownership CAS. Takeover may /// retry a changed predecessor, so an adapter must explicitly accept or /// reject each input; it must never silently overwrite an earlier basis. + /// Resuming an already claimed takeover reconfirms the original input here + /// before materialization; it performs no additional ownership CAS. fn before_claim<'a>(&'a self, input: &'a Control) -> AcquisitionObservation<'a>; /// Records the exact published recovery/root position before actor activation. diff --git a/crates/cellule-runtime/src/cell/actor/inventory/mod.rs b/crates/cellule-runtime/src/cell/actor/inventory/mod.rs index 424be5dd..f639da41 100644 --- a/crates/cellule-runtime/src/cell/actor/inventory/mod.rs +++ b/crates/cellule-runtime/src/cell/actor/inventory/mod.rs @@ -58,6 +58,9 @@ pub struct OwnedCellObservation { pub generation: u64, /// Authority-pinned incarnation loaded by this activation. pub incarnation: IncarnationId, + /// Immutable activation identity, retained while a publication owns the + /// publisher. This is advisory and grants no release or durability proof. + pub owner_fence: crate::control::OwnerFence, /// Current executable contract identity. pub code: Digest, /// Current schema version. @@ -296,6 +299,7 @@ fn observe_owned(active: &ActiveCell) -> crate::Result { target: active.catalog.target()?, generation: active.generation, incarnation: active.incarnation, + owner_fence: active.admission.owner_fence, code: active.code, schema: active.schema, role: active.role, diff --git a/crates/cellule-runtime/src/cell/actor/maintenance/release.rs b/crates/cellule-runtime/src/cell/actor/maintenance/release.rs index e0760bdc..0278ffaa 100644 --- a/crates/cellule-runtime/src/cell/actor/maintenance/release.rs +++ b/crates/cellule-runtime/src/cell/actor/maintenance/release.rs @@ -23,11 +23,10 @@ pub(in crate::cell::actor) fn begin( let _ = reply.send(Err(Error::Fenced)); return None; } - let Some(publisher) = active.publisher.as_ref() else { - refuse(reply, DrainBlocker::BusyExecution, None); - return None; - }; - if publisher.control().value().epoch != epoch { + // This exact activation fence survives temporary publisher ownership by an + // accepted publication. Preflight still waits for that owner to return and + // for complete native settlement before the canonical final release. + if active.admission.owner_fence.epoch != epoch { let _ = reply.send(Err(Error::Fenced)); return None; } diff --git a/crates/cellule-runtime/src/cell/actor/mod.rs b/crates/cellule-runtime/src/cell/actor/mod.rs index 4a90278f..4ed0b7a5 100644 --- a/crates/cellule-runtime/src/cell/actor/mod.rs +++ b/crates/cellule-runtime/src/cell/actor/mod.rs @@ -21,6 +21,7 @@ pub use inventory::{ use state::*; use task::*; mod acquire; +mod acquire_resume; mod acquisition_observer; mod prefix; mod serving; diff --git a/crates/cellule-runtime/src/cell/actor/prefix.rs b/crates/cellule-runtime/src/cell/actor/prefix.rs index e2bc0396..def83263 100644 --- a/crates/cellule-runtime/src/cell/actor/prefix.rs +++ b/crates/cellule-runtime/src/cell/actor/prefix.rs @@ -61,6 +61,28 @@ impl CellRuntime { Ok(proof) } + /// Verifies a selected unowned Idle root against the original sealed suffix + /// through the same bounded lineage/origin walk and shared I/O admission. + /// The complete Idle control is rechecked before and after origin I/O. + /// This does not acquire a writer or certify current serving or settlement. + pub async fn verify_recovered_idle_prefix( + &self, + catalog: &CatalogProof, + authority: &CellAuthority, + replica: cellule_ltx::CellReplica, + required: &crate::recovery::manifest::PinnedRecoveryCell, + observed: &VersionedControl, + limit: usize, + ) -> crate::Result { + let root = observed.value().ltx_root().ok_or(Error::Fenced)?; + let (_metadata, replica) = self.prefix_replica(catalog, replica, root, limit)?; + let proof = authority + .verify_recovered_idle_prefix(required, observed, &replica, limit) + .await?; + self.ensure_running()?; + Ok(proof) + } + fn prefix_replica( &self, catalog: &CatalogProof, diff --git a/crates/cellule-runtime/src/cell/actor/task.rs b/crates/cellule-runtime/src/cell/actor/task.rs index f31f43a9..b33d783e 100644 --- a/crates/cellule-runtime/src/cell/actor/task.rs +++ b/crates/cellule-runtime/src/cell/actor/task.rs @@ -763,8 +763,10 @@ pub(super) fn handle_message( if active.generation != generation || active.incarnation != incarnation { return Err(Error::Fenced); } - let publisher = active.publisher.as_ref().ok_or(Error::CellDraining)?; - if publisher.control().value().epoch != epoch { + // Publication temporarily owns the publisher. The activation's + // immutable admission fence still identifies this writer, so + // closing foreground admission must not wait for a quiet gap. + if active.admission.owner_fence.epoch != epoch { return Err(Error::Fenced); } match active diff --git a/crates/cellule-runtime/src/client/mod.rs b/crates/cellule-runtime/src/client/mod.rs index a33821c0..ea6c046e 100644 --- a/crates/cellule-runtime/src/client/mod.rs +++ b/crates/cellule-runtime/src/client/mod.rs @@ -365,7 +365,28 @@ impl PreparedCommand { /// Executes the prepared request once against its validated owner incarnation. /// /// Returns pending evidence when acceptance is unknown and rejects an expired identity. + /// A configured Blob namespace also requires the original artifact store's + /// admission. Clones and restored requests cannot dispatch after its closure; + /// accepted dispatch remains owned through native completion after waiter loss. pub async fn execute( + self, + ) -> std::result::Result, InvocationError> { + if self + .client + .registry + .namespace_contract(self.evidence.target.namespace()) + .is_some_and(|(_, namespace)| namespace.role == crate::cell::catalog::CatalogRole::Blob) + && let Some(store) = self.client.blob_artifact_store() + { + return store.run_invocation(self.execute_native()).await; + } + self.execute_native().await + } + + // Namespace convenience mutation already owns staging and dispatch in one + // original artifact lifetime. Re-admission here could refuse its accepted + // manifest write after closure or consume a second slot at the job bound. + pub(crate) async fn execute_native( mut self, ) -> std::result::Result, InvocationError> { let now_ms = unix_time_ms().map_err(InvocationError::NotStarted)?; diff --git a/crates/cellule-runtime/src/control/authority/acquisition/prefix.rs b/crates/cellule-runtime/src/control/authority/acquisition/prefix.rs index 46493af0..b4a0b1fe 100644 --- a/crates/cellule-runtime/src/control/authority/acquisition/prefix.rs +++ b/crates/cellule-runtime/src/control/authority/acquisition/prefix.rs @@ -46,6 +46,50 @@ impl CellAuthority { replica: &cellule_ltx::CellReplica, limit: usize, ) -> Result { + let first = self.recovered_prefix_start(required, root, replica, limit)?; + let selected = self.load(required.cell).await?.ok_or(Error::Fenced)?; + if selected.value().state != ControlState::Serving { + return Err(Error::Fenced); + } + self.verify_recovered_prefix_selected(required, root, replica, limit, selected, first) + .await + } + + /// Verifies an exact unowned Idle rollback root against its original sealed + /// suffix and native acquisition history. This grants no ownership, actor + /// admission or serving proof. The complete selected control must still + /// equal the caller's observation before and after the bounded origin walk. + pub async fn verify_recovered_idle_prefix( + &self, + required: &PinnedRecoveryCell, + observed: &VersionedControl, + replica: &cellule_ltx::CellReplica, + limit: usize, + ) -> Result { + if observed.value().state != ControlState::Idle + || observed.value().owner.is_some() + || observed.value().recovery.is_some() + || observed.value().cell != required.cell + { + return Err(Error::Fenced); + } + let root = observed.value().ltx_root().ok_or(Error::Fenced)?; + let first = self.recovered_prefix_start(required, root, replica, limit)?; + let selected = self.load(required.cell).await?.ok_or(Error::Fenced)?; + if selected.value() != observed.value() { + return Err(Error::Fenced); + } + self.verify_recovered_prefix_selected(required, root, replica, limit, selected, first) + .await + } + + fn recovered_prefix_start( + &self, + required: &PinnedRecoveryCell, + root: cellule_ltx::RootRef, + replica: &cellule_ltx::CellReplica, + limit: usize, + ) -> Result { if limit == 0 || limit > MAX_LINEAGE_ROOTS { return Err(Error::Capacity("invalid Cell root lineage traversal bound")); } @@ -61,10 +105,20 @@ impl CellAuthority { .cell_epoch .checked_add(1) .ok_or(Error::Control("recovered acquisition epoch overflow"))?; - let selected = self.load(required.cell).await?.ok_or(Error::Fenced)?; + Ok(first) + } + + async fn verify_recovered_prefix_selected( + &self, + required: &PinnedRecoveryCell, + root: cellule_ltx::RootRef, + replica: &cellule_ltx::CellReplica, + limit: usize, + selected: VersionedControl, + first: u64, + ) -> Result { if selected.value().incarnation != required.incarnation || selected.value().ltx_root() != Some(root) - || selected.value().state != ControlState::Serving || selected.value().epoch < first { return Err(Error::Fenced); @@ -139,8 +193,10 @@ impl CellAuthority { if confirmed.value().owner != selected.value().owner || confirmed.value().epoch != selected.value().epoch || confirmed.value().incarnation != required.incarnation - || confirmed.value().state != ControlState::Serving + || confirmed.value().state != selected.value().state || confirmed.value().ltx_root() != Some(root) + || (selected.value().state == ControlState::Idle + && confirmed.value() != selected.value()) { return Err(Error::Fenced); } diff --git a/crates/cellule-runtime/src/control/authority/acquisition/tests.rs b/crates/cellule-runtime/src/control/authority/acquisition/tests.rs index 5fc7d533..47ad6585 100644 --- a/crates/cellule-runtime/src/control/authority/acquisition/tests.rs +++ b/crates/cellule-runtime/src/control/authority/acquisition/tests.rs @@ -461,3 +461,66 @@ async fn matching_endpoint_cannot_replace_a_sealed_recovery_input() { )) )); } + +#[tokio::test] +async fn idle_suffix_verification_requires_exact_control_and_does_not_certify_serving() { + let authority = authority(Arc::new(InMemory::new())); + let required = suffix(); + let mut selected = successor(&input()); + selected.state = ControlState::Idle; + selected.owner = None; + selected.recovery = None; + selected.root = Some(RootRef { + digest: Digest::from_bytes([8; 32]), + txid: required.recovery.final_txid, + checksum: required.recovery.final_checksum, + commit_sequence: required.recovery.final_commit_sequence, + }); + original_suffix_scope(&authority, &required, &selected).await; + let observed = authority.load(required.cell).await.unwrap().unwrap(); + let root = observed.value().ltx_root().unwrap(); + let replica = cellule_ltx::CellReplica::new( + authority.layout.clone(), + root.cell, + root.incarnation, + cellule_ltx::Limits::default(), + ) + .unwrap(); + assert!(matches!( + authority + .verify_recovered_prefix(&required, root, &replica, 8) + .await, + Err(Error::Fenced) + )); + assert!(matches!( + authority + .verify_recovered_idle_prefix(&required, &observed, &replica, 0) + .await, + Err(Error::Capacity(_)) + )); + assert!(matches!( + authority + .verify_recovered_idle_prefix(&required, &observed, &replica, 8) + .await, + Err(Error::AcquisitionHistoryIncomplete { epoch: 2, .. }) + )); + // Even an unchanged root at a newer revision cannot replace the complete + // originally selected Idle control. No origin read or acquisition follows. + selected.revision += 1; + let path = authority.layout.control_path(required.cell.as_bytes()); + authority.layout.store().delete(&path).await.unwrap(); + authority + .layout + .store() + .create_strict(&path, Bytes::from(selected.encode().unwrap())) + .await + .unwrap(); + assert!(matches!( + authority + .verify_recovered_idle_prefix(&required, &observed, &replica, 8) + .await, + Err(Error::Fenced) + )); + let latest = authority.load(required.cell).await.unwrap().unwrap(); + assert_eq!(latest.value(), &selected); +} diff --git a/crates/cellule-runtime/src/fleet/operations/accepted.rs b/crates/cellule-runtime/src/fleet/operations/accepted.rs index af4ccb72..54db78f8 100644 --- a/crates/cellule-runtime/src/fleet/operations/accepted.rs +++ b/crates/cellule-runtime/src/fleet/operations/accepted.rs @@ -2,7 +2,7 @@ use crate::identity::{NodeId, SessionId}; use super::{ FleetAction, FleetActionKind, FleetActionOutcome, FleetHead, MovementAction, OperationError, - Result, nonzero, + RegistryVersion, Result, nonzero, }; /// Immutable acceptance of one exact local fleet effect. @@ -39,6 +39,27 @@ impl AcceptedFleetAction { Ok(accepted) } + /// Checks a continuation against the current head and exact registry in + /// the same transaction that durably accepts the receiver action. + pub fn new_with_registry( + action: FleetAction, + head: &FleetHead, + registry: RegistryVersion, + node: NodeId, + session: SessionId, + now_ms: i64, + ) -> Result { + action.authorize_against_registry(head, registry, now_ms)?; + let accepted = Self { + action, + node, + session, + accepted_at_ms: now_ms, + }; + accepted.validate()?; + Ok(accepted) + } + /// Returns the originally accepted envelope, including its authorization. #[must_use] pub const fn action(&self) -> &FleetAction { @@ -138,7 +159,11 @@ impl FleetAction { action: replay, attempt: b, }, - ) => original == replay && a.spec() == b.spec(), + ) => { + original == replay + && a.spec() == b.spec() + && self.receiver_route == action.receiver_route + } ( FleetActionKind::Maintenance { action: original, @@ -176,7 +201,7 @@ impl FleetAction { FleetActionKind::Movement { action, attempt } => { let spec = attempt.spec(); let source = node == spec.source_node && session == spec.source; - let receiver = node == spec.destination_node && session == spec.destination; + let receiver = self.receiver_endpoint() == Some((node, session)); match action { MovementAction::Release | MovementAction::ReleaseMaintenance => source, MovementAction::Prepare diff --git a/crates/cellule-runtime/src/fleet/operations/actions.rs b/crates/cellule-runtime/src/fleet/operations/actions.rs index 6cd5e507..9a6b4c0d 100644 --- a/crates/cellule-runtime/src/fleet/operations/actions.rs +++ b/crates/cellule-runtime/src/fleet/operations/actions.rs @@ -2,10 +2,218 @@ use crate::identity::{Digest, NodeId, SessionId}; use super::{ ActivationEvidence, AttemptId, AttemptPhase, DrainBlocker, DrainEvidence, FleetHead, - FleetScope, MaintenanceOperation, MaintenancePhase, MoveAttempt, MovementAction, - OperationError, PublishedPosition, ReceiverReservation, RecoveredActivation, Result, nonzero, + FleetScope, MAX_RECEIVER_HANDOFFS, MaintenanceOperation, MaintenancePhase, MoveAttempt, + MovementAction, OperationError, PublishedPosition, ReceiverReservation, RecoveredActivation, + RegistryVersion, Result, nonzero, }; +/// One exact, process-closed receiver replacement in an immutable attempt. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct ReceiverHandoff { + pub(super) previous_node: NodeId, + pub(super) previous_session: SessionId, + pub(super) target_node: NodeId, + pub(super) target_session: SessionId, + pub(super) process_closure: Digest, + pub(super) registry: RegistryVersion, +} + +impl ReceiverHandoff { + /// Returns the exact receiver boot whose process closure permits this hop. + #[must_use] + pub const fn previous(&self) -> (NodeId, SessionId) { + (self.previous_node, self.previous_session) + } + + /// Returns the exact newly selected receiver boot. + #[must_use] + pub const fn target(&self) -> (NodeId, SessionId) { + (self.target_node, self.target_session) + } + + /// Returns the digest of the checked process-closure evidence. + #[must_use] + pub const fn process_closure(&self) -> Digest { + self.process_closure + } + + /// Returns the current registry barrier captured with the handoff. + #[must_use] + pub const fn registry(&self) -> RegistryVersion { + self.registry + } +} + +/// Bounded, ordered receiver replacement history for one movement dispatch. +/// +/// The embedding host constructs this from a fresh closed-boot proof and +/// current placement observations. This record is shape-only: the receiving +/// journal must match its final hop to the opaque proof and current registry +/// before first acceptance. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct ReceiverRoute { + pub(super) hops: Vec, +} + +impl ReceiverRoute { + /// Starts a route after the originally preferred receiver process closed. + pub fn begin( + scope: FleetScope, + spec: &super::MoveAttemptSpec, + target_node: NodeId, + target_session: SessionId, + process_closure: Digest, + registry: RegistryVersion, + ) -> Result { + let route = Self { + hops: vec![ReceiverHandoff { + previous_node: spec.destination_node, + previous_session: spec.destination, + target_node, + target_session, + process_closure, + registry, + }], + }; + route.validate_for(scope, spec)?; + Ok(route) + } + + /// Extends a route after the current receiver process has closed. + pub fn extend( + &self, + scope: FleetScope, + spec: &super::MoveAttemptSpec, + target_node: NodeId, + target_session: SessionId, + process_closure: Digest, + registry: RegistryVersion, + ) -> Result { + self.validate_for(scope, spec)?; + if self.hops.len() >= MAX_RECEIVER_HANDOFFS { + return Err(OperationError::Invalid("receiver handoff limit reached")); + } + let previous = self.endpoint(spec); + let route = Self { + hops: self + .hops + .iter() + .cloned() + .chain(std::iter::once(ReceiverHandoff { + previous_node: previous.0, + previous_session: previous.1, + target_node, + target_session, + process_closure, + registry, + })) + .collect(), + }; + route.validate_for(scope, spec)?; + Ok(route) + } + + /// Returns the latest selected receiver boot. + #[must_use] + pub fn target(&self, spec: &super::MoveAttemptSpec) -> (NodeId, SessionId) { + self.endpoint(spec) + } + + /// Returns the number of process-closed hops represented by this route. + #[must_use] + pub const fn hop_count(&self) -> usize { + self.hops.len() + } + + /// Returns the final closed-boot proof in this route. + #[must_use] + pub fn latest_handoff(&self) -> Option<&ReceiverHandoff> { + self.hops.last() + } + + /// Checks whether this route is the same route or appends exactly one hop. + #[must_use] + pub fn follows(&self, previous: &Self) -> bool { + self.hops == previous.hops + || (self.hops.len() == previous.hops.len() + 1 && self.hops.starts_with(&previous.hops)) + } + + /// Returns the registry barrier for the latest handoff. + #[must_use] + pub fn registry(&self) -> Option { + self.hops.last().map(|hop| hop.registry) + } + + pub(super) fn validate_for( + &self, + scope: FleetScope, + spec: &super::MoveAttemptSpec, + ) -> Result<()> { + scope.validate()?; + spec.validate()?; + if self.hops.is_empty() || self.hops.len() > MAX_RECEIVER_HANDOFFS { + return Err(OperationError::Invalid("invalid receiver route length")); + } + let mut previous = (spec.destination_node, spec.destination); + let mut visited = vec![spec.source_node, spec.destination_node]; + let mut previous_revision = 0; + for hop in &self.hops { + hop.registry.validate()?; + if hop.previous_node != previous.0 + || hop.previous_session != previous.1 + || !nonzero(hop.target_node.as_bytes()) + || !nonzero(hop.target_session.as_bytes()) + || !nonzero(hop.process_closure.as_bytes()) + || hop.target_node == spec.source_node + || visited.contains(&hop.target_node) + || hop.target_session == spec.source + || hop.registry.scope() != scope + || hop.registry.revision() == 0 + || hop.registry.bootstrap_revision().is_none() + || hop.registry.revision() <= previous_revision + { + return Err(OperationError::Invalid("invalid receiver handoff")); + } + previous = (hop.target_node, hop.target_session); + visited.push(hop.target_node); + previous_revision = hop.registry.revision(); + } + Ok(()) + } + + fn endpoint(&self, spec: &super::MoveAttemptSpec) -> (NodeId, SessionId) { + self.hops + .last() + .map_or((spec.destination_node, spec.destination), |hop| { + (hop.target_node, hop.target_session) + }) + } + + fn digest(&self) -> Digest { + let mut hash = blake3::Hasher::new(); + hash.update(b"cellule.fleet-receiver-route.v1\0"); + hash.update(&(self.hops.len() as u64).to_be_bytes()); + for hop in &self.hops { + hash.update(hop.previous_node.as_bytes()); + hash.update(hop.previous_session.as_bytes()); + hash.update(hop.target_node.as_bytes()); + hash.update(hop.target_session.as_bytes()); + hash.update(hop.process_closure.as_bytes()); + hash.update(hop.registry.scope().fleet.as_bytes()); + hash.update(hop.registry.scope().application.as_bytes()); + hash.update(&hop.registry.revision().to_be_bytes()); + hash.update( + &hop.registry + .bootstrap_revision() + .unwrap_or_default() + .to_be_bytes(), + ); + hash.update(&[u8::from(hop.registry.scheduling_enabled())]); + } + Digest::from_bytes(*hash.finalize().as_bytes()) + } +} + /// Node lifecycle work recorded by one maintenance operation. #[derive(Clone, Copy, Debug, PartialEq, Eq)] #[repr(u8)] @@ -52,6 +260,7 @@ pub struct FleetAction { pub(super) controller_epoch: u64, pub(super) issued_at_ms: i64, pub(super) kind: FleetActionKind, + pub(super) receiver_route: Option, } impl FleetAction { @@ -86,27 +295,69 @@ impl FleetAction { &self.kind } + /// Returns the process-closed receiver route, when this is a continuation. + #[must_use] + pub const fn receiver_route(&self) -> Option<&ReceiverRoute> { + self.receiver_route.as_ref() + } + + /// Returns the exact receiver endpoint used by this action, if it is movement. + #[must_use] + pub fn receiver_endpoint(&self) -> Option<(NodeId, SessionId)> { + match &self.kind { + FleetActionKind::Movement { attempt, .. } => Some(self.receiver_endpoint_for(attempt)), + FleetActionKind::Maintenance { .. } => None, + } + } + + fn receiver_endpoint_for(&self, attempt: &MoveAttempt) -> (NodeId, SessionId) { + self.receiver_route.as_ref().map_or( + (attempt.spec.destination_node, attempt.spec.destination), + |route| route.endpoint(&attempt.spec), + ) + } + /// Returns a stable deduplication key across controller adoption and retries. /// Authorization revision/time are deliberately excluded; the current /// journal is still required to accept an effect for the first time. pub fn key(&self) -> Result { self.validate()?; - let mut hash = blake3::Hasher::new(); - hash.update(b"cellule.fleet-action-key.v1\0"); - hash.update(self.scope.fleet.as_bytes()); - hash.update(self.scope.application.as_bytes()); match &self.kind { FleetActionKind::Movement { action, attempt } => { - return Ok(movement_key(self.scope, *action, &attempt.spec)); + let base = movement_key(self.scope, *action, &attempt.spec); + Ok(self.receiver_route.as_ref().map_or(base, |route| { + let mut hash = blake3::Hasher::new(); + hash.update(b"cellule.fleet-routed-action-key.v1\0"); + hash.update(base.as_bytes()); + hash.update(route.digest().as_bytes()); + Digest::from_bytes(*hash.finalize().as_bytes()) + })) } FleetActionKind::Maintenance { action, operation } => { - hash.update(&[2, *action as u8]); - hash.update(operation.id.as_bytes()); - hash.update(operation.node.as_bytes()); - hash.update(operation.session.as_bytes()); - hash.update(&operation.intent_revision.to_be_bytes()); + Self::maintenance_action_key(self.scope, *action, operation) } } + } + + /// Returns a maintenance action's stable key without requiring the action + /// to be dispatchable in the operation's current phase. Journals use this + /// to verify the exact SettleRoles receipt atomically with ReadyToClose. + pub fn maintenance_action_key( + scope: FleetScope, + action: MaintenanceAction, + operation: &MaintenanceOperation, + ) -> Result { + scope.validate()?; + operation.validate()?; + let mut hash = blake3::Hasher::new(); + hash.update(b"cellule.fleet-action-key.v1\0"); + hash.update(scope.fleet.as_bytes()); + hash.update(scope.application.as_bytes()); + hash.update(&[2, action as u8]); + hash.update(operation.id.as_bytes()); + hash.update(operation.node.as_bytes()); + hash.update(operation.session.as_bytes()); + hash.update(&operation.intent_revision.to_be_bytes()); Ok(Digest::from_bytes(*hash.finalize().as_bytes())) } @@ -114,6 +365,11 @@ impl FleetAction { /// Caller identity and backend authenticity are checked by the application. pub fn authorize_against(&self, head: &FleetHead, now_ms: i64) -> Result<()> { self.validate()?; + if self.receiver_route.is_some() { + return Err(OperationError::Invalid( + "receiver continuation requires a current registry", + )); + } head.check_action_controller(now_ms)?; if self.scope != head.scope || self.journal_revision != head.revision @@ -135,6 +391,58 @@ impl FleetAction { self.check_admission_deadline(now_ms) } + /// Verifies routed authorization against the exact current registry. + /// Ordinary actions retain the registry-free validator and byte format. + /// This shape check does not replace the journal's typed closure check. + pub fn authorize_against_registry( + &self, + head: &FleetHead, + registry: RegistryVersion, + now_ms: i64, + ) -> Result<()> { + self.validate()?; + let Some(route) = &self.receiver_route else { + return self.authorize_against(head, now_ms); + }; + registry.validate()?; + if route.registry() != Some(registry) { + return Err(OperationError::Conflict); + } + let FleetActionKind::Movement { action, attempt } = &self.kind else { + return Err(OperationError::Invalid("receiver route is not movement")); + }; + route.validate_for(self.scope, &attempt.spec)?; + if !matches!( + action, + MovementAction::Activate + | MovementAction::Recover + | MovementAction::Inspect + | MovementAction::Cancel + ) || attempt.released().is_none() + { + return Err(OperationError::Invalid( + "receiver continuation is not a post-release action", + )); + } + head.check_action_controller(now_ms)?; + if self.scope != head.scope + || self.journal_revision != head.revision + || self.issued_at_ms > now_ms + { + return Err(OperationError::Conflict); + } + let expected = head.movement_action_with_receiver_route( + attempt.spec.id, + *action, + route.clone(), + self.issued_at_ms, + )?; + if *self != expected { + return Err(OperationError::Fenced); + } + self.check_admission_deadline(now_ms) + } + pub(super) fn check_admission_deadline(&self, now_ms: i64) -> Result<()> { if let FleetActionKind::Movement { action, attempt } = &self.kind { if matches!( @@ -172,7 +480,7 @@ impl FleetAction { FleetActionKind::Movement { action, attempt } => { attempt.validate()?; if attempt.spec.target.application() != self.scope.application - || !movement_allowed(attempt, *action) + || !movement_allowed(attempt, *action, self.receiver_route.is_some()) { return Err(OperationError::Invalid("movement dispatch phase mismatch")); } @@ -186,6 +494,24 @@ impl FleetAction { } } } + if let Some(route) = &self.receiver_route { + let FleetActionKind::Movement { action, attempt } = &self.kind else { + return Err(OperationError::Invalid("receiver route is not movement")); + }; + route.validate_for(self.scope, &attempt.spec)?; + if !matches!( + action, + MovementAction::Activate + | MovementAction::Recover + | MovementAction::Inspect + | MovementAction::Cancel + ) || attempt.released().is_none() + { + return Err(OperationError::Invalid( + "receiver continuation is not a post-release action", + )); + } + } self.check_admission_deadline(self.issued_at_ms) } } @@ -242,6 +568,31 @@ impl FleetHead { ) } + /// Builds a post-release action bound to a closed receiver and its current + /// replacement. First acceptance also requires the journal's typed closure + /// check; a routed action must not reach a remote effect before that commit. + pub fn movement_action_with_receiver_route( + &self, + id: AttemptId, + action: MovementAction, + route: ReceiverRoute, + now_ms: i64, + ) -> Result { + let attempt = self + .attempts + .iter() + .find(|attempt| attempt.spec.id == id) + .ok_or(OperationError::NotFound)?; + self.make_action_with_route( + FleetActionKind::Movement { + action, + attempt: Box::new(attempt.clone()), + }, + Some(route), + now_ms, + ) + } + /// Builds lifecycle work only from the current committed maintenance operation. pub fn maintenance_action( &self, @@ -259,6 +610,15 @@ impl FleetHead { } fn make_action(&self, kind: FleetActionKind, now_ms: i64) -> Result { + self.make_action_with_route(kind, None, now_ms) + } + + fn make_action_with_route( + &self, + kind: FleetActionKind, + receiver_route: Option, + now_ms: i64, + ) -> Result { self.check_action_controller(now_ms)?; let controller = self.controller.ok_or(OperationError::Fenced)?; let action = FleetAction { @@ -268,17 +628,18 @@ impl FleetHead { controller_epoch: controller.epoch, issued_at_ms: now_ms, kind, + receiver_route, }; action.validate()?; Ok(action) } } -fn movement_allowed(attempt: &MoveAttempt, action: MovementAction) -> bool { +fn movement_allowed(attempt: &MoveAttempt, action: MovementAction, routed: bool) -> bool { if action == MovementAction::Inspect { return true; } - if attempt.blocker == Some(DrainBlocker::OutcomeUnknown) { + if attempt.blocker == Some(DrainBlocker::OutcomeUnknown) && !routed { return false; } match action { @@ -330,11 +691,23 @@ pub enum FleetOutcome { ReceiverCleaned, /// The exact local node session has applied its persistent admission closure. Cordoned, - /// Complete role inventory at the named barrier reports zero obligations. + /// Legacy role inventory result retained for decoding; it lacks the registry + /// barrier and therefore cannot authorize a new Closing transition. RolesSettled { /// Digest of the authoritative inventory used at the final barrier. inventory: Digest, }, + /// Role settlement bound to the exact registry version observed at completion. + /// The digest must bind complete native/foreign inventories, joined accepted + /// work, and the checked replacement-policy evidence for this same version. + RolesSettledAt { + /// Digest of the complete settlement evidence at `registry`. + inventory: Digest, + /// Exact journal head revision observed with the complete inventory. + head_revision: u64, + /// Exact intent/enrollment/evidence version captured at the barrier. + registry: RegistryVersion, + }, /// Complete relocation, runtime/facility shutdown, and withdrawal evidence. Stopped(DrainEvidence), } @@ -397,6 +770,20 @@ impl FleetActionOutcome { FleetOutcome::RolesSettled { inventory } if !nonzero(inventory.as_bytes()) => { Err(OperationError::Invalid("role result lacks inventory")) } + FleetOutcome::RolesSettledAt { + inventory, + head_revision, + registry, + } if !nonzero(inventory.as_bytes()) + || *head_revision == 0 + || registry.scope() != self.scope + || registry.revision() == 0 + || registry.bootstrap_revision().is_none() => + { + Err(OperationError::Invalid( + "role result lacks a bootstrapped inventory barrier", + )) + } FleetOutcome::Stopped(e) if e.node != self.node || e.session != self.session @@ -424,15 +811,24 @@ impl FleetActionOutcome { if self.scope != action.scope || self.action_key != action.key()? { return Err(OperationError::Invalid("fleet reply action mismatch")); } + if let FleetOutcome::RolesSettledAt { head_revision, .. } = &self.outcome + && *head_revision < action.journal_revision() + { + return Err(OperationError::Invalid( + "role result predates its accepted journal head", + )); + } let permitted = match &action.kind { FleetActionKind::Movement { action: kind, attempt, } => { + let receiver_endpoint = action + .receiver_endpoint() + .ok_or(OperationError::Invalid("movement action lacks receiver"))?; let source = self.node == attempt.spec.source_node && self.session == attempt.spec.source; - let receiver = self.node == attempt.spec.destination_node - && self.session == attempt.spec.destination; + let receiver = receiver_endpoint == (self.node, self.session); match &self.outcome { FleetOutcome::Reserved(_) => { receiver @@ -455,6 +851,7 @@ impl FleetActionOutcome { && e.position.epoch > attempt.spec.source_epoch && e.session != attempt.spec.source && (e.node != attempt.spec.destination_node || receiver) + && (action.receiver_route.is_none() || receiver) } FleetOutcome::Recovered(e) => { receiver @@ -489,14 +886,23 @@ impl FleetActionOutcome { FleetOutcome::Cordoned => { matches!(kind, MaintenanceAction::Cordon | MaintenanceAction::Inspect) } - FleetOutcome::RolesSettled { .. } => matches!( - kind, - MaintenanceAction::SettleRoles | MaintenanceAction::Inspect - ), - FleetOutcome::Stopped(_) => matches!( - kind, - MaintenanceAction::Finalize | MaintenanceAction::Inspect - ), + FleetOutcome::RolesSettled { .. } | FleetOutcome::RolesSettledAt { .. } => { + matches!( + kind, + MaintenanceAction::SettleRoles | MaintenanceAction::Inspect + ) + } + FleetOutcome::Stopped(evidence) => { + matches!( + kind, + MaintenanceAction::Finalize | MaintenanceAction::Inspect + ) && operation.drain_evidence().is_some_and(|mut expected| { + expected.facilities_closed = true; + expected.stopped = true; + expected.withdrawn = true; + evidence == expected + }) + } FleetOutcome::Rejected(_) | FleetOutcome::Blocked(_) | FleetOutcome::Unknown => true, diff --git a/crates/cellule-runtime/src/fleet/operations/codec/actions.rs b/crates/cellule-runtime/src/fleet/operations/codec/actions.rs index 17e1e501..1a7c538a 100644 --- a/crates/cellule-runtime/src/fleet/operations/codec/actions.rs +++ b/crates/cellule-runtime/src/fleet/operations/codec/actions.rs @@ -1,5 +1,41 @@ +use super::registry::{read_version, write_version}; use super::*; +fn write_receiver_route(e: &mut BoundedEncoder, route: &ReceiverRoute) -> Result<()> { + e.write_u8( + u8::try_from(route.hops.len()) + .map_err(|_| OperationError::Invalid("receiver route length overflow"))?, + )?; + for hop in &route.hops { + e.write_bytes(hop.previous_node.as_bytes())?; + e.write_bytes(hop.previous_session.as_bytes())?; + e.write_bytes(hop.target_node.as_bytes())?; + e.write_bytes(hop.target_session.as_bytes())?; + e.write_bytes(hop.process_closure.as_bytes())?; + write_version(e, hop.registry)?; + } + Ok(()) +} + +fn read_receiver_route(d: &mut BoundedDecoder<'_>) -> Result { + let count = usize::from(d.read_u8()?); + if count == 0 || count > MAX_RECEIVER_HANDOFFS { + return Err(OperationError::Invalid("invalid receiver route length")); + } + let mut hops = Vec::with_capacity(count); + for _ in 0..count { + hops.push(ReceiverHandoff { + previous_node: NodeId::from_bytes(fixed(d)?), + previous_session: SessionId::from_bytes(fixed(d)?), + target_node: NodeId::from_bytes(fixed(d)?), + target_session: SessionId::from_bytes(fixed(d)?), + process_closure: Digest::from_bytes(fixed(d)?), + registry: read_version(d)?, + }); + } + Ok(ReceiverRoute { hops }) +} + fn movement(d: &mut BoundedDecoder<'_>) -> Result { match d.read_u8()? { 1 => Ok(MovementAction::Prepare), @@ -34,6 +70,14 @@ impl FleetAction { e.write_u64(self.controller_epoch)?; e.write_i64(self.issued_at_ms)?; match &self.kind { + FleetActionKind::Movement { action, attempt } + if let Some(route) = &self.receiver_route => + { + e.write_u8(3)?; + e.write_u8(*action as u8)?; + write_attempt(&mut e, attempt)?; + write_receiver_route(&mut e, route)?; + } FleetActionKind::Movement { action, attempt } => { e.write_u8(1)?; e.write_u8(*action as u8)?; @@ -57,15 +101,28 @@ impl FleetAction { let controller = SessionId::from_bytes(fixed(&mut d)?); let controller_epoch = d.read_u64()?; let issued_at_ms = d.read_i64()?; - let kind = match d.read_u8()? { - 1 => FleetActionKind::Movement { - action: movement(&mut d)?, - attempt: Box::new(read_attempt(&mut d)?), - }, - 2 => FleetActionKind::Maintenance { - action: maintenance(&mut d)?, - operation: Box::new(read_maintenance(&mut d)?), - }, + let (kind, receiver_route) = match d.read_u8()? { + 1 => ( + FleetActionKind::Movement { + action: movement(&mut d)?, + attempt: Box::new(read_attempt(&mut d)?), + }, + None, + ), + 2 => ( + FleetActionKind::Maintenance { + action: maintenance(&mut d)?, + operation: Box::new(read_maintenance(&mut d)?), + }, + None, + ), + 3 => ( + FleetActionKind::Movement { + action: movement(&mut d)?, + attempt: Box::new(read_attempt(&mut d)?), + }, + Some(read_receiver_route(&mut d)?), + ), _ => return Err(OperationError::Invalid("unknown fleet action family")), }; d.finish()?; @@ -76,6 +133,7 @@ impl FleetAction { controller_epoch, issued_at_ms, kind, + receiver_route, }; action.validate()?; Ok(action) @@ -123,6 +181,16 @@ impl FleetActionOutcome { e.write_u8(9)?; e.write_bytes(inventory.as_bytes())?; } + FleetOutcome::RolesSettledAt { + inventory, + head_revision, + registry, + } => { + e.write_u8(12)?; + e.write_bytes(inventory.as_bytes())?; + e.write_u64(*head_revision)?; + write_version(&mut e, *registry)?; + } FleetOutcome::Recovered(evidence) => { e.write_u8(11)?; super::recovery::write_recovered(&mut e, evidence)?; @@ -168,6 +236,11 @@ impl FleetActionOutcome { }, 10 => FleetOutcome::Stopped(read_drain_evidence(&mut d)?), 11 => FleetOutcome::Recovered(Box::new(super::recovery::read_recovered(&mut d)?)), + 12 => FleetOutcome::RolesSettledAt { + inventory: Digest::from_bytes(fixed(&mut d)?), + head_revision: d.read_u64()?, + registry: read_version(&mut d)?, + }, _ => return Err(OperationError::Invalid("unknown fleet action outcome")), }; d.finish()?; diff --git a/crates/cellule-runtime/src/fleet/operations/codec/mod.rs b/crates/cellule-runtime/src/fleet/operations/codec/mod.rs index d44a1e5f..3facbc66 100644 --- a/crates/cellule-runtime/src/fleet/operations/codec/mod.rs +++ b/crates/cellule-runtime/src/fleet/operations/codec/mod.rs @@ -15,6 +15,7 @@ mod follower_evacuation; mod inspection; mod maintenance_enrollments; mod reader_evacuation; +mod receiver_recovery; mod recovery; mod registry; mod writer_inventory; @@ -53,6 +54,8 @@ const WRITER_INVENTORY_BASIS: u8 = 26; const MAINTENANCE_ENROLLMENTS: u8 = 27; const MAINTENANCE_ENROLLMENTS_PAGE: u8 = 28; const MAINTENANCE_ENROLLMENTS_BASIS: u8 = 29; +const RECEIVER_RECOVERY_BASIS: u8 = 30; +const RECEIVER_RECOVERY_EVIDENCE: u8 = 31; fn encoder(kind: u8) -> Result { encoder_limited(kind, MAX_RECORD_BYTES) diff --git a/crates/cellule-runtime/src/fleet/operations/codec/receiver_recovery.rs b/crates/cellule-runtime/src/fleet/operations/codec/receiver_recovery.rs new file mode 100644 index 00000000..6661bd41 --- /dev/null +++ b/crates/cellule-runtime/src/fleet/operations/codec/receiver_recovery.rs @@ -0,0 +1,60 @@ +//! Separate envelopes preserve all existing source-recovery record bytes. +use super::*; + +impl ReceiverRecoveryBasis { + /// Encodes bounded receiver recovery input; it grants no takeover authority. + pub fn to_bytes(&self) -> Result> { + self.validate()?; + let mut e = encoder(RECEIVER_RECOVERY_BASIS)?; + e.write_bytes(&self.accepted.to_bytes()?)?; + e.write_bytes( + &self + .control + .encode() + .map_err(|e| OperationError::Control(Box::new(e)))?, + )?; + e.write_i64(self.observed_at_ms)?; + Ok(e.finish()) + } + /// Validates shape only; the trusted journal retains checked provenance. + pub fn from_bytes(bytes: &[u8]) -> Result { + let mut d = decoder(bytes, RECEIVER_RECOVERY_BASIS)?; + let value = Self { + accepted: AcceptedFleetAction::from_bytes(d.read_bytes()?)?, + control: crate::control::Control::decode(d.read_bytes()?) + .map_err(|e| OperationError::Control(Box::new(e)))?, + observed_at_ms: d.read_i64()?, + }; + d.finish()?; + value.validate()?; + Ok(value) + } +} +impl ReceiverRecoveryEvidence { + /// Encodes the canonical recovered receiver position before actor admission. + pub fn to_bytes(&self) -> Result> { + self.validate()?; + let mut e = encoder(RECEIVER_RECOVERY_EVIDENCE)?; + e.write_bytes(&self.basis.to_bytes()?)?; + e.write_bytes( + &self + .restored + .encode() + .map_err(|e| OperationError::Control(Box::new(e)))?, + )?; + e.write_i64(self.recorded_at_ms)?; + Ok(e.finish()) + } + /// Validates transition shape without creating runtime recovery permission. + pub fn from_bytes(bytes: &[u8]) -> Result { + let mut d = decoder(bytes, RECEIVER_RECOVERY_EVIDENCE)?; + let value = Self::new( + ReceiverRecoveryBasis::from_bytes(d.read_bytes()?)?, + crate::control::Control::decode(d.read_bytes()?) + .map_err(|e| OperationError::Control(Box::new(e)))?, + d.read_i64()?, + )?; + d.finish()?; + Ok(value) + } +} diff --git a/crates/cellule-runtime/src/fleet/operations/codec/registry.rs b/crates/cellule-runtime/src/fleet/operations/codec/registry.rs index 1f4ceaab..44d2166f 100644 --- a/crates/cellule-runtime/src/fleet/operations/codec/registry.rs +++ b/crates/cellule-runtime/src/fleet/operations/codec/registry.rs @@ -1,7 +1,7 @@ use super::*; use crate::node::NodeMode; -fn write_version(e: &mut BoundedEncoder, version: RegistryVersion) -> Result<()> { +pub(super) fn write_version(e: &mut BoundedEncoder, version: RegistryVersion) -> Result<()> { write_scope(e, version.scope)?; e.write_u64(version.revision)?; e.write_bool(version.bootstrap_revision.is_some())?; @@ -12,7 +12,7 @@ fn write_version(e: &mut BoundedEncoder, version: RegistryVersion) -> Result<()> Ok(()) } -fn read_version(d: &mut BoundedDecoder<'_>) -> Result { +pub(super) fn read_version(d: &mut BoundedDecoder<'_>) -> Result { let version = RegistryVersion { scope: read_scope(d)?, revision: d.read_u64()?, diff --git a/crates/cellule-runtime/src/fleet/operations/mod.rs b/crates/cellule-runtime/src/fleet/operations/mod.rs index 1581f5fc..2e5d21b6 100644 --- a/crates/cellule-runtime/src/fleet/operations/mod.rs +++ b/crates/cellule-runtime/src/fleet/operations/mod.rs @@ -24,11 +24,13 @@ pub use maintenance_enrollments::{ }; mod journal; mod reader_evacuation; +mod receiver_recovery; mod records; mod recovery; pub use reader_evacuation::{ ReaderEvacuationPage, ReaderEvacuationRecord, ReaderReplacementWitness, }; +pub use receiver_recovery::{ReceiverRecoveryBasis, ReceiverRecoveryEvidence}; mod registry; mod writer_inventory; pub use recovery::{RecoveredActivation, RecoveryBasis, RecoveryEvidence}; @@ -42,6 +44,7 @@ pub use accepted::AcceptedFleetAction; pub use acquisition::AcquisitionBasis; pub use actions::{ FleetAction, FleetActionKind, FleetActionOutcome, FleetOutcome, MaintenanceAction, + ReceiverHandoff, ReceiverRoute, }; pub use attempt::{ ActivationEvidence, AttemptEvent, AttemptPhase, MoveAttempt, MoveAttemptSpec, MovementAction, @@ -73,6 +76,8 @@ pub const MAX_PAGE_ENTRIES: usize = 128; pub const MAX_ACTIVE_ATTEMPTS: usize = 2; /// Hard initial bound on disk demand reserved by unresolved attempts. pub const MAX_RESTORE_BYTES: u64 = 8 * 1024 * 1024 * 1024; +/// Maximum closed receiver boots retained before reconciliation blocks. +pub const MAX_RECEIVER_HANDOFFS: usize = 2; /// Failure of a pure fleet operation transition or its bounded codec. #[derive(Debug, thiserror::Error)] diff --git a/crates/cellule-runtime/src/fleet/operations/receiver_recovery.rs b/crates/cellule-runtime/src/fleet/operations/receiver_recovery.rs new file mode 100644 index 00000000..30ac3ce2 --- /dev/null +++ b/crates/cellule-runtime/src/fleet/operations/receiver_recovery.rs @@ -0,0 +1,209 @@ +//! Recovery of a process-closed receiver after a proven source release. +use super::*; +use crate::control::{Control, ControlState, Transition}; +use crate::node::NodeTakeoverProof; + +/// Exact failed-receiver input retained before canonical takeover. +/// This is separate from failed-source recovery and never erases clean release. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct ReceiverRecoveryBasis { + pub(super) accepted: AcceptedFleetAction, + pub(super) control: Control, + pub(super) observed_at_ms: i64, +} + +impl ReceiverRecoveryBasis { + /// Requires canonical failed-session proof for a closed boot in this route. + pub fn new( + accepted: AcceptedFleetAction, + control: Control, + takeover: NodeTakeoverProof, + observed_at_ms: i64, + ) -> Result { + if control + .owner + .as_ref() + .is_none_or(|owner| owner.session != takeover.session()) + || takeover.claimant() != accepted.session() + { + return Err(OperationError::Fenced); + } + let basis = Self { + accepted, + control, + observed_at_ms, + }; + basis.validate()?; + Ok(basis) + } + /// Returns the exact accepted routed activation. + pub const fn accepted(&self) -> &AcceptedFleetAction { + &self.accepted + } + /// Returns the failed ownership control read before takeover. + pub const fn control(&self) -> &Control { + &self.control + } + /// Returns the immutable input capture time. + pub const fn observed_at_ms(&self) -> i64 { + self.observed_at_ms + } + + pub(super) fn validate(&self) -> Result<()> { + self.accepted.validate()?; + self.control + .encode() + .map_err(|e| OperationError::Control(Box::new(e)))?; + let FleetActionKind::Movement { + action: MovementAction::Activate, + attempt, + } = self.accepted.action().kind() + else { + return Err(OperationError::Invalid( + "receiver recovery requires activation", + )); + }; + let route = self + .accepted + .action() + .receiver_route() + .ok_or(OperationError::Invalid( + "receiver recovery lacks closed route", + ))?; + let owner = self.control.owner.as_ref().ok_or(OperationError::Invalid( + "receiver recovery lacks failed owner", + ))?; + let released = attempt.released().ok_or(OperationError::Invalid( + "receiver recovery lacks clean release", + ))?; + if !route + .hops + .iter() + .any(|hop| hop.previous().1 == owner.session) + || self.control.cell != attempt.spec().target.cell_id() + || self.control.incarnation != attempt.spec().incarnation + || self.control.epoch <= released.epoch + || !matches!( + self.control.state, + ControlState::Serving | ControlState::Recovering + ) + || self.control.root.is_none() + || self.observed_at_ms < self.accepted.accepted_at_ms() + { + return Err(OperationError::Invalid( + "receiver recovery scope or state mismatch", + )); + } + let position = PublishedPosition { + incarnation: self.control.incarnation, + epoch: self.control.epoch, + root: self + .control + .root + .clone() + .ok_or(OperationError::Invalid("receiver recovery root missing"))?, + }; + if !super::attempt::successor_position(&position, released) { + return Err(OperationError::Invalid( + "receiver recovery precedes release", + )); + } + position.validate() + } +} + +/// Canonical recovered receiver position retained before actor admission. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct ReceiverRecoveryEvidence { + pub(super) basis: ReceiverRecoveryBasis, + pub(super) restored: Control, + pub(super) recorded_at_ms: i64, +} +impl ReceiverRecoveryEvidence { + /// Validates the actual takeover and optional pinned-overlay publication. + pub fn new( + basis: ReceiverRecoveryBasis, + restored: Control, + recorded_at_ms: i64, + ) -> Result { + let evidence = Self { + basis, + restored, + recorded_at_ms, + }; + evidence.validate()?; + Ok(evidence) + } + /// Returns the exact original receiver recovery input. + pub const fn basis(&self) -> &ReceiverRecoveryBasis { + &self.basis + } + /// Returns the canonical pre-activation recovery result. + pub const fn restored(&self) -> &Control { + &self.restored + } + /// Returns the original evidence recording time. + pub const fn recorded_at_ms(&self) -> i64 { + self.recorded_at_ms + } + /// Checks an activation result against this immutable recovery record. + pub fn validate_result(&self, result: &FleetActionOutcome) -> Result<()> { + self.validate()?; + self.basis.accepted.validate_result(result)?; + let FleetOutcome::Activated(serving) = &result.outcome else { + return Err(OperationError::Invalid( + "receiver recovery result is not activated", + )); + }; + let required = PublishedPosition { + incarnation: self.restored.incarnation, + epoch: self.restored.epoch, + root: self + .restored + .root + .clone() + .ok_or(OperationError::Invalid("receiver result lacks root"))?, + }; + if result.observed_at_ms < self.recorded_at_ms + || !super::attempt::successor_position(&serving.position, &required) + { + return Err(OperationError::Invalid( + "activation precedes receiver recovery", + )); + } + Ok(()) + } + pub(super) fn validate(&self) -> Result<()> { + self.basis.validate()?; + self.restored + .encode() + .map_err(|e| OperationError::Control(Box::new(e)))?; + let owner = self + .restored + .owner + .clone() + .ok_or(OperationError::Invalid("receiver recovery lacks claimant"))?; + if owner.session != self.basis.accepted.session() + || self.recorded_at_ms < self.basis.observed_at_ms + { + return Err(OperationError::Invalid( + "receiver recovery claimant or time mismatch", + )); + } + let claimed = self + .basis + .control + .takeover(owner) + .map_err(|e| OperationError::Control(Box::new(e)))?; + if claimed.recovery.is_some() { + claimed + .validate_transition(&self.restored, Transition::PublishRecovery) + .map_err(|e| OperationError::Control(Box::new(e)))?; + } else if claimed != self.restored { + return Err(OperationError::Invalid( + "receiver recovery changed root without overlay", + )); + } + Ok(()) + } +} diff --git a/crates/cellule-runtime/src/fleet/operations/records.rs b/crates/cellule-runtime/src/fleet/operations/records.rs index 4c16c07e..2ad07e47 100644 --- a/crates/cellule-runtime/src/fleet/operations/records.rs +++ b/crates/cellule-runtime/src/fleet/operations/records.rs @@ -323,6 +323,8 @@ impl MaintenanceOperation { && (evidence.node != self.node || evidence.session != self.session || !evidence.ready_to_close() + || (self.phase == MaintenancePhase::Closing + && (evidence.facilities_closed || evidence.stopped || evidence.withdrawn)) || (self.phase == MaintenancePhase::Completed && !(evidence.facilities_closed && evidence.stopped && evidence.withdrawn))) { @@ -350,6 +352,9 @@ impl MaintenanceOperation { MaintenanceEvent::ReadyToClose(e) if matches(e) && e.ready_to_close() + && !e.facilities_closed + && !e.stopped + && !e.withdrawn && matches!( self.phase, MaintenancePhase::Evacuating | MaintenancePhase::Closing diff --git a/crates/cellule-runtime/src/fleet/operations/tests/contracts.rs b/crates/cellule-runtime/src/fleet/operations/tests/contracts.rs index a9b1443d..3d4604c3 100644 --- a/crates/cellule-runtime/src/fleet/operations/tests/contracts.rs +++ b/crates/cellule-runtime/src/fleet/operations/tests/contracts.rs @@ -167,7 +167,7 @@ fn node_intent_retains_cordon_until_exact_completed_revision_and_new_boot() { .apply(MaintenanceEvent::BeginEvacuation, 0) .unwrap(); operation - .apply(MaintenanceEvent::ReadyToClose(drain_evidence()), 0) + .apply(MaintenanceEvent::ReadyToClose(closing_evidence()), 0) .unwrap(); operation .apply(MaintenanceEvent::Stopped(drain_evidence()), 0) @@ -414,9 +414,45 @@ fn maintenance_actions_and_results_require_exact_phase_and_terminal_proof() { .to_bytes() .is_err() ); + let registry = RegistryVersion::new(settle.scope()) + .unwrap() + .bootstrap(0) + .unwrap(); + let settled = reply( + &settle, + true, + FleetOutcome::RolesSettledAt { + inventory: Digest::from_bytes([88; 32]), + head_revision: settle.journal_revision(), + registry, + }, + ); + assert_eq!( + FleetActionOutcome::from_bytes(&settled.to_bytes().unwrap()).unwrap(), + settled + ); + assert!( + reply( + &settle, + true, + FleetOutcome::RolesSettledAt { + inventory: Digest::from_bytes([88; 32]), + head_revision: settle.journal_revision(), + registry: RegistryVersion::new(FleetScope { + fleet: Digest::from_bytes([17; 32]), + application: ApplicationId::from_bytes([18; 16]), + }) + .unwrap() + .bootstrap(0) + .unwrap(), + }, + ) + .to_bytes() + .is_err() + ); let head = transition( &head, - JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(drain_evidence())), + JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(closing_evidence())), ); let finalize = head .maintenance_action(MaintenanceAction::Finalize, 0) @@ -441,6 +477,32 @@ fn maintenance_actions_and_results_require_exact_phase_and_terminal_proof() { ); } +#[test] +fn ready_to_close_cannot_claim_host_shutdown_or_withdrawal_before_finalize() { + let head = transition(&head(), JournalTransition::BeginMaintenance(maintenance())); + let head = transition( + &head, + JournalTransition::Maintenance(MaintenanceEvent::Cordoned), + ); + let head = transition( + &head, + JournalTransition::Maintenance(MaintenanceEvent::BeginEvacuation), + ); + let mut evidence = drain_evidence(); + evidence.facilities_closed = true; + let epoch = head.controller().unwrap().epoch; + assert!( + head.transition( + FleetProfile::default(), + head.revision(), + epoch, + 0, + JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(evidence)), + ) + .is_err() + ); +} + #[test] fn action_and_intent_codecs_reject_truncation_extra_bytes_and_other_record_families() { let head = transition(&head(), JournalTransition::BeginMaintenance(maintenance())); @@ -473,3 +535,127 @@ fn action_and_intent_codecs_reject_truncation_extra_bytes_and_other_record_famil assert!(FleetAction::from_bytes(&intent.to_bytes().unwrap()).is_err()); assert!(NodeIntent::from_bytes(&action.to_bytes().unwrap()).is_err()); } + +#[test] +fn receiver_continuation_is_registry_bound_endpoint_bound_and_versioned() { + let original = spec(1); + let head = transition(&head(), JournalTransition::Allocate(original.clone())); + let head = attempt(&head, original.id, AttemptEvent::BeginPrepare); + let head = attempt( + &head, + original.id, + AttemptEvent::Reserved(ReceiverReservation { + session: original.destination, + expires_at_ms: 5_000, + }), + ); + let head = attempt(&head, original.id, AttemptEvent::BeginRelease); + let head = attempt(&head, original.id, AttemptEvent::Released(release())); + let head = attempt(&head, original.id, AttemptEvent::BeginActivate); + let registry = RegistryVersion::new(head.scope()) + .unwrap() + .bootstrap(0) + .unwrap(); + let first = ReceiverRoute::begin( + head.scope(), + &original, + NodeId::from_bytes([7; 16]), + SessionId::from_bytes([77; 16]), + Digest::from_bytes([31; 32]), + registry, + ) + .unwrap(); + let next_registry = registry.advance(registry.revision()).unwrap(); + let route = first + .extend( + head.scope(), + &original, + NodeId::from_bytes([8; 16]), + SessionId::from_bytes([78; 16]), + Digest::from_bytes([32; 32]), + next_registry, + ) + .unwrap(); + assert_eq!( + route.target(&original), + (NodeId::from_bytes([8; 16]), SessionId::from_bytes([78; 16])) + ); + assert!( + route + .extend( + head.scope(), + &original, + NodeId::from_bytes([9; 16]), + SessionId::from_bytes([79; 16]), + Digest::from_bytes([33; 32]), + next_registry.advance(next_registry.revision()).unwrap(), + ) + .is_err() + ); + + let action = head + .movement_action_with_receiver_route( + original.id, + MovementAction::Activate, + route.clone(), + 0, + ) + .unwrap(); + assert!(action.authorize_against(&head, 1).is_err()); + action + .authorize_against_registry(&head, next_registry, 1) + .unwrap(); + assert!( + action + .authorize_against_registry(&head, registry, 1) + .is_err() + ); + let unknown_head = attempt(&head, original.id, AttemptEvent::OutcomeUnknown); + let adopted = unknown_head + .movement_action_with_receiver_route(original.id, MovementAction::Activate, route, 0) + .unwrap(); + adopted + .authorize_against_registry(&unknown_head, next_registry, 1) + .unwrap(); + assert!( + action + .validate_endpoint(NodeId::from_bytes([2; 16]), original.destination) + .is_err() + ); + assert!( + action + .validate_endpoint(NodeId::from_bytes([8; 16]), SessionId::from_bytes([78; 16])) + .is_ok() + ); + assert_eq!( + FleetAction::from_bytes(&action.to_bytes().unwrap()).unwrap(), + action + ); + + let accepted = AcceptedFleetAction::new_with_registry( + action.clone(), + &head, + next_registry, + NodeId::from_bytes([8; 16]), + SessionId::from_bytes([78; 16]), + 1, + ) + .unwrap(); + let result = FleetActionOutcome { + scope: head.scope(), + action_key: action.key().unwrap(), + node: NodeId::from_bytes([8; 16]), + session: SessionId::from_bytes([78; 16]), + observed_at_ms: 2, + outcome: FleetOutcome::Activated(ActivationEvidence { + node: NodeId::from_bytes([8; 16]), + session: SessionId::from_bytes([78; 16]), + position: PublishedPosition { + incarnation: original.incarnation, + epoch: original.source_epoch + 1, + root: release().root, + }, + }), + }; + accepted.validate_result(&result).unwrap(); +} diff --git a/crates/cellule-runtime/src/fleet/operations/tests/mod.rs b/crates/cellule-runtime/src/fleet/operations/tests/mod.rs index 48146c71..10c363da 100644 --- a/crates/cellule-runtime/src/fleet/operations/tests/mod.rs +++ b/crates/cellule-runtime/src/fleet/operations/tests/mod.rs @@ -11,6 +11,7 @@ mod cleanup; mod contracts; mod inspection; mod maintenance_release; +mod receiver_recovery; mod recovery; mod registry; @@ -500,6 +501,15 @@ fn drain_evidence() -> DrainEvidence { } } +fn closing_evidence() -> DrainEvidence { + DrainEvidence { + facilities_closed: false, + stopped: false, + withdrawn: false, + ..drain_evidence() + } +} + #[test] fn maintenance_relocation_foreign_tails_and_shutdown_all_gate_completion() { let head = transition(&head(), JournalTransition::BeginMaintenance(maintenance())); @@ -546,7 +556,7 @@ fn maintenance_relocation_foreign_tails_and_shutdown_all_gate_completion() { } let head = transition( &head, - JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(drain_evidence())), + JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(closing_evidence())), ); let missing_withdrawal = DrainEvidence { withdrawn: false, @@ -671,7 +681,7 @@ fn maintenance_idempotency_and_unresolved_attempts_prevent_false_close() { head.revision(), 1, 0, - JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(drain_evidence())) + JournalTransition::Maintenance(MaintenanceEvent::ReadyToClose(closing_evidence())) ) .is_err() ); diff --git a/crates/cellule-runtime/src/fleet/operations/tests/receiver_recovery.rs b/crates/cellule-runtime/src/fleet/operations/tests/receiver_recovery.rs new file mode 100644 index 00000000..66602f6e --- /dev/null +++ b/crates/cellule-runtime/src/fleet/operations/tests/receiver_recovery.rs @@ -0,0 +1,156 @@ +use super::*; +use crate::control::{Control, ControlState, Owner}; + +fn evidence() -> ReceiverRecoveryEvidence { + let (head, id) = reserved(); + let head = attempt(&head, id, AttemptEvent::BeginRelease); + let head = attempt(&head, id, AttemptEvent::Released(release())); + let head = attempt(&head, id, AttemptEvent::BeginActivate); + let spec = spec(1); + let registry = RegistryVersion::new(head.scope()) + .unwrap() + .bootstrap(0) + .unwrap(); + let node = NodeId::from_bytes([7; 16]); + let session = SessionId::from_bytes([77; 16]); + let route = ReceiverRoute::begin( + head.scope(), + &spec, + node, + session, + Digest::from_bytes([31; 32]), + registry, + ) + .unwrap(); + let action = head + .movement_action_with_receiver_route(id, MovementAction::Activate, route, 0) + .unwrap(); + let accepted = + AcceptedFleetAction::new_with_registry(action, &head, registry, node, session, 0).unwrap(); + // Unit records check shape. Real NodeDirectory capability construction, + // persistence and observed takeover are covered by minion's public cases. + let mut control = Control::initial( + spec.target.cell_id(), + spec.incarnation, + Owner { + session: spec.destination, + endpoint: "https://receiver.internal".into(), + }, + Digest::from_bytes([35; 32]), + 1, + ) + .unwrap(); + control.epoch = release().epoch + 1; + control.revision = 9; + control.progress = 9; + control.root = Some(release().root); + control.state = ControlState::Serving; + let basis = ReceiverRecoveryBasis { + accepted, + control, + observed_at_ms: 0, + }; + basis.validate().unwrap(); + let restored = basis + .control + .takeover(Owner { + session, + endpoint: "https://successor.internal".into(), + }) + .unwrap(); + ReceiverRecoveryEvidence::new(basis, restored, 0).unwrap() +} + +#[test] +fn receiver_recovery_requires_exact_closed_owner_and_clean_release() { + let original = evidence(); + for mutation in 0..7 { + let mut basis = original.basis.clone(); + match mutation { + 0 => basis.control.owner.as_mut().unwrap().session = spec(1).source, + 1 => basis.control.owner.as_mut().unwrap().session = basis.accepted.session(), + 2 => basis.control.epoch = release().epoch, + 3 => basis.control.incarnation = IncarnationId::from_bytes([99; 16]), + 4 => basis.control.cell = spec(2).target.cell_id(), + 5 => basis.control.state = ControlState::Idle, + _ => basis.observed_at_ms = -1, + } + assert!(basis.to_bytes().is_err(), "mutation {mutation}"); + } + for mutation in 0..6 { + let mut restored = original.restored.clone(); + match mutation { + 0 => restored.epoch += 1, + 1 => restored.revision += 1, + 2 => restored.root.as_mut().unwrap().digest = Digest::from_bytes([99; 32]), + 3 => restored.owner.as_mut().unwrap().session = spec(1).destination, + 4 => restored.state = ControlState::Serving, + _ => restored.progress += 1, + } + assert!(ReceiverRecoveryEvidence::new(original.basis.clone(), restored, 0).is_err()); + } + let accepted = original.basis.accepted(); + let result = FleetActionOutcome { + scope: accepted.action().scope(), + action_key: accepted.action().key().unwrap(), + node: accepted.node(), + session: accepted.session(), + observed_at_ms: 0, + outcome: FleetOutcome::Activated(ActivationEvidence { + node: accepted.node(), + session: accepted.session(), + position: PublishedPosition { + incarnation: original.restored.incarnation, + epoch: original.restored.epoch, + root: original.restored.root.clone().unwrap(), + }, + }), + }; + original.validate_result(&result).unwrap(); + let mut changed = result; + let FleetOutcome::Activated(ref mut serving) = changed.outcome else { + unreachable!() + }; + serving.position.root.digest = Digest::from_bytes([99; 32]); + assert!(original.validate_result(&changed).is_err()); +} + +#[test] +fn receiver_recovery_records_reject_incomplete_wrong_version_kind_and_oversize() { + let evidence = evidence(); + let basis = evidence.basis(); + let basis_bytes = basis.to_bytes().unwrap(); + let evidence_bytes = evidence.to_bytes().unwrap(); + assert_eq!( + ReceiverRecoveryBasis::from_bytes(&basis_bytes).unwrap(), + *basis + ); + assert_eq!( + ReceiverRecoveryEvidence::from_bytes(&evidence_bytes).unwrap(), + evidence + ); + for end in 0..basis_bytes.len() { + assert!(ReceiverRecoveryBasis::from_bytes(&basis_bytes[..end]).is_err()); + } + for end in 0..evidence_bytes.len() { + assert!(ReceiverRecoveryEvidence::from_bytes(&evidence_bytes[..end]).is_err()); + } + let mut trailing = basis_bytes.clone(); + trailing.push(0); + assert!(ReceiverRecoveryBasis::from_bytes(&trailing).is_err()); + let mut trailing = evidence_bytes.clone(); + trailing.push(0); + assert!(ReceiverRecoveryEvidence::from_bytes(&trailing).is_err()); + for bytes in [&basis_bytes, &evidence_bytes] { + let mut version = bytes.clone(); + version[4 + b"cellule.fleet-operation\0".len()] = 255; + assert!(ReceiverRecoveryBasis::from_bytes(&version).is_err()); + assert!(ReceiverRecoveryEvidence::from_bytes(&version).is_err()); + } + assert!(ReceiverRecoveryBasis::from_bytes(&evidence_bytes).is_err()); + assert!(ReceiverRecoveryEvidence::from_bytes(&basis_bytes).is_err()); + assert!(RecoveryBasis::from_bytes(&basis_bytes).is_err()); + assert!(RecoveryEvidence::from_bytes(&evidence_bytes).is_err()); + assert!(ReceiverRecoveryBasis::from_bytes(&vec![0; MAX_RECORD_BYTES as usize + 1]).is_err()); + assert!(ReceiverRecoveryEvidence::from_bytes(&vec![0; MAX_RECORD_BYTES as usize + 1]).is_err()); +} diff --git a/crates/cellule-runtime/src/fleet/operations/tests/registry.rs b/crates/cellule-runtime/src/fleet/operations/tests/registry.rs index 2017b26f..a4efbb77 100644 --- a/crates/cellule-runtime/src/fleet/operations/tests/registry.rs +++ b/crates/cellule-runtime/src/fleet/operations/tests/registry.rs @@ -277,7 +277,7 @@ fn source_maintenance_phase_fences_new_roles_without_invalidating_replay() { .unwrap(); assert!(admit(&operation).is_ok()); operation - .apply(MaintenanceEvent::ReadyToClose(drain_evidence()), 3) + .apply(MaintenanceEvent::ReadyToClose(closing_evidence()), 3) .unwrap(); assert!(matches!(admit(&operation), Err(OperationError::Conflict))); assert_eq!(source.advance_maintenance(&operation).unwrap(), source); diff --git a/crates/cellule-runtime/src/node/advertisement/mod.rs b/crates/cellule-runtime/src/node/advertisement/mod.rs index f94b6b7a..81c7d400 100644 --- a/crates/cellule-runtime/src/node/advertisement/mod.rs +++ b/crates/cellule-runtime/src/node/advertisement/mod.rs @@ -291,10 +291,10 @@ impl NodeAdvertisement { self.canonical_bytes() } - // Both callers first verify this immutable value. Re-encoding for canonical - // byte equality needs the same serializer, without a second signature pass. - // Storage producers still enter through encode and verify before emission. - fn canonical_bytes(&self) -> Result> { + // Callers first verify this immutable value. Decoding and directory + // create/refresh use the same serializer after authentication; no field may + // change between that check and emission. Unverified producers use encode. + pub(super) fn canonical_bytes(&self) -> Result> { let encoded = serde_json::to_vec(&RawAdvertisement::from(self))?; if encoded.len() as u64 > MAX_NODE_BYTES { return Err(Error::Node("advertisement exceeds 64 KiB")); diff --git a/crates/cellule-runtime/src/node/directory/advertisement.rs b/crates/cellule-runtime/src/node/directory/advertisement.rs index 7aa98cad..8d23ff87 100644 --- a/crates/cellule-runtime/src/node/directory/advertisement.rs +++ b/crates/cellule-runtime/src/node/directory/advertisement.rs @@ -376,16 +376,18 @@ impl NodeDirectory { return Ok(None); }; validate_record_path(&self.layout, advertisement.session, &meta.location)?; + // Canonical decoding above verifies the signatures on each + // fresh body. These policies neither mutate nor reuse it. match scan { AdvertisementScan::LiveRelease => { if advertisement.expires_at_ms <= now_ms { return Ok(None); } - self.validate(&advertisement, now_ms)?; + advertisement.validate_at(now_ms)?; + self.validate_scope(&advertisement)?; } AdvertisementScan::AdvertisedFleet => { advertisement.validate_shape()?; - advertisement.verify_signature()?; if advertisement.fleet != self.fleet || advertisement.issued_at_ms > now_ms.saturating_add(MAX_CLOCK_SKEW_MS) @@ -668,11 +670,13 @@ impl NodeDirectory { candidate.log.clone_from(&base.advertisement.log); self.validate(&candidate, now_ms)?; validate_successor(&base.advertisement, &candidate)?; + // The validated candidate remains immutable through its CAS body. + let encoded = candidate.canonical_bytes()?; let path = self.layout.node_path(candidate.session.as_bytes()); match self .layout .store() - .update(&path, Bytes::from(candidate.encode()?), base.token.clone()) + .update(&path, Bytes::from(encoded), base.token.clone()) .await { Ok(token) => { diff --git a/crates/cellule-runtime/src/node/directory/enrollment.rs b/crates/cellule-runtime/src/node/directory/enrollment.rs index a434c17b..a10cd00f 100644 --- a/crates/cellule-runtime/src/node/directory/enrollment.rs +++ b/crates/cellule-runtime/src/node/directory/enrollment.rs @@ -138,8 +138,13 @@ fn enrollment_digest( ) -> Result { let mut digest = blake3::Hasher::new(); digest.update(domain); - digest.update(&prepared.source.encode()?); - digest.update(&source.advertisement.encode()?); + // These private, immutable proof inputs were authenticated by preparation, + // canonical inspection, or the checked conditional write. Historical + // fingerprints serialize those exact bytes; every fresh provider read still + // authenticates independently. Re-verifying here repeated all boot signatures + // for each producer page without observing any new authority. + digest.update(&prepared.source.canonical_bytes()?); + digest.update(&source.advertisement.canonical_bytes()?); // Preserve the immutable conditional-write token as well as signed bytes. // Lengths and option markers keep arbitrary provider tokens unambiguous. for token in [&source.token.e_tag, &source.token.version] { @@ -156,7 +161,7 @@ fn enrollment_digest( } digest.update(&prepared.log.epoch().to_le_bytes()); for follower in &prepared.followers { - digest.update(&follower.encode()?); + digest.update(&follower.canonical_bytes()?); } Ok(Digest::from_bytes(*digest.finalize().as_bytes())) } diff --git a/crates/cellule-runtime/src/node/directory/inventory/batch.rs b/crates/cellule-runtime/src/node/directory/inventory/batch.rs new file mode 100644 index 00000000..226d1986 --- /dev/null +++ b/crates/cellule-runtime/src/node/directory/inventory/batch.rs @@ -0,0 +1,174 @@ +//! One fresh canonical traversal feeds independently bound follower windows. +use super::*; +use std::collections::BTreeMap; + +pub(super) async fn collect( + directory: &NodeDirectory, + requests: &[(NodeId, Option)], + limit: usize, + now_ms: i64, +) -> Result> { + let mut windows = requests + .iter() + .map(|(member, cursor)| Window::new(*member, *cursor)) + .collect::>(); + let prefix = directory.layout.node_directory_path(); + let mut objects = directory.layout.store().inner().list(Some(&prefix)); + let mut seen = HashSet::with_capacity(MAX_LIVE_NODE_RECORDS); + while let Some(object) = objects.next().await { + if seen.len() == MAX_LIVE_NODE_RECORDS { + return Err(Error::Capacity("follower log inventory record bound")); + } + let meta = object.map_err(|error| map_object_store_error(error, prefix.as_ref()))?; + let Some((record, _)) = directory.load_record_at(&meta.location).await? else { + return Err(Error::Node("log inventory record changed during scan")); + }; + let session = record.session(); + validate_record_path(&directory.layout, session, &meta.location)?; + if !seen.insert(session) { + return Err(Error::Node("duplicate log inventory session")); + } + let (node, leader_state) = match &record { + NodeRecord::Advertisement(advertisement) => { + // Authenticate every fresh body before filtering any follower, + // including expired and foreign-release log obligations. + advertisement.validate_shape()?; + if advertisement.fleet != directory.fleet + || advertisement.issued_at_ms > now_ms.saturating_add(MAX_CLOCK_SKEW_MS) + { + return Err(Error::Node("log inventory fleet or issue time differs")); + } + ( + advertisement.node, + if advertisement.expires_at_ms > now_ms { + LogLeaderState::Live + } else { + LogLeaderState::Expired + }, + ) + } + NodeRecord::Tombstone(tombstone) => (tombstone.node, LogLeaderState::Fenced), + }; + if let Some(log) = record.log() { + for window in &mut windows { + window.accept(session, node, leader_state, log, limit); + } + } + } + // A malformed cursor in any window rejects the entire traversal; no caller + // can consume a successful prefix as proof that the full request completed. + windows + .into_iter() + .map(|window| window.finish(directory.inventory_scope, limit, now_ms)) + .collect() +} + +struct Window { + member: NodeId, + cursor: Option, + entries: BTreeMap<[u8; 16], FollowerLogObservation>, + combined: [u8; 32], + total_logs: usize, + cursor_found: bool, +} +impl Window { + fn new(member: NodeId, cursor: Option) -> Self { + Self { + member, + cursor, + entries: BTreeMap::new(), + combined: [0; 32], + total_logs: 0, + cursor_found: cursor.is_none(), + } + } + fn accept( + &mut self, + session: SessionId, + node: NodeId, + state: LogLeaderState, + log: &NodeLogStatus, + limit: usize, + ) { + if !log.members().contains(&self.member) { + return; + } + self.total_logs += 1; + // Preserve the existing commutative topology digest byte-for-byte. + // Coverage/heartbeat changes remain visible in exact recheck rows. + let mut hash = blake3::Hasher::new(); + hash.update(session.as_bytes()); + hash.update(node.as_bytes()); + hash.update(&log.epoch().to_le_bytes()); + hash.update(&[match log.phase() { + NodeLogPhase::Open => 1, + NodeLogPhase::Recovering => 2, + NodeLogPhase::Sealed => 3, + NodeLogPhase::Retired => 4, + }]); + for member in log.members() { + hash.update(member.as_bytes()); + } + for (combined, byte) in self.combined.iter_mut().zip(hash.finalize().as_bytes()) { + *combined ^= byte; + } + if self.cursor.is_some_and(|cursor| cursor.after == session) { + self.cursor_found = true; + } + if self + .cursor + .is_some_and(|cursor| session.as_bytes() <= cursor.after.as_bytes()) + { + return; + } + self.entries.insert( + *session.as_bytes(), + FollowerLogObservation { + leader: session, + leader_node: node, + leader_state: state, + log: log.clone(), + }, + ); + if self.entries.len() > limit + 1 { + self.entries.pop_last(); + } + } + fn finish(mut self, scope: [u8; 16], limit: usize, now_ms: i64) -> Result { + let mut hash = blake3::Hasher::new(); + hash.update(b"cellule-authoritative-log-inventory-v1"); + hash.update(&scope); + hash.update(self.member.as_bytes()); + hash.update(&(self.total_logs as u64).to_le_bytes()); + hash.update(&self.combined); + let topology = Digest::from_bytes(*hash.finalize().as_bytes()); + if !self.cursor_found + || self + .cursor + .is_some_and(|cursor| cursor.topology != topology) + { + return Err(Error::Node("log inventory topology changed; restart scan")); + } + let more = self.entries.len() > limit; + if more { + self.entries.pop_last(); + } + let entries = self.entries.into_values().collect::>(); + let next = if more { + entries.last().map(|last| LogInventoryCursor { + topology, + after: last.leader, + }) + } else { + None + }; + Ok(LogInventoryPage { + member: self.member, + topology, + observed_at_ms: now_ms, + total_logs: self.total_logs, + entries, + next, + }) + } +} diff --git a/crates/cellule-runtime/src/node/directory/inventory/mod.rs b/crates/cellule-runtime/src/node/directory/inventory/mod.rs index c7699645..9e4bcb60 100644 --- a/crates/cellule-runtime/src/node/directory/inventory/mod.rs +++ b/crates/cellule-runtime/src/node/directory/inventory/mod.rs @@ -1,6 +1,8 @@ //! Bounded discovery of authoritative log references, including failed owners. -use std::collections::{BTreeMap, HashSet}; +use std::collections::HashSet; + +mod batch; use super::*; @@ -109,6 +111,35 @@ impl LogInventoryPage { } impl NodeDirectory { + /// Traverses the directory once for bounded physical-follower windows. + /// Returns pages in request order with the same cursor/digest contract as + /// `follower_logs_page`. The sum of requested page limits is at most 128; + /// each window retains at most one additional lookahead row. Any invalid + /// record or cursor rejects the entire batch. This is interval evidence, + /// not atomic membership or permission to finalize a node. + pub async fn follower_logs_pages( + &self, + requests: &[(NodeId, Option)], + limit: usize, + now_ms: i64, + ) -> Result> { + if requests.is_empty() + || requests.len() > MAX_PAGE_ENTRIES + || limit == 0 + || limit > MAX_PAGE_ENTRIES / requests.len() + || now_ms < 0 + { + return Err(Error::Node("invalid follower log inventory batch bounds")); + } + let mut members = HashSet::new(); + for (member, _) in requests { + if member.as_bytes() == &[0; 16] || !members.insert(*member) { + return Err(Error::Node("invalid follower log inventory batch member")); + } + } + batch::collect(self, requests, limit, now_ms).await + } + /// Discovers current log references to a follower, including expired records /// and fenced/recovering tombstones that `live` intentionally omits. /// @@ -127,120 +158,11 @@ impl NodeDirectory { if member.as_bytes() == &[0; 16] || !(1..=MAX_PAGE_ENTRIES).contains(&limit) || now_ms < 0 { return Err(Error::Node("invalid follower log inventory bounds")); } - let prefix = self.layout.node_directory_path(); - let mut objects = self.layout.store().inner().list(Some(&prefix)); - let mut seen = HashSet::with_capacity(MAX_LIVE_NODE_RECORDS); - let mut window = BTreeMap::new(); - let mut combined = [0_u8; 32]; - let mut total_logs = 0; - let mut cursor_found = cursor.is_none(); - while let Some(object) = objects.next().await { - if seen.len() == MAX_LIVE_NODE_RECORDS { - return Err(Error::Capacity("follower log inventory record bound")); - } - let meta = object.map_err(|error| map_object_store_error(error, prefix.as_ref()))?; - let Some((record, _)) = self.load_record_at(&meta.location).await? else { - // A disappearing record does not establish a complete scan. - return Err(Error::Node("log inventory record changed during scan")); - }; - let session = record.session(); - validate_record_path(&self.layout, session, &meta.location)?; - if !seen.insert(session) { - return Err(Error::Node("duplicate log inventory session")); - } - let (node, leader_state) = match &record { - NodeRecord::Advertisement(advertisement) => { - advertisement.validate_shape()?; - advertisement.verify_signature()?; - if advertisement.fleet != self.fleet - || advertisement.issued_at_ms > now_ms.saturating_add(MAX_CLOCK_SKEW_MS) - { - return Err(Error::Node("log inventory fleet or issue time differs")); - } - ( - advertisement.node, - if advertisement.expires_at_ms > now_ms { - LogLeaderState::Live - } else { - LogLeaderState::Expired - }, - ) - } - NodeRecord::Tombstone(tombstone) => (tombstone.node, LogLeaderState::Fenced), - }; - let Some(log) = record.log().filter(|log| log.members().contains(&member)) else { - continue; - }; - total_logs += 1; - // Commutative digest makes arbitrary object listing order harmless. - // Duplicate sessions are rejected separately. Volatile coverage and - // heartbeat times do not reset topology; exact actions recheck them. - let mut hash = blake3::Hasher::new(); - hash.update(session.as_bytes()); - hash.update(node.as_bytes()); - hash.update(&log.epoch().to_le_bytes()); - hash.update(&[match log.phase() { - NodeLogPhase::Open => 1, - NodeLogPhase::Recovering => 2, - NodeLogPhase::Sealed => 3, - NodeLogPhase::Retired => 4, - }]); - for enrolled in log.members() { - hash.update(enrolled.as_bytes()); - } - for (combined, byte) in combined.iter_mut().zip(hash.finalize().as_bytes()) { - *combined ^= byte; - } - if cursor.is_some_and(|cursor| cursor.after == session) { - cursor_found = true; - } - if cursor.is_some_and(|cursor| session.as_bytes() <= cursor.after.as_bytes()) { - continue; - } - window.insert( - *session.as_bytes(), - FollowerLogObservation { - leader: session, - leader_node: node, - leader_state, - log: log.clone(), - }, - ); - if window.len() > limit + 1 { - window.pop_last(); - } - } - let mut hash = blake3::Hasher::new(); - hash.update(b"cellule-authoritative-log-inventory-v1"); - hash.update(&self.inventory_scope); - hash.update(member.as_bytes()); - hash.update(&(total_logs as u64).to_le_bytes()); - hash.update(&combined); - let topology = Digest::from_bytes(*hash.finalize().as_bytes()); - if !cursor_found || cursor.is_some_and(|cursor| cursor.topology != topology) { - return Err(Error::Node("log inventory topology changed; restart scan")); - } - let more = window.len() > limit; - if more { - window.pop_last(); - } - let entries = window.into_values().collect::>(); - let next = if more { - entries.last().map(|last| LogInventoryCursor { - topology, - after: last.leader, - }) - } else { - None - }; - Ok(LogInventoryPage { - member, - topology, - observed_at_ms: now_ms, - total_logs, - entries, - next, - }) + self.follower_logs_pages(&[(member, cursor)], limit, now_ms) + .await? + .into_iter() + .next() + .ok_or(Error::Node("follower log inventory page missing")) } } diff --git a/crates/cellule-runtime/src/node/directory/inventory/tests.rs b/crates/cellule-runtime/src/node/directory/inventory/tests.rs index f921317b..c6993980 100644 --- a/crates/cellule-runtime/src/node/directory/inventory/tests.rs +++ b/crates/cellule-runtime/src/node/directory/inventory/tests.rs @@ -5,6 +5,179 @@ use object_store::{memory::InMemory, path::Path}; const NOW: i64 = 1_000_000; +#[tokio::test] +async fn shared_windows_keep_existing_topology_domain_and_cursor_bytes() { + let mut directory = directory(); + directory.inventory_scope = [42; 16]; + for id in [3, 1, 2] { + enroll(&directory, id, NOW).await; + } + // Independent encoding of the existing v1 domain: session, physical owner, + // little-endian epoch, Open tag, complete ensemble; XOR then scoped count. + let mut combined = [0; 32]; + for id in 1..=3 { + let mut row = Vec::new(); + row.extend_from_slice(&[id; 16]); + row.extend_from_slice(&[id; 16]); + row.extend_from_slice(&4u64.to_le_bytes()); + row.push(1); + row.extend_from_slice(&[9; 16]); + for (byte, value) in combined.iter_mut().zip(blake3::hash(&row).as_bytes()) { + *byte ^= value; + } + } + let mut domain = b"cellule-authoritative-log-inventory-v1".to_vec(); + domain.extend_from_slice(&[42; 16]); + domain.extend_from_slice(&[9; 16]); + domain.extend_from_slice(&3u64.to_le_bytes()); + domain.extend_from_slice(&combined); + let expected = *blake3::hash(&domain).as_bytes(); + let pages = directory + .follower_logs_pages( + &[(NodeId::from_bytes([8; 16]), None), (member(), None)], + 1, + NOW + 2, + ) + .await + .unwrap(); + assert_eq!(pages[1].topology().as_bytes(), &expected); + let mut cursor = [0; 48]; + cursor[..32].copy_from_slice(&expected); + cursor[32..].copy_from_slice(&[1; 16]); + assert_eq!(pages[1].next().unwrap().to_bytes(), cursor); + let next = directory + .follower_logs_page( + member(), + Some(LogInventoryCursor::from_bytes(&cursor).unwrap()), + 128, + NOW + 3, + ) + .await + .unwrap(); + assert_eq!(next.total_logs(), 3); + assert_eq!(next.entries().len(), 2); + assert!(next.next().is_none()); +} + +#[tokio::test] +async fn shared_windows_preserve_member_cursors_and_exact_expired_rows() { + let directory = directory(); + for id in [3, 1, 2] { + enroll(&directory, id, NOW).await; + } + let other = NodeId::from_bytes([8; 16]); + let requests = [(other, None), (member(), None)]; + let pages = directory + .follower_logs_pages(&requests, 1, NOW + 2) + .await + .unwrap(); + assert_eq!(pages[0].member(), other); + assert_eq!(pages[0].total_logs(), 0); + assert!(pages[0].entries().is_empty()); + let single = directory + .follower_logs_page(member(), None, 1, NOW + 2) + .await + .unwrap(); + assert_eq!(pages[1].topology(), single.topology()); + assert_eq!(pages[1].entries(), single.entries()); + assert_eq!(pages[1].total_logs(), single.total_logs()); + let cursor = pages[1].next().unwrap(); + assert_eq!(cursor.to_bytes(), single.next().unwrap().to_bytes()); + let decoded = LogInventoryCursor::from_bytes(&cursor.to_bytes()).unwrap(); + let pages = directory + .follower_logs_pages(&[(other, None), (member(), Some(decoded))], 2, NOW + 20_000) + .await + .unwrap(); + assert!( + pages[1] + .entries() + .iter() + .all(|row| row.leader_state == LogLeaderState::Expired) + ); + assert_eq!( + pages[1] + .entries() + .iter() + .map(|row| row.leader) + .collect::>(), + [ + SessionId::from_bytes([2; 16]), + SessionId::from_bytes([3; 16]) + ] + ); + assert!(pages[1].next().is_none()); + assert!( + directory + .follower_logs_pages(&[(member(), None), (other, Some(cursor))], 1, NOW + 3) + .await + .is_err() + ); + let second = directory + .load(SessionId::from_bytes([2; 16]), NOW + 3) + .await + .unwrap() + .unwrap(); + directory + .advance_log_coverage(&second, 27, NOW + 4) + .await + .unwrap(); + let pages = directory + .follower_logs_pages(&[(other, None), (member(), Some(cursor))], 2, NOW + 5) + .await + .unwrap(); + assert_eq!(pages[1].topology(), single.topology()); + assert_eq!(pages[1].entries()[0].log.tiered_through(), 27); +} + +#[tokio::test] +async fn shared_windows_enforce_total_page_and_unique_member_bounds() { + let directory = directory(); + let other = NodeId::from_bytes([8; 16]); + for requests in [ + vec![], + vec![(member(), None), (member(), None)], + vec![(NodeId::from_bytes([0; 16]), None)], + ] { + assert!( + directory + .follower_logs_pages(&requests, 1, NOW) + .await + .is_err() + ); + } + for limit in [0, 65, usize::MAX] { + assert!( + directory + .follower_logs_pages(&[(member(), None), (other, None)], limit, NOW) + .await + .is_err() + ); + } + let requests = (1..=128) + .map(|id| (NodeId::from_bytes([id; 16]), None)) + .collect::>(); + let pages = directory + .follower_logs_pages(&requests, 1, NOW) + .await + .unwrap(); + assert_eq!(pages.len(), 128); + assert!(pages.iter().all(|page| page.entries().is_empty())); + assert!( + directory + .follower_logs_pages(&requests, 1, -1) + .await + .is_err() + ); + let mut too_many = requests; + too_many.push((NodeId::from_bytes([129; 16]), None)); + assert!( + directory + .follower_logs_pages(&too_many, 1, NOW) + .await + .is_err() + ); +} + fn directory() -> NodeDirectory { NodeDirectory::new( CellStorageLayout::new( diff --git a/crates/cellule-runtime/src/node/directory/log.rs b/crates/cellule-runtime/src/node/directory/log.rs index d24bce64..14363b75 100644 --- a/crates/cellule-runtime/src/node/directory/log.rs +++ b/crates/cellule-runtime/src/node/directory/log.rs @@ -61,9 +61,9 @@ impl NodeDirectory { return Err(Error::Node("node advertisement path and session differ")); } if let NodeRecord::Advertisement(advertisement) = &record { + // load_record_at already verified this exact canonical body. self.validate_scope(advertisement)?; advertisement.validate_shape()?; - advertisement.verify_signature()?; } Ok(record .log() diff --git a/crates/cellule-runtime/src/node/directory/mod.rs b/crates/cellule-runtime/src/node/directory/mod.rs index 146f70e8..771579df 100644 --- a/crates/cellule-runtime/src/node/directory/mod.rs +++ b/crates/cellule-runtime/src/node/directory/mod.rs @@ -111,7 +111,9 @@ impl NodeDirectory { now_ms: i64, ) -> Result { self.validate(&advertisement, now_ms)?; - let encoded = advertisement.encode()?; + // validate authenticated both signature sets; serialize this exact + // immutable value without repeating their cryptographic verification. + let encoded = advertisement.canonical_bytes()?; let path = self.layout.node_path(advertisement.session.as_bytes()); let result = match self .layout @@ -141,7 +143,10 @@ impl NodeDirectory { let Some((advertisement, token)) = self.load_canonical(session).await? else { return Ok(None); }; - self.validate(&advertisement, now_ms)?; + // This fresh canonical decode already verified both immutable signature + // sets. Apply current time/scope policy to those same bytes once. + advertisement.validate_at(now_ms)?; + self.validate_scope(&advertisement)?; Ok(Some(VersionedNodeAdvertisement { advertisement, token, @@ -169,7 +174,8 @@ impl NodeDirectory { if advertisement.expires_at_ms <= now_ms { return Ok(None); } - self.validate(&advertisement, now_ms)?; + // load_canonical authenticated this exact read, including placement. + advertisement.validate_at(now_ms)?; Ok(Some(VersionedNodeAdvertisement { advertisement, token, @@ -189,7 +195,7 @@ impl NodeDirectory { return Ok(None); }; advertisement.validate_shape()?; - advertisement.verify_signature()?; + // load_canonical authenticated the exact immutable record, even expired. self.validate_scope(&advertisement)?; if advertisement.issued_at_ms > now_ms.saturating_add(MAX_CLOCK_SKEW_MS) { return Err(Error::Node("advertisement issue time is in the future")); diff --git a/crates/cellule-runtime/src/node/tests/enrollment.rs b/crates/cellule-runtime/src/node/tests/enrollment.rs index 7b80a00e..44bb9561 100644 --- a/crates/cellule-runtime/src/node/tests/enrollment.rs +++ b/crates/cellule-runtime/src/node/tests/enrollment.rs @@ -675,3 +675,82 @@ async fn enrollment_evidence_is_replay_stable_and_separates_native_outcomes() { refused.clone().evidence_digest().unwrap() ); } + +// The original recipe is independent of the immutable proof's serializer path. +fn original_evidence_digest( + prepared: &PreparedNodeLogEnrollment, + source: &VersionedNodeAdvertisement, + domain: &[u8], +) -> Digest { + let mut hash = blake3::Hasher::new(); + hash.update(domain); + hash.update(&prepared.source().encode().unwrap()); + hash.update(&source.advertisement().encode().unwrap()); + for token in [&source.token.e_tag, &source.token.version] { + if let Some(token) = token { + hash.update(&[1]); + hash.update(&(token.len() as u64).to_le_bytes()); + hash.update(token.as_bytes()); + } else { + hash.update(&[0]); + } + } + hash.update(&prepared.log().epoch().to_le_bytes()); + for follower in prepared.followers() { + hash.update(&follower.encode().unwrap()); + } + Digest::from_bytes(*hash.finalize().as_bytes()) +} + +#[tokio::test] +async fn immutable_enrollment_evidence_preserves_bytes_without_reauthenticating_history() { + for commit in [false, true] { + let (directory, source, _) = ensemble().await; + let prepared = directory + .prepare_log_enrollment(&source, 7, 1, 3, NOW_MS + 1) + .await + .unwrap() + .unwrap(); + let attempt = directory + .prepare_log_enrollment_attempt(&prepared, NOW_MS + 2) + .await + .unwrap(); + let expected = original_evidence_digest( + attempt.prepared(), + attempt.observed(), + b"cellule.node-log.attempt.v1\0", + ); + let before = signature_passes(); + assert_eq!(attempt.evidence_digest().unwrap(), expected); + assert_eq!(signature_passes() - before, 0); + if commit { + let proof = directory + .commit_log_enrollment(&attempt, NOW_MS + 3) + .await + .unwrap(); + let expected = original_evidence_digest( + proof.prepared(), + proof.enrollment(), + b"cellule.node-log.enrolled.v1\0", + ); + let before = signature_passes(); + assert_eq!(proof.evidence_digest().unwrap(), expected); + assert_eq!(proof.clone().evidence_digest().unwrap(), expected); + assert_eq!(signature_passes() - before, 0); + } else { + let proof = directory + .fence_log_enrollment(&attempt, NOW_MS + 3) + .await + .unwrap(); + let expected = original_evidence_digest( + proof.prepared(), + proof.refusal(), + b"cellule.node-log.refused.v1\0", + ); + let before = signature_passes(); + assert_eq!(proof.evidence_digest().unwrap(), expected); + assert_eq!(proof.clone().evidence_digest().unwrap(), expected); + assert_eq!(signature_passes() - before, 0); + } + } +} diff --git a/crates/cellule-runtime/src/node/tests/mod.rs b/crates/cellule-runtime/src/node/tests/mod.rs index 545fc09f..42ee050e 100644 --- a/crates/cellule-runtime/src/node/tests/mod.rs +++ b/crates/cellule-runtime/src/node/tests/mod.rs @@ -28,6 +28,7 @@ mod operational; mod placement; mod records; mod recovered_retirement; +mod scanned_records; mod sessions; // Thread-local operation counts let canonical codec tests assert a verification diff --git a/crates/cellule-runtime/src/node/tests/records.rs b/crates/cellule-runtime/src/node/tests/records.rs index ad06e473..32951b63 100644 --- a/crates/cellule-runtime/src/node/tests/records.rs +++ b/crates/cellule-runtime/src/node/tests/records.rs @@ -224,7 +224,7 @@ fn canonical_advertisement_decode_verifies_each_immutable_signature_set_once() { } } -fn canonical_advertisements() -> [NodeAdvertisement; 3] { +pub(super) fn canonical_advertisements() -> [NodeAdvertisement; 3] { let key = SigningKey::from_bytes(&[7; 32]); let legacy = advertisement(&key, 1, NOW_MS); let placement = NodePlacementCapacity { diff --git a/crates/cellule-runtime/src/node/tests/scanned_records.rs b/crates/cellule-runtime/src/node/tests/scanned_records.rs new file mode 100644 index 00000000..43885bdf --- /dev/null +++ b/crates/cellule-runtime/src/node/tests/scanned_records.rs @@ -0,0 +1,439 @@ +//! Fresh directory reads authenticate immutable bytes once, including scans. +use super::*; + +#[tokio::test] +async fn shared_follower_windows_authenticate_each_fresh_record_once() { + let requests = (9..=12) + .map(|n| (NodeId::from_bytes([n; 16]), None)) + .collect::>(); + for original in records::canonical_advertisements() { + let directory = directory(); + directory.create(original.clone(), NOW_MS).await.unwrap(); + let before = signature_passes(); + let pages = directory + .follower_logs_pages(&requests, 32, NOW_MS + 1) + .await + .unwrap(); + assert_eq!(signature_passes() - before, 1); + assert_eq!(pages.len(), requests.len()); + for (page, (member, _)) in pages.iter().zip(&requests) { + assert_eq!(page.member(), *member); + assert_eq!(page.total_logs(), 0); + assert!(page.entries().is_empty()); + assert!(page.next().is_none()); + } + let before = signature_passes(); + directory + .follower_logs_pages(&requests, 32, NOW_MS + 2) + .await + .unwrap(); + assert_eq!(signature_passes() - before, 1); + } +} + +#[tokio::test] +async fn canonical_producers_verify_once_before_creating_or_refreshing_bytes() { + let key = SigningKey::from_bytes(&[7; 32]); + for original in records::canonical_advertisements() { + let directory = directory(); + let before = signature_passes(); + let current = directory.create(original.clone(), NOW_MS).await.unwrap(); + assert_eq!(signature_passes() - before, 1); + let next = advertisement(&key, 2, NOW_MS + 1_000); + let next = match original.placement_version { + 0 => next, + PLACEMENT_SCHEMA_VERSION => next + .with_placement_capacity(original.placement.unwrap(), &key) + .unwrap(), + OPERATIONAL_PLACEMENT_SCHEMA_VERSION => next + .with_operational_placement( + original.placement.unwrap(), + NodeOperationalSample { + sequence: 2, + observed_at_ms: NOW_MS + 1_000, + mode: NodeMode::Draining, + pressure: NodePressure::Critical, + }, + &key, + ) + .unwrap(), + _ => unreachable!(), + }; + let before = signature_passes(); + let refreshed = directory + .refresh(¤t, next.clone(), NOW_MS + 1_000) + .await + .unwrap(); + assert_eq!(signature_passes() - before, 1); + assert_eq!( + refreshed.advertisement().placement_version, + original.placement_version + ); + let bytes = directory + .layout + .store() + .get_with_etag(&directory.layout.node_path(original.session().as_bytes())) + .await + .unwrap() + .0; + let decoded = NodeAdvertisement::decode_canonical(&bytes).unwrap(); + assert_eq!(&decoded, refreshed.advertisement()); + assert_eq!(decoded.encode().unwrap(), bytes); + } +} + +#[tokio::test] +async fn exact_canonical_reads_verify_once_per_fresh_record() { + for original in records::canonical_advertisements() { + let directory = directory(); + directory.create(original.clone(), NOW_MS).await.unwrap(); + let before = signature_passes(); + let loaded = directory + .load(original.session(), NOW_MS + 1) + .await + .unwrap() + .unwrap(); + assert_eq!(loaded.advertisement(), &original); + assert_eq!(signature_passes() - before, 1); + let before = signature_passes(); + let loaded = directory + .load_if_live(original.session(), NOW_MS + 1) + .await + .unwrap() + .unwrap(); + assert_eq!(loaded.advertisement(), &original); + assert_eq!(signature_passes() - before, 1); + let before = signature_passes(); + assert_eq!( + directory + .inspect_advertisement(original.session(), NOW_MS + 1) + .await + .unwrap(), + Some(original) + ); + assert_eq!(signature_passes() - before, 1); + } +} + +#[tokio::test] +async fn directory_scans_verify_once_and_preserve_expired_obligations() { + for original in records::canonical_advertisements() { + let directory = directory(); + directory.create(original.clone(), NOW_MS).await.unwrap(); + let before = signature_passes(); + assert_eq!( + directory.live(NOW_MS + 1, 128).await.unwrap(), + vec![original.clone()] + ); + assert_eq!(signature_passes() - before, 1); + let before = signature_passes(); + assert_eq!( + directory + .advertised_sessions(NOW_MS + 20_000, 128) + .await + .unwrap(), + vec![original.session()] + ); + assert_eq!(signature_passes() - before, 1); + let current = directory + .load(original.session(), NOW_MS + 1) + .await + .unwrap() + .unwrap(); + let mut with_log = original.clone(); + with_log.log = Some( + NodeLogStatus::open(original.node(), 7, vec![NodeId::from_bytes([9; 16])]).unwrap(), + ); + directory + .update_advertisement(¤t, with_log, NOW_MS + 1) + .await + .unwrap(); + let before = signature_passes(); + let page = directory + .follower_logs_page(NodeId::from_bytes([9; 16]), None, 128, NOW_MS + 20_000) + .await + .unwrap(); + assert_eq!(page.entries().len(), 1); + assert_eq!(page.entries()[0].leader, original.session()); + assert_eq!(page.entries()[0].leader_state, LogLeaderState::Expired); + assert_eq!(page.entries()[0].log.epoch(), 7); + assert_eq!(signature_passes() - before, 1); + let before = signature_passes(); + assert!( + directory + .log_epoch_referenced(original.session(), 7) + .await + .unwrap() + ); + assert_eq!(signature_passes() - before, 1); + } +} + +#[tokio::test] +async fn every_fresh_scan_rejects_forged_identity_and_operational_signatures() { + for original in records::canonical_advertisements() { + let bytes = original.encode().unwrap(); + for placement in [false, true] { + if placement && !original.has_signed_placement() { + continue; + } + let directory = directory(); + let mut raw: serde_json::Value = serde_json::from_slice(&bytes).unwrap(); + if placement { + raw["placement_signature"] = serde_json::Value::String("00".repeat(64)); + } else { + raw["identity"]["signature"] = serde_json::Value::String("00".repeat(64)); + } + directory + .layout + .store() + .create_strict( + &directory.layout.node_path(original.session().as_bytes()), + Bytes::from(serde_json::to_vec(&raw).unwrap()), + ) + .await + .unwrap(); + assert!( + directory + .load(original.session(), NOW_MS + 1) + .await + .is_err() + ); + assert!( + directory + .log_epoch_referenced(original.session(), 7) + .await + .is_err() + ); + assert!( + directory + .load_if_live(original.session(), NOW_MS + 1) + .await + .is_err() + ); + assert!( + directory + .inspect_advertisement(original.session(), NOW_MS + 20_000) + .await + .is_err() + ); + assert!(directory.live(NOW_MS + 1, 128).await.is_err()); + assert!( + directory + .advertised_sessions(NOW_MS + 20_000, 128) + .await + .is_err() + ); + assert!( + directory + .follower_logs_page(NodeId::from_bytes([9; 16]), None, 128, NOW_MS + 20_000) + .await + .is_err() + ); + assert!( + directory + .follower_logs_pages( + &[ + (NodeId::from_bytes([9; 16]), None), + (NodeId::from_bytes([10; 16]), None) + ], + 64, + NOW_MS + 20_000 + ) + .await + .is_err() + ); + } + } +} + +#[tokio::test] +async fn later_canonical_read_observes_new_bytes_and_authenticates_them_again() { + let [original, bridge, operational] = records::canonical_advertisements(); + for next in [bridge, operational] { + let directory = directory(); + let current = directory.create(original.clone(), NOW_MS).await.unwrap(); + let before = signature_passes(); + assert_eq!( + directory + .load(original.session(), NOW_MS + 1) + .await + .unwrap() + .unwrap() + .advertisement(), + &original + ); + assert_eq!(signature_passes() - before, 1); + // Replace the same path with different valid signed bytes. No process, + // page or ETag cache may reuse an earlier signature decision. + let body = Bytes::from(next.encode().unwrap()); + directory + .layout + .store() + .update( + &directory.layout.node_path(original.session().as_bytes()), + body, + current.token.clone(), + ) + .await + .unwrap(); + let before = signature_passes(); + let loaded = directory + .load(original.session(), NOW_MS + 1) + .await + .unwrap() + .unwrap(); + assert_eq!(loaded.advertisement(), &next); + assert_eq!(signature_passes() - before, 1); + } +} + +#[tokio::test] +async fn authenticated_reads_keep_path_scope_and_time_policies() { + for original in records::canonical_advertisements() { + let directory = directory(); + directory.create(original.clone(), NOW_MS).await.unwrap(); + assert!( + directory + .load(original.session(), NOW_MS + 20_000) + .await + .is_err() + ); + assert!( + directory + .load_if_live(original.session(), NOW_MS + 20_000) + .await + .unwrap() + .is_none() + ); + assert!( + directory + .live(NOW_MS + 20_000, 128) + .await + .unwrap() + .is_empty() + ); + let future = NOW_MS - MAX_CLOCK_SKEW_MS - 1; + assert!(directory.load(original.session(), future).await.is_err()); + assert!( + directory + .load_if_live(original.session(), future) + .await + .is_err() + ); + assert!( + directory + .inspect_advertisement(original.session(), future) + .await + .is_err() + ); + assert!(directory.advertised_sessions(future, 128).await.is_err()); + assert!( + directory + .follower_logs_page(NodeId::from_bytes([9; 16]), None, 128, future) + .await + .is_err() + ); + let foreign = NodeDirectory::new( + directory.layout.clone(), + Digest::from_bytes([88; 32]), + directory.image, + directory.release, + ); + assert!(foreign.load(original.session(), NOW_MS + 1).await.is_err()); + assert!( + foreign + .load_if_live(original.session(), NOW_MS + 1) + .await + .is_err() + ); + assert!( + foreign + .inspect_advertisement(original.session(), NOW_MS + 1) + .await + .is_err() + ); + assert!(foreign.live(NOW_MS + 1, 128).await.is_err()); + assert!( + foreign + .advertised_sessions(NOW_MS + 20_000, 128) + .await + .is_err() + ); + assert!( + foreign + .follower_logs_page(NodeId::from_bytes([9; 16]), None, 128, NOW_MS + 20_000) + .await + .is_err() + ); + let body = Bytes::from(original.encode().unwrap()); + directory + .layout + .store() + .create_strict(&directory.layout.node_path(&[77; 16]), body) + .await + .unwrap(); + assert!( + directory + .load(SessionId::from_bytes([77; 16]), NOW_MS + 1) + .await + .is_err() + ); + assert!(directory.live(NOW_MS + 1, 128).await.is_err()); + assert!( + directory + .advertised_sessions(NOW_MS + 20_000, 128) + .await + .is_err() + ); + assert!( + directory + .follower_logs_page(NodeId::from_bytes([9; 16]), None, 128, NOW_MS + 20_000) + .await + .is_err() + ); + } +} + +#[tokio::test] +async fn invalid_producer_signatures_never_reach_create_or_refresh_cas() { + for original in records::canonical_advertisements() { + for placement in [false, true] { + if placement && !original.has_signed_placement() { + continue; + } + let directory = directory(); + let mut invalid = original.clone(); + if placement { + invalid.placement_signature[0] ^= 1; + } else { + invalid.signature[0] ^= 1; + } + assert!(matches!( + directory.create(invalid.clone(), NOW_MS).await, + Err(Error::PeerSignature(_)) + )); + assert!( + directory + .load(original.session(), NOW_MS) + .await + .unwrap() + .is_none() + ); + let current = directory.create(original.clone(), NOW_MS).await.unwrap(); + invalid.issued_at_ms += 1_000; + invalid.expires_at_ms += 1_000; + assert!(matches!( + directory.refresh(¤t, invalid, NOW_MS + 1_000).await, + Err(Error::PeerSignature(_)) + )); + let loaded = directory + .load(original.session(), NOW_MS + 1_000) + .await + .unwrap() + .unwrap(); + assert_eq!(loaded.advertisement(), &original); + assert_eq!(loaded.token, current.token); + } + } +} diff --git a/crates/cellule-runtime/src/primitives/blob/api/mod.rs b/crates/cellule-runtime/src/primitives/blob/api/mod.rs index 571b5988..d926457e 100644 --- a/crates/cellule-runtime/src/primitives/blob/api/mod.rs +++ b/crates/cellule-runtime/src/primitives/blob/api/mod.rs @@ -123,15 +123,25 @@ impl BlobNamespace { } /// Applies one multipart or conditional mutation on the key's shard. + /// Once accepted, the original operation owns staging through command + /// response, including after caller cancellation or artifact-store closure. + /// Prepare separately and retain evidence when cancellation may require + /// resolving an uncertain command outcome. pub async fn mutate( &self, identity: crate::cell::executor::MutationIdentity, mutation: BlobMutation, ) -> std::result::Result, InvocationError> { - self.prepare_mutation(identity, mutation) - .await? - .execute() + let namespace = self.clone(); + self.artifact_store + .run_invocation(async move { + namespace + .prepare_mutation_native(identity, mutation) + .await? + .execute_native() + .await + }) .await } @@ -141,6 +151,10 @@ impl BlobNamespace { /// manifest reference or make an object visible. Retain the returned /// command's evidence before executing; resolve it after cancellation or /// an uncertain reply before deciding whether to retry the same command. + /// The returned command is caller-owned and no longer a store job. Its + /// execution acquires the same original store's admission, so closure refuses + /// new dispatch while already accepted dispatch remains owned through its + /// response. This does not settle unknown results from earlier execution. pub async fn prepare_mutation( &self, identity: crate::cell::executor::MutationIdentity, @@ -148,6 +162,22 @@ impl BlobNamespace { ) -> std::result::Result< crate::client::PreparedCommand>, InvocationError, + > { + let namespace = self.clone(); + self.artifact_store + .run_invocation( + async move { namespace.prepare_mutation_native(identity, mutation).await }, + ) + .await + } + + async fn prepare_mutation_native( + &self, + identity: crate::cell::executor::MutationIdentity, + mutation: BlobMutation, + ) -> std::result::Result< + crate::client::PreparedCommand>, + InvocationError, > { let target = self .target(mutation_key(&mutation)) @@ -165,8 +195,9 @@ impl BlobNamespace { ))); } let digest = super::part_digest(&payload); + let size = payload.len() as u32; self.artifact_store - .put_part(digest, &payload) + .put_part_native(digest, bytes::Bytes::from(payload)) .await .map_err(InvocationError::NotStarted)?; BlobMutation::PutPartRef { @@ -174,7 +205,7 @@ impl BlobNamespace { upload_id, part_number, digest, - size: payload.len() as u32, + size, } } mutation => mutation, @@ -185,10 +216,23 @@ impl BlobNamespace { } /// Reads metadata or a bounded range from one key shard. + /// The original accepted lifetime covers metadata and all part reads; + /// closing admission or cancelling the caller does not truncate that work. pub async fn query( &self, query: BlobQuery, minimum: Option, + ) -> std::result::Result, InvocationError> { + let namespace = self.clone(); + self.artifact_store + .run_invocation(async move { namespace.query_native(query, minimum).await }) + .await + } + + async fn query_native( + &self, + query: BlobQuery, + minimum: Option, ) -> std::result::Result, InvocationError> { let key = query_key(&query).ok_or_else(|| { InvocationError::NotStarted(crate::Error::Identity( @@ -220,16 +264,21 @@ impl BlobNamespace { let target = self .shard_target(shard) .map_err(InvocationError::NotStarted)?; - self.client - .query::>( - &target, - minimum, - BlobQuery::List { - prefix, - after, - limit, - }, - ) + let client = self.client.clone(); + self.artifact_store + .run_invocation(async move { + client + .query::>( + &target, + minimum, + BlobQuery::List { + prefix, + after, + limit, + }, + ) + .await + }) .await } @@ -276,7 +325,9 @@ async fn hydrate_read(store: &BlobArtifactStore, read: &mut BlobRead) -> crate:: let end = read.end.min(read.metadata.size); let mut bytes = Vec::with_capacity((end.saturating_sub(read.offset)) as usize); for part in &read.parts { - let payload = store.read_part(part.digest, part.size).await?; + // All parts belong to the public query's original accepted lifetime. + // Closing admission between parts must not interrupt its remaining I/O. + let payload = store.read_part_native(part.digest, part.size).await?; let part_end = part.offset.saturating_add(u64::from(part.size)); let start = read.offset.saturating_sub(part.offset) as usize; let take_end = end.saturating_sub(part.offset).min(u64::from(part.size)) as usize; diff --git a/crates/cellule-runtime/src/primitives/blob/mod.rs b/crates/cellule-runtime/src/primitives/blob/mod.rs index c18a3eb7..45178845 100644 --- a/crates/cellule-runtime/src/primitives/blob/mod.rs +++ b/crates/cellule-runtime/src/primitives/blob/mod.rs @@ -19,7 +19,7 @@ mod store; mod tests; pub use api::{BlobCommand, BlobModule, BlobNamespace, BlobQueryCommand, register_blob}; -pub use store::{BlobArtifactStore, BlobGarbageCollectionReport}; +pub use store::{BlobArtifactLifecycleObservation, BlobArtifactStore, BlobGarbageCollectionReport}; use sql::*; use store::part_digest; diff --git a/crates/cellule-runtime/src/primitives/blob/store/lifecycle.rs b/crates/cellule-runtime/src/primitives/blob/store/lifecycle.rs new file mode 100644 index 00000000..73655342 --- /dev/null +++ b/crates/cellule-runtime/src/primitives/blob/store/lifecycle.rs @@ -0,0 +1,182 @@ +//! One admission word and original native owner shared by every store capability. +use super::*; +use std::{ + future::Future, + sync::{ + Mutex, + atomic::{AtomicUsize, Ordering}, + }, +}; +use tokio::sync::{Notify, oneshot}; + +const CLOSED: usize = 1 << (usize::BITS - 1); +const MAX_JOBS: usize = 64; + +/// Local closure of one original artifact-store owner. This does not prove +/// global Blob references, remote success, retention pins or Cell authority. +#[derive(Clone, Debug)] +pub struct BlobArtifactLifecycleObservation { + closed: bool, + accepted_jobs: usize, + unjoined_jobs: usize, + first_failure: Option>, +} +impl BlobArtifactLifecycleObservation { + /// Whether the original shared admission word is irreversibly closed. + #[must_use] + pub const fn admission_closed(&self) -> bool { + self.closed + } + /// Accepted operations with retained supervisors, including cancelled + /// callers' work. Inspect `unjoined_jobs` for lost original supervisors. + #[must_use] + pub const fn accepted_jobs(&self) -> usize { + self.accepted_jobs + } + /// Original supervisors lost before observing their native task join. + /// Their native work may still exist; zero running jobs cannot close it. + #[must_use] + pub const fn unjoined_jobs(&self) -> usize { + self.unjoined_jobs + } + /// Closed admission with no accepted or unjoined original work. Operation + /// success and remote outcome still require their own result evidence. + #[must_use] + pub const fn locally_joined(&self) -> bool { + self.closed && self.accepted_jobs == 0 && self.unjoined_jobs == 0 + } + /// Original first native failure, retained after waiter loss and closure. + /// This diagnostic cannot classify an ambiguous remote result as absent. + #[must_use] + pub fn first_failure(&self) -> Option<&Arc> { + self.first_failure.as_ref() + } +} + +#[derive(Default)] +pub(super) struct ArtifactLifetime { + state: AtomicUsize, + changed: Notify, + first_failure: Mutex>>, + unjoined_jobs: AtomicUsize, +} +struct Accepted { + lifetime: Arc, + joined: bool, +} +impl Drop for Accepted { + fn drop(&mut self) { + if !self.joined { + // Runtime teardown can drop this supervisor before the separately + // owned native task (or its provider job) joins. Retain unknown + // joining and close admission; releasing a task count is no proof. + self.lifetime.unjoined_jobs.fetch_add(1, Ordering::AcqRel); + self.lifetime.close(); + self.lifetime.retain_failure(Error::Control( + "Blob artifact supervisor ended before original native join", + )); + } + self.lifetime.state.fetch_sub(1, Ordering::AcqRel); + self.lifetime.changed.notify_waiters(); + } +} +impl ArtifactLifetime { + fn accept(self: &Arc) -> Result { + self.state + .try_update(Ordering::AcqRel, Ordering::Acquire, |state| { + (state & CLOSED == 0 && state < MAX_JOBS).then_some(state + 1) + }) + .map_err(|state| { + if state & CLOSED != 0 { + Error::CellDraining + } else { + Error::Capacity("Blob artifact jobs") + } + })?; + Ok(Accepted { + lifetime: self.clone(), + joined: false, + }) + } + pub(super) fn close(&self) { + self.state.fetch_or(CLOSED, Ordering::AcqRel); + } + pub(super) async fn join(&self) -> Result<()> { + loop { + let changed = self.changed.notified(); + tokio::pin!(changed); + changed.as_mut().enable(); + if self.state.load(Ordering::Acquire) & !CLOSED == 0 { + // A final guard may have recorded an unproven join before its + // release; acquire its state before checking that flag again. + if self.unjoined_jobs.load(Ordering::Acquire) != 0 { + return Err(Error::Control( + "Blob artifact original native join is unproven", + )); + } + return Ok(()); + } + changed.await; + } + } + pub(super) fn observe(&self) -> Result { + // Closed-plus-zero acquires every final result write before reading + // diagnostics; reading the failure first could miss the last job's error. + let state = self.state.load(Ordering::Acquire); + let first_failure = self + .first_failure + .lock() + .map_err(|_| Error::Control("Blob artifact failure lock poisoned"))? + .clone(); + Ok(BlobArtifactLifecycleObservation { + closed: state & CLOSED != 0, + accepted_jobs: state & !CLOSED, + unjoined_jobs: self.unjoined_jobs.load(Ordering::Acquire), + first_failure, + }) + } + pub(super) fn retain_failure(&self, error: Error) -> Error { + let original = match error { + Error::Shared(original) => original, + error => Arc::new(error), + }; + match self.first_failure.lock() { + Ok(mut first) => { + first.get_or_insert_with(|| original.clone()); + } + Err(poisoned) => { + poisoned + .into_inner() + .get_or_insert_with(|| original.clone()); + } + } + Error::Shared(original) + } + pub(super) async fn run(self: &Arc, work: F) -> Result + where + F: Future> + Send + 'static, + T: Send + 'static, + { + let runtime = tokio::runtime::Handle::try_current().map_err(Error::RuntimeStart)?; + let mut accepted = self.accept()?; + let (reply, waiter) = oneshot::channel(); + // The retained supervisor joins the original native task even if its + // caller cancels. This also preserves a provider panic's JoinError. + runtime.spawn(async move { + let result = match tokio::spawn(work).await { + Ok(result) => result, + Err(source) => Err(Error::Facility { + name: "blob-artifacts", + source: Box::new(source), + }), + }; + let result = result.map_err(|error| accepted.lifetime.retain_failure(error)); + let _ = reply.send(result); + // Inputs, native I/O and undelivered output have all completed or + // dropped before waking an irreversible closed-plus-zero waiter. + accepted.joined = true; + drop(accepted); + }); + waiter.await.map_err(|_| Error::RuntimeClosed)? + } +} diff --git a/crates/cellule-runtime/src/primitives/blob/store.rs b/crates/cellule-runtime/src/primitives/blob/store/mod.rs similarity index 52% rename from crates/cellule-runtime/src/primitives/blob/store.rs rename to crates/cellule-runtime/src/primitives/blob/store/mod.rs index 853a9b79..c87d0a5d 100644 --- a/crates/cellule-runtime/src/primitives/blob/store.rs +++ b/crates/cellule-runtime/src/primitives/blob/store/mod.rs @@ -1,6 +1,14 @@ //! Durable Blob part store backed by the configured object store. use super::*; +use crate::client::InvocationError; +use std::sync::Arc; + +mod lifecycle; +#[cfg(test)] +mod tests; +use lifecycle::ArtifactLifetime; +pub use lifecycle::BlobArtifactLifecycleObservation; /// Object-store backing for Blob parts. /// @@ -10,6 +18,7 @@ use super::*; #[derive(Clone)] pub struct BlobArtifactStore { store: Store, + lifetime: Arc, } /// Result from one Blob part reachability sweep. @@ -41,23 +50,117 @@ impl BlobGarbageCollectionReport { } impl BlobArtifactStore { - /// Wraps a configured object store for Blob artifact data. + /// Wraps a configured object store with shared, irreversible admission. + /// At most 64 original operations are accepted concurrently. Namespace + /// operations span metadata, provider I/O and their command/query response; + /// GC retains the complete reference set. Clones share closure and keep + /// accepted operations running after caller cancellation. #[must_use] pub fn new(store: Store) -> Self { - Self { store } + Self { + store, + lifetime: Arc::new(ArtifactLifetime::default()), + } + } + + /// Irreversibly refuses new artifact I/O through every clone. Accepted jobs + /// continue; use `close_and_join` before disposing of their provider. + pub fn close(&self) { + self.lifetime.close(); + } + + /// Closes admission and waits for known original operations, returning an + /// error if any original supervisor was lost before observing native join. + /// Cancelling this waiter never cancels accepted work or reopens admission. + /// This proves local joining, not cross-Cell reachability, remote operation + /// success, Cell release, pin retirement or permission to finalize a node. + /// Returned prepared commands acquire this same owner's admission when + /// executed by its configured client. Closure refuses new dispatch; it does + /// not prove absence of uncertain earlier commands or writes through other + /// client/provider capabilities. + pub async fn close_and_join(&self) -> Result { + self.close(); + self.lifetime.join().await?; + self.lifecycle_observation() + } + + /// Reads original local admission, accepted-work and failure diagnostics. + /// Open counts are advisory; closed plus zero is irreversible. + pub fn lifecycle_observation(&self) -> Result { + self.lifetime.observe() } + pub(crate) async fn run_invocation( + &self, + work: F, + ) -> std::result::Result> + where + T: Send + 'static, + U: Send + 'static, + F: std::future::Future>> + + Send + + 'static, + { + let lifetime = self.lifetime.clone(); + self.lifetime + .run(async move { + // Preserve typed rejection/pending evidence. Only source-bearing + // failures enter local diagnostics; joining a Pending result does + // not prove the remote command completed or failed to execute. + Ok(work.await.map_err(|error| match error { + InvocationError::NotStarted(source) => { + InvocationError::NotStarted(lifetime.retain_failure(source)) + } + InvocationError::InvalidPublishedResult { receipt, source } => { + InvocationError::InvalidPublishedResult { + receipt, + source: Box::new(lifetime.retain_failure(*source)), + } + } + other => other, + })) + }) + .await + .map_err(InvocationError::NotStarted)? + } + + #[cfg(test)] pub(super) async fn put_part(&self, digest: [u8; 32], payload: &[u8]) -> Result<()> { + if payload.len() > MAX_BLOB_PART_BYTES { + return Err(Error::Command("blob part exceeds 256 KiB")); + } if part_digest(payload) != digest { return Err(Error::Command("blob part digest does not match payload")); } - self.store - .put(&self.part_path(&digest), Bytes::copy_from_slice(payload)) - .await?; + // Copy only bounded input; the accepted native owner retains it even + // when preparation's caller disappears before immutable publication. + let payload = Bytes::copy_from_slice(payload); + let store = self.clone(); + self.lifetime + .run(async move { store.put_part_native(digest, payload).await }) + .await + } + + pub(super) async fn put_part_native(&self, digest: [u8; 32], payload: Bytes) -> Result<()> { + self.store.put(&self.part_path(&digest), payload).await?; Ok(()) } + #[cfg(test)] pub(super) async fn read_part(&self, digest: [u8; 32], size: u32) -> Result> { + if usize::try_from(size) + .ok() + .is_none_or(|size| size > MAX_BLOB_PART_BYTES) + { + return Err(Error::Command("invalid stored blob part size")); + } + let store = self.clone(); + self.lifetime + .run(async move { store.read_part_native(digest, size).await }) + .await + } + + pub(super) async fn read_part_native(&self, digest: [u8; 32], size: u32) -> Result> { if usize::try_from(size) .ok() .is_none_or(|size| size > MAX_BLOB_PART_BYTES) @@ -82,15 +185,28 @@ impl BlobArtifactStore { /// Objects newer than `cutoff_ms` are retained for uploads whose SQLite /// manifests have not committed. Each call inspects the complete unordered /// listing until 128 parts have been deleted; schedule another call while - /// [`BlobGarbageCollectionReport::has_more`] is true. + /// [`BlobGarbageCollectionReport::has_more`] is true. The complete reference + /// set is shared ownership because native deletion outlives a cancelled + /// waiter; never truncate it to satisfy an admission or deletion budget. pub async fn sweep_unreferenced( &self, - live_digests: &BTreeSet<[u8; 32]>, + live_digests: Arc>, cutoff_ms: i64, ) -> Result { if cutoff_ms < 0 { return Err(Error::Command("negative blob garbage-collection cutoff")); } + let store = self.clone(); + self.lifetime + .run(async move { store.sweep_native(&live_digests, cutoff_ms).await }) + .await + } + + async fn sweep_native( + &self, + live_digests: &BTreeSet<[u8; 32]>, + cutoff_ms: i64, + ) -> Result { let prefix = self .store .storage_scope() diff --git a/crates/cellule-runtime/src/primitives/blob/store/tests.rs b/crates/cellule-runtime/src/primitives/blob/store/tests.rs new file mode 100644 index 00000000..ca5e3174 --- /dev/null +++ b/crates/cellule-runtime/src/primitives/blob/store/tests.rs @@ -0,0 +1,506 @@ +//! Real provider entrypoints under caller loss, closure, failure and capacity. +use super::*; +use futures_util::stream::BoxStream; +use object_store::{ + CopyOptions, GetOptions, GetResult, ListResult, MultipartUpload, ObjectMeta, ObjectStore, + ObjectStoreExt, PutMultipartOptions, PutOptions, PutPayload, PutResult, +}; +use std::{ + fmt, + sync::atomic::{AtomicBool, AtomicU8, AtomicUsize, Ordering}, +}; +use tokio::sync::Notify; + +#[derive(Default, Debug)] +struct Gate { + kind: AtomicU8, + entered: AtomicUsize, + released: AtomicBool, + changed: Notify, + resume: Notify, + fail: AtomicBool, + panic: AtomicBool, +} +impl Gate { + async fn wait(&self, count: usize) { + tokio::time::timeout(std::time::Duration::from_secs(5), async { + loop { + let changed = self.changed.notified(); + tokio::pin!(changed); + changed.as_mut().enable(); + if self.entered.load(Ordering::Acquire) >= count { + break; + } + changed.await; + } + }) + .await + .unwrap(); + } + fn release(&self) { + self.released.store(true, Ordering::Release); + self.resume.notify_waiters(); + } + async fn before(&self, kind: u8) { + if self.kind.load(Ordering::Acquire) != kind { + return; + } + self.entered.fetch_add(1, Ordering::AcqRel); + self.changed.notify_waiters(); + loop { + let resumed = self.resume.notified(); + tokio::pin!(resumed); + resumed.as_mut().enable(); + if self.released.load(Ordering::Acquire) { + return; + } + resumed.await; + } + } +} +#[derive(Debug)] +struct ProviderFailure(Arc); +impl fmt::Display for ProviderFailure { + fn fmt(&self, out: &mut fmt::Formatter<'_>) -> fmt::Result { + out.write_str("original Blob provider failure") + } +} +impl std::error::Error for ProviderFailure {} +#[derive(Debug)] +struct Provider { + inner: Arc, + gate: Arc, + failure: Arc, + blocking: Option>, +} +#[derive(Default, Debug)] +struct BlockingJob { + entered: AtomicBool, + finished: AtomicBool, + released: std::sync::Mutex, + resume: std::sync::Condvar, +} +impl BlockingJob { + fn release(&self) { + *self.released.lock().unwrap() = true; + self.resume.notify_all(); + } +} +impl fmt::Display for Provider { + fn fmt(&self, out: &mut fmt::Formatter<'_>) -> fmt::Result { + out.write_str("blob-native-test-provider") + } +} +#[async_trait::async_trait] +impl ObjectStore for Provider { + async fn put_opts( + &self, + location: &ObjectPath, + payload: PutPayload, + options: PutOptions, + ) -> object_store::Result { + self.gate.before(1).await; + if let Some(blocking) = &self.blocking { + let blocking = blocking.clone(); + let inner = self.inner.clone(); + let path = location.clone(); + return tokio::task::spawn_blocking(move || { + use futures_util::FutureExt; + blocking.entered.store(true, Ordering::Release); + let mut released = blocking.released.lock().unwrap(); + while !*released { + released = blocking.resume.wait(released).unwrap(); + } + drop(released); + // This uncontended InMemory put is immediately ready. Actual + // publication occurs on the original blocking worker, which + // survives teardown of its waiting Tokio runtime. + let result = inner + .put_opts(&path, payload, options) + .now_or_never() + .unwrap(); + blocking.finished.store(true, Ordering::Release); + result + }) + .await + .unwrap(); + } + assert!( + !self.gate.panic.load(Ordering::Acquire), + "original Blob provider panic" + ); + if self.gate.fail.load(Ordering::Acquire) { + // Store preserves NotSupported's source and does not retry it. + // Auth classification intentionally maps PermissionDenied to a + // domain variant without a source before Blob receives the error. + return Err(object_store::Error::NotSupported { + source: Box::new(ProviderFailure(self.failure.clone())), + }); + } + self.inner.put_opts(location, payload, options).await + } + async fn put_multipart_opts( + &self, + location: &ObjectPath, + options: PutMultipartOptions, + ) -> object_store::Result> { + self.inner.put_multipart_opts(location, options).await + } + async fn get_opts( + &self, + location: &ObjectPath, + options: GetOptions, + ) -> object_store::Result { + self.gate.before(2).await; + self.inner.get_opts(location, options).await + } + fn delete_stream( + &self, + locations: BoxStream<'static, object_store::Result>, + ) -> BoxStream<'static, object_store::Result> { + let gate = self.gate.clone(); + let inner = self.inner.clone(); + Box::pin(locations.then(move |location| { + let gate = gate.clone(); + let inner = inner.clone(); + async move { + let location = location?; + gate.before(3).await; + inner.delete(&location).await?; + Ok(location) + } + })) + } + fn list( + &self, + prefix: Option<&ObjectPath>, + ) -> BoxStream<'static, object_store::Result> { + self.inner.list(prefix) + } + async fn list_with_delimiter( + &self, + prefix: Option<&ObjectPath>, + ) -> object_store::Result { + self.inner.list_with_delimiter(prefix).await + } + async fn copy_opts( + &self, + from: &ObjectPath, + to: &ObjectPath, + options: CopyOptions, + ) -> object_store::Result<()> { + self.inner.copy_opts(from, to, options).await + } +} +fn fixture() -> (BlobArtifactStore, Arc) { + let provider = Arc::new(Provider { + inner: Arc::new(object_store::memory::InMemory::new()), + gate: Arc::new(Gate::default()), + failure: Arc::new(37), + blocking: None, + }); + ( + BlobArtifactStore::new(Store::new(provider.clone())), + provider, + ) +} +async fn join(store: &BlobArtifactStore) -> BlobArtifactLifecycleObservation { + tokio::time::timeout(std::time::Duration::from_secs(5), store.close_and_join()) + .await + .unwrap() + .unwrap() +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn cancelled_original_put_read_and_gc_remain_owned_through_repeated_close() { + for kind in [1, 2, 3] { + let (store, provider) = fixture(); + let payload = b"retained immutable Blob bytes"; + let digest = part_digest(payload); + if kind != 1 { + provider + .inner + .put( + &store.part_path(&digest), + Bytes::from_static(payload).into(), + ) + .await + .unwrap(); + } + let refs = Arc::new(BTreeSet::new()); + let retained_refs = Arc::downgrade(&refs); + provider.gate.kind.store(kind, Ordering::Release); + let caller = { + let store = store.clone(); + tokio::spawn(async move { + match kind { + 1 => { + store.put_part(digest, payload).await?; + Ok(Vec::new()) + } + 2 => store.read_part(digest, payload.len() as u32).await, + _ => { + let result = store.sweep_unreferenced(refs, i64::MAX).await?; + Ok(result.deleted().to_be_bytes().to_vec()) + } + } + }) + }; + provider.gate.wait(1).await; + assert_eq!(store.lifecycle_observation().unwrap().accepted_jobs(), 1); + caller.abort(); + assert!(caller.await.unwrap_err().is_cancelled()); + let closing = { + let store = store.clone(); + tokio::spawn(async move { store.close_and_join().await }) + }; + tokio::task::yield_now().await; + store.close(); + let closed = store.lifecycle_observation().unwrap(); + assert!(closed.admission_closed() && !closed.locally_joined()); + assert_eq!(closed.accepted_jobs(), 1); + assert!(matches!( + store.put_part(digest, payload).await, + Err(Error::CellDraining) + )); + assert!(matches!( + store.read_part(digest, payload.len() as u32).await, + Err(Error::CellDraining) + )); + assert!(!closing.is_finished()); + closing.abort(); + assert!(closing.await.unwrap_err().is_cancelled()); + if kind == 3 { + assert!(retained_refs.upgrade().is_some()); + } + provider.gate.release(); + let closed = join(&store).await; + assert!(closed.locally_joined()); + assert!(closed.first_failure().is_none()); + assert!(join(&store.clone()).await.locally_joined()); + assert!(retained_refs.upgrade().is_none()); + let native = provider.inner.get(&store.part_path(&digest)).await; + if kind == 3 { + assert!(matches!(native, Err(object_store::Error::NotFound { .. }))); + } else { + assert_eq!(native.unwrap().bytes().await.unwrap(), payload.as_slice()); + } + } +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn failure_and_panic_keep_original_native_sources_after_caller_loss() { + for (panic, cancel) in [(false, false), (false, true), (true, true)] { + let (store, provider) = fixture(); + provider.gate.kind.store(1, Ordering::Release); + provider.gate.fail.store(!panic, Ordering::Release); + provider.gate.panic.store(panic, Ordering::Release); + let caller = { + let store = store.clone(); + tokio::spawn(async move { store.put_part(part_digest(b"fail"), b"fail").await }) + }; + provider.gate.wait(1).await; + let caller = if cancel { + caller.abort(); + assert!(caller.await.unwrap_err().is_cancelled()); + None + } else { + Some(caller) + }; + provider.gate.release(); + let observed = join(&store).await; + assert!(observed.locally_joined()); + let original = observed.first_failure().unwrap(); + if let Some(caller) = caller { + let Error::Shared(shared) = caller.await.unwrap().unwrap_err() else { + panic!("missing retained original error") + }; + assert!(Arc::ptr_eq(&shared, original)); + } + let mut cause: &(dyn std::error::Error + 'static) = original.as_ref(); + loop { + if let Some(failure) = cause.downcast_ref::() { + assert!(!panic); + assert!(Arc::ptr_eq(&failure.0, &provider.failure)); + break; + } + if let Some(failure) = cause.downcast_ref::() { + assert!(panic && failure.is_panic()); + break; + } + cause = cause.source().expect("native source was replaced"); + } + assert!(Arc::ptr_eq( + join(&store).await.first_failure().unwrap(), + original + )); + } +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn native_capacity_is_retained_after_waiter_cancellation_until_actual_join() { + let (store, provider) = fixture(); + provider.gate.kind.store(1, Ordering::Release); + let mut callers = Vec::new(); + for index in 0..64_u8 { + let store = store.clone(); + callers.push(tokio::spawn(async move { + let bytes = [index; 4]; + store.put_part(part_digest(&bytes), &bytes).await + })); + } + provider.gate.wait(64).await; + assert_eq!(store.lifecycle_observation().unwrap().accepted_jobs(), 64); + assert!(matches!( + store.put_part(part_digest(b"spill"), b"spill").await, + Err(Error::Capacity("Blob artifact jobs")) + )); + for caller in &callers { + caller.abort(); + } + for caller in callers { + assert!(caller.await.unwrap_err().is_cancelled()); + } + assert_eq!(store.lifecycle_observation().unwrap().accepted_jobs(), 64); + store.close(); + provider.gate.release(); + assert!(join(&store).await.locally_joined()); + assert_eq!( + provider.inner.list(None).collect::>().await.len(), + 64 + ); + assert!(matches!( + store.put_part(part_digest(b"spill"), b"spill").await, + Err(Error::CellDraining) + )); +} + +#[test] +fn forced_runtime_loss_cannot_turn_an_unjoined_provider_job_into_local_closure() { + let blocking = Arc::new(BlockingJob::default()); + let provider = Arc::new(Provider { + inner: Arc::new(object_store::memory::InMemory::new()), + gate: Arc::new(Gate::default()), + failure: Arc::new(38), + blocking: Some(blocking.clone()), + }); + let store = BlobArtifactStore::new(Store::new(provider.clone())); + let former = tokio::runtime::Builder::new_multi_thread() + .worker_threads(2) + .enable_all() + .build() + .unwrap(); + let caller = { + let store = store.clone(); + former.spawn(async move { + store + .put_part(part_digest(b"unfinished"), b"unfinished") + .await + }) + }; + former.block_on(async { + tokio::time::timeout(std::time::Duration::from_secs(5), async { + while !blocking.entered.load(Ordering::Acquire) { + tokio::task::yield_now().await; + } + }) + .await + .unwrap(); + }); + former.shutdown_background(); + store.close(); + let current = tokio::runtime::Builder::new_multi_thread() + .worker_threads(2) + .enable_all() + .build() + .unwrap(); + let supervisors_stopped = current.block_on(async { + tokio::time::timeout(std::time::Duration::from_secs(5), async { + while store.lifecycle_observation().unwrap().accepted_jobs() != 0 { + tokio::task::yield_now().await; + } + }) + .await + }); + let observed = store.lifecycle_observation().unwrap(); + let early_close = current.block_on(async { + tokio::time::timeout(std::time::Duration::from_secs(5), store.close_and_join()).await + }); + let finished_early = blocking.finished.load(Ordering::Acquire); + // Always let the original native worker finish before any failure assertion. + blocking.release(); + let native_finished = current.block_on(async { + tokio::time::timeout(std::time::Duration::from_secs(5), async { + while !blocking.finished.load(Ordering::Acquire) { + tokio::task::yield_now().await; + } + }) + .await + }); + drop(caller); + assert!(supervisors_stopped.is_ok() && native_finished.is_ok() && !finished_early); + assert!(observed.admission_closed()); + assert_eq!(observed.accepted_jobs(), 0); + assert_eq!(observed.unjoined_jobs(), 1); + assert!(!observed.locally_joined()); + assert!(matches!( + early_close, + Ok(Err(Error::Control( + "Blob artifact original native join is unproven" + ))) + )); + assert!(matches!( + observed.first_failure().unwrap().as_ref(), + Error::Control("Blob artifact supervisor ended before original native join") + )); + current.block_on(async { + let raw = provider + .inner + .get(&store.part_path(&part_digest(b"unfinished"))) + .await + .unwrap() + .bytes() + .await + .unwrap(); + assert_eq!(raw.as_ref(), b"unfinished"); + // Provider completion after abandonment cannot reconstruct the original + // join or erase its retained unknown result/ownership obligation. + assert!(store.close_and_join().await.is_err()); + assert!(!store.lifecycle_observation().unwrap().locally_joined()); + }); +} + +#[tokio::test] +async fn malformed_inputs_never_start_native_work_or_consume_admission() { + let (store, provider) = fixture(); + assert!(store.put_part([0; 32], b"mismatch").await.is_err()); + let oversized = vec![0; MAX_BLOB_PART_BYTES + 1]; + assert!( + store + .put_part(part_digest(&oversized), &oversized) + .await + .is_err() + ); + assert!( + store + .read_part([0; 32], (MAX_BLOB_PART_BYTES + 1) as u32) + .await + .is_err() + ); + assert!( + store + .sweep_unreferenced(Arc::new(BTreeSet::new()), -1) + .await + .is_err() + ); + assert_eq!(store.lifecycle_observation().unwrap().accepted_jobs(), 0); + assert!( + store + .lifecycle_observation() + .unwrap() + .first_failure() + .is_none() + ); + assert_eq!(provider.inner.list(None).collect::>().await.len(), 0); + assert!(join(&store).await.locally_joined()); +} diff --git a/crates/cellule-runtime/src/primitives/blob/tests.rs b/crates/cellule-runtime/src/primitives/blob/tests.rs index fb3b035a..e523ecd9 100644 --- a/crates/cellule-runtime/src/primitives/blob/tests.rs +++ b/crates/cellule-runtime/src/primitives/blob/tests.rs @@ -133,7 +133,7 @@ async fn object_store_sweep_keeps_live_parts_and_reclaims_old_orphans() { .unwrap(); let report = artifacts - .sweep_unreferenced(&BTreeSet::from([live_digest]), i64::MAX) + .sweep_unreferenced(std::sync::Arc::new(BTreeSet::from([live_digest])), i64::MAX) .await .unwrap(); assert_eq!(report.scanned(), 2); @@ -174,7 +174,7 @@ async fn object_store_sweep_reaches_orphans_beyond_live_entries() { .unwrap(); let report = artifacts - .sweep_unreferenced(&live_digests, i64::MAX) + .sweep_unreferenced(std::sync::Arc::new(live_digests), i64::MAX) .await .unwrap(); assert_eq!(report.scanned(), 130); @@ -198,14 +198,14 @@ async fn object_store_sweep_bounds_deletions_and_finishes_on_retry() { } let first = artifacts - .sweep_unreferenced(&BTreeSet::new(), i64::MAX) + .sweep_unreferenced(std::sync::Arc::new(BTreeSet::new()), i64::MAX) .await .unwrap(); assert_eq!(first.deleted(), MAX_BLOB_GC_DELETIONS); assert!(first.has_more()); let second = artifacts - .sweep_unreferenced(&BTreeSet::new(), i64::MAX) + .sweep_unreferenced(std::sync::Arc::new(BTreeSet::new()), i64::MAX) .await .unwrap(); assert_eq!(second.deleted(), 1); diff --git a/crates/cellule-runtime/tests/primitives/blob_cron/lifecycle/mod.rs b/crates/cellule-runtime/tests/primitives/blob_cron/lifecycle/mod.rs new file mode 100644 index 00000000..2d5cc0f7 --- /dev/null +++ b/crates/cellule-runtime/tests/primitives/blob_cron/lifecycle/mod.rs @@ -0,0 +1,505 @@ +//! Original public Blob operations across metadata, provider I/O and caller loss. +use super::*; + +mod prepared; +use cellule_runtime::BlobArtifactStore; +use cellule_runtime::client::InvocationError; +use futures_util::stream::BoxStream; +use object_store::ObjectStore; +use std::{ + fmt, + sync::atomic::{AtomicBool, AtomicU8, AtomicUsize, Ordering}, + time::Duration, +}; +use tokio::sync::Notify; + +#[derive(Default, Debug)] +struct Gate { + kind: AtomicU8, + entered: AtomicUsize, + released: AtomicBool, + changed: Notify, + resume: Notify, +} +impl Gate { + async fn wait(&self, count: usize) -> bool { + tokio::time::timeout(Duration::from_secs(5), async { + loop { + let changed = self.changed.notified(); + tokio::pin!(changed); + changed.as_mut().enable(); + if self.entered.load(Ordering::Acquire) >= count { + return; + } + changed.await; + } + }) + .await + .is_ok() + } + fn release(&self) { + self.released.store(true, Ordering::Release); + self.resume.notify_waiters(); + } + async fn before(&self, kind: u8) { + if self.kind.load(Ordering::Acquire) != kind { + return; + } + self.entered.fetch_add(1, Ordering::AcqRel); + self.changed.notify_waiters(); + loop { + let resume = self.resume.notified(); + tokio::pin!(resume); + resume.as_mut().enable(); + if self.released.load(Ordering::Acquire) { + return; + } + resume.await; + } + } +} +#[derive(Debug)] +struct Provider { + inner: InMemory, + gate: Gate, +} +impl fmt::Display for Provider { + fn fmt(&self, out: &mut fmt::Formatter<'_>) -> fmt::Result { + out.write_str("public-blob-lifecycle") + } +} +#[async_trait::async_trait] +impl ObjectStore for Provider { + async fn put_opts( + &self, + path: &Path, + payload: object_store::PutPayload, + options: object_store::PutOptions, + ) -> object_store::Result { + self.gate.before(1).await; + self.inner.put_opts(path, payload, options).await + } + async fn put_multipart_opts( + &self, + path: &Path, + options: object_store::PutMultipartOptions, + ) -> object_store::Result> { + self.inner.put_multipart_opts(path, options).await + } + async fn get_opts( + &self, + path: &Path, + options: object_store::GetOptions, + ) -> object_store::Result { + self.gate.before(2).await; + self.inner.get_opts(path, options).await + } + fn delete_stream( + &self, + paths: BoxStream<'static, object_store::Result>, + ) -> BoxStream<'static, object_store::Result> { + self.inner.delete_stream(paths) + } + fn list( + &self, + prefix: Option<&Path>, + ) -> BoxStream<'static, object_store::Result> { + self.inner.list(prefix) + } + async fn list_with_delimiter( + &self, + prefix: Option<&Path>, + ) -> object_store::Result { + self.inner.list_with_delimiter(prefix).await + } + async fn copy_opts( + &self, + from: &Path, + to: &Path, + options: object_store::CopyOptions, + ) -> object_store::Result<()> { + self.inner.copy_opts(from, to, options).await + } +} +struct Fixture { + runtime: CellRuntime, + handle: cellule_runtime::cell::actor::CellHandle, + client: CellClient, + target: CellTarget, + blobs: BlobNamespace, + artifacts: BlobArtifactStore, + provider: Arc, + _directory: tempfile::TempDir, +} +impl Fixture { + async fn new() -> Self { + let registry = registry(); + let tenant = TenantId::from_bytes([100; 16]); + let application = ApplicationId::from_bytes([101; 16]); + let layout = CellStorageLayout::new( + Store::new(Arc::new(InMemory::new())), + Path::from("blob-local-lifecycle"), + *application.as_bytes(), + ); + let target = + CellTarget::new(tenant, application, BLOB_NAMESPACE, &0_u32.to_be_bytes()).unwrap(); + let incarnation = IncarnationId::from_bytes([102; 16]); + let proof = CellCatalog::new(layout.clone(), tenant) + .provision( + CatalogEntry::new( + &target, + CatalogRole::Blob, + registry.module_code(BLOB_MODULE).unwrap(), + 1, + ) + .unwrap(), + ) + .await + .unwrap(); + let authority = CellAuthority::new(layout.clone()); + let session = SessionId::from_bytes([103; 16]); + let control = authority + .create_initial( + &proof, + incarnation, + Owner { + session, + endpoint: "https://blob-lifecycle.internal:8081".into(), + }, + ) + .await + .unwrap(); + let runtime = super::maintenance::node_runtime(session); + let directory = tempfile::TempDir::new().unwrap(); + let handle = runtime + .bootstrap( + proof, + CellReplica::new( + layout, + *target.cell_id().as_bytes(), + *incarnation.as_bytes(), + Limits::default(), + ) + .unwrap(), + authority, + control, + directory.path().join("blob.sqlite"), + cellule_runtime::primitives::blob::install_blob_schema, + ) + .await + .unwrap(); + let provider = Arc::new(Provider { + inner: InMemory::new(), + gate: Gate::default(), + }); + let artifacts = BlobArtifactStore::new(Store::new(provider.clone())); + let client = + CellClient::local(registry, handle.clone()).with_blob_artifact_store(artifacts.clone()); + let blobs = BlobNamespace::new(client.clone(), tenant, application).unwrap(); + Self { + runtime, + handle, + client, + target, + blobs, + artifacts, + provider, + _directory: directory, + } + } + async fn mutate( + &self, + id: u8, + mutation: BlobMutation, + ) -> cellule_runtime::Committed { + let now = now_ms(); + let committed = self + .blobs + .mutate(mutation_identity_window(id, now, now + 60_000), mutation) + .await + .unwrap(); + idle(&self.artifacts).await; + committed + } + async fn begin(&self) { + self.mutate( + 104, + BlobMutation::Begin { + key: b"key".to_vec(), + upload_id: [105; 16], + condition: BlobCondition::Any, + content_type: None, + metadata: vec![], + expires_at_ms: now_ms() + 120_000, + }, + ) + .await; + } + async fn finish(&self) { + self.handle.drain().await.unwrap(); + self.runtime.shutdown().await.unwrap(); + assert_eq!(self.runtime.stats().active_cells(), 0); + assert_eq!(self.runtime.stats().worker_jobs(), 0); + assert_eq!(self.runtime.stats().retained_bytes(), 0); + assert_eq!(self.runtime.stats().local_disk_reserved_bytes(), 0); + } +} +// A reply can arrive before its original supervisor drops undelivered output +// and releases its job. Join prior fixture work before attributing a later +// paused job's exact count; no other producer is running during this setup. +async fn idle(store: &BlobArtifactStore) { + tokio::time::timeout(Duration::from_secs(5), async { + while store.lifecycle_observation().unwrap().accepted_jobs() != 0 { + tokio::task::yield_now().await; + } + }) + .await + .unwrap(); +} + +async fn joined( + store: &BlobArtifactStore, +) -> cellule_runtime::primitives::blob::BlobArtifactLifecycleObservation { + tokio::time::timeout(Duration::from_secs(5), store.close_and_join()) + .await + .unwrap() + .unwrap() +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn original_range_lifetime_covers_metadata_and_all_parts_after_close_and_waiter_loss() { + for cancel in [false, true] { + let fixture = Fixture::new().await; + fixture.begin().await; + for (part_number, payload, id) in [(1, b"abc".to_vec(), 106), (2, b"def".to_vec(), 107)] { + fixture + .mutate( + id, + BlobMutation::PutPart { + key: b"key".to_vec(), + upload_id: [105; 16], + part_number, + payload, + }, + ) + .await; + } + let committed = fixture + .mutate( + 108, + BlobMutation::Complete { + key: b"key".to_vec(), + upload_id: [105; 16], + part_count: 2, + }, + ) + .await; + let started = Arc::new(Notify::new()); + let (release, receive) = std::sync::mpsc::channel(); + let blocked_metadata = { + let handle = fixture.handle.clone(); + let started = started.clone(); + tokio::spawn(async move { + handle + .query(1, 1, move |_| { + started.notify_one(); + receive.recv().unwrap(); + Ok(vec![1]) + }) + .await + }) + }; + started.notified().await; + fixture.provider.gate.kind.store(2, Ordering::Release); + let caller = { + let blobs = fixture.blobs.clone(); + tokio::spawn(async move { + blobs + .query( + BlobQuery::Read { + key: b"key".to_vec(), + offset: 1, + limit: 4, + }, + Some(committed.receipt), + ) + .await + }) + }; + // Observe admission before the held real SQL query's existing deadline. + let accepted = tokio::time::timeout(Duration::from_millis(500), async { + while fixture + .artifacts + .lifecycle_observation() + .unwrap() + .accepted_jobs() + != 1 + { + tokio::task::yield_now().await; + } + }) + .await; + let reads_before_metadata = fixture.provider.gate.entered.load(Ordering::Acquire); + fixture.artifacts.close(); + let caller = if cancel { + caller.abort(); + assert!(caller.await.unwrap_err().is_cancelled()); + None + } else { + Some(caller) + }; + release.send(()).unwrap(); + blocked_metadata.await.unwrap().unwrap(); + let provider_entered = fixture.provider.gate.wait(1).await; + let close = { + let store = fixture.artifacts.clone(); + tokio::spawn(async move { store.close_and_join().await }) + }; + close.abort(); + let close_cancelled = close.await.is_err_and(|error| error.is_cancelled()); + let pending = fixture + .artifacts + .lifecycle_observation() + .unwrap() + .accepted_jobs(); + fixture.provider.gate.release(); + let observed = joined(&fixture.artifacts).await; + let output = match caller { + Some(caller) => Some(caller.await), + None => None, + }; + fixture.finish().await; + // Resume/join the original worker and all provider I/O before failure + // assertions, including when ownership admission regresses. + assert!(accepted.is_ok() && provider_entered && close_cancelled); + assert_eq!(reads_before_metadata, 0); + assert_eq!(pending, 1); + assert!(observed.locally_joined() && observed.first_failure().is_none()); + assert_eq!(fixture.provider.gate.entered.load(Ordering::Acquire), 2); + if let Some(output) = output { + assert!( + matches!(output.unwrap().unwrap().output, BlobQueryResult::Read(Some(read)) if read.bytes == b"bcde") + ); + } + assert!(matches!( + fixture + .blobs + .query( + BlobQuery::Head { + key: b"key".to_vec() + }, + None + ) + .await, + Err(InvocationError::NotStarted( + cellule_runtime::Error::CellDraining + )) + )); + assert!(matches!( + fixture.blobs.list_shard(0, vec![], None, 1, None).await, + Err(InvocationError::NotStarted( + cellule_runtime::Error::CellDraining + )) + )); + } +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn original_part_mutation_joins_manifest_publication_after_cancelled_upload_waiter() { + let fixture = Fixture::new().await; + fixture.begin().await; + let payload = b"accepted-part".to_vec(); + let mut digest = blake3::Hasher::new(); + digest.update(b"crab.blob-part.v1\0"); + digest.update(&payload); + let now = now_ms(); + let identity = mutation_identity_window(109, now, now + 60_000); + let prepared = fixture + .client + .prepare_command::>( + &fixture.target, + identity, + BlobMutation::PutPartRef { + key: b"key".to_vec(), + upload_id: [105; 16], + part_number: 1, + digest: *digest.finalize().as_bytes(), + size: payload.len() as u32, + }, + ) + .await + .unwrap(); + let evidence = prepared.evidence().clone(); + fixture.provider.gate.kind.store(1, Ordering::Release); + let caller = { + let blobs = fixture.blobs.clone(); + tokio::spawn(async move { + blobs + .mutate( + identity, + BlobMutation::PutPart { + key: b"key".to_vec(), + upload_id: [105; 16], + part_number: 1, + payload, + }, + ) + .await + }) + }; + let provider_entered = fixture.provider.gate.wait(1).await; + caller.abort(); + assert!(caller.await.unwrap_err().is_cancelled()); + fixture.artifacts.close(); + let pending = fixture + .artifacts + .lifecycle_observation() + .unwrap() + .accepted_jobs(); + fixture.provider.gate.release(); + assert!(joined(&fixture.artifacts).await.first_failure().is_none()); + assert!(provider_entered); + assert_eq!(pending, 1); + // Observe before any duplicate dispatch: the retained original operation + // must have published its manifest after its public caller disappeared. + assert!(matches!( + fixture.client.resolve(&evidence).await.unwrap(), + cellule_runtime::Resolution::Committed(_) + )); + // Closure fences even retained exact prepared replays. Original committed + // bytes and sequence are still resolved through the canonical request log. + assert!(matches!( + prepared.execute().await, + Err(InvocationError::NotStarted( + cellule_runtime::Error::CellDraining + )) + )); + let cellule_runtime::Resolution::Committed(outcome) = + fixture.client.resolve(&evidence).await.unwrap() + else { + panic!("original accepted manifest outcome"); + }; + let mut decoder = cellule_runtime::codec::BoundedDecoder::new(outcome.result(), 1024).unwrap(); + let output = + ::decode(&mut decoder).unwrap(); + decoder.finish().unwrap(); + assert!(matches!(output, BlobMutationOutcome::PartStored { .. })); + assert_eq!(outcome.commit_sequence(), 2); + assert_eq!(fixture.provider.gate.entered.load(Ordering::Acquire), 1); + assert!(matches!( + fixture + .blobs + .prepare_mutation( + identity, + BlobMutation::Abort { + key: b"key".to_vec(), + upload_id: [105; 16] + } + ) + .await, + Err(InvocationError::NotStarted( + cellule_runtime::Error::CellDraining + )) + )); + fixture.finish().await; +} diff --git a/crates/cellule-runtime/tests/primitives/blob_cron/lifecycle/prepared.rs b/crates/cellule-runtime/tests/primitives/blob_cron/lifecycle/prepared.rs new file mode 100644 index 00000000..413c38f7 --- /dev/null +++ b/crates/cellule-runtime/tests/primitives/blob_cron/lifecycle/prepared.rs @@ -0,0 +1,189 @@ +//! Prepared Blob dispatch shares the original store admission and native owner. +use super::*; +use cellule_runtime::primitives::blob::BlobCommand; + +async fn part( + fixture: &Fixture, + id: u8, +) -> cellule_runtime::client::PreparedCommand> { + let now = now_ms(); + let prepared = fixture + .blobs + .prepare_mutation( + mutation_identity_window(id, now, now + 60_000), + BlobMutation::PutPart { + key: b"key".to_vec(), + upload_id: [105; 16], + part_number: 1, + payload: b"prepared-part".to_vec(), + }, + ) + .await + .unwrap(); + idle(&fixture.artifacts).await; + prepared +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn closed_original_store_refuses_prepared_clones_and_restored_dispatch() { + let fixture = Fixture::new().await; + fixture.begin().await; + fixture.provider.gate.kind.store(1, Ordering::Release); + fixture.provider.gate.release(); + let prepared = part(&fixture, 110).await; + let evidence = prepared.evidence().clone(); + let body = prepared.input_bytes().to_vec(); + let header = prepared.snapshot().to_bytes().unwrap(); + let snapshot = cellule_runtime::client::PreparedCommandSnapshot::from_bytes(&header).unwrap(); + assert_eq!(snapshot.to_bytes().unwrap(), header); + let restored = fixture + .client + .restore_command::>(snapshot.clone(), body.clone()) + .unwrap(); + assert_eq!(restored.evidence(), &evidence); + assert_eq!(restored.input_bytes(), body); + let clone = prepared.clone(); + let closed = joined(&fixture.artifacts).await; + // Import still preserves evidence after closure without dispatching or + // changing the original identity, digest, body or durable snapshot bytes. + let after_close = fixture + .client + .restore_command::>(snapshot, body) + .unwrap(); + let mut refusals = Vec::new(); + for command in [prepared, clone, restored, after_close] { + refusals.push(matches!( + command.execute().await, + Err(InvocationError::NotStarted( + cellule_runtime::Error::CellDraining + )) + )); + } + let resolution = fixture.client.resolve(&evidence).await.unwrap(); + let final_observation = fixture.artifacts.lifecycle_observation().unwrap(); + fixture.finish().await; + assert!(closed.locally_joined() && final_observation.locally_joined()); + assert!(refusals.into_iter().all(|refused| refused)); + assert_eq!(resolution, cellule_runtime::Resolution::Absent); + assert!(final_observation.first_failure().is_none()); + assert_eq!(fixture.provider.gate.entered.load(Ordering::Acquire), 1); +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn accepted_prepared_blob_dispatch_survives_caller_loss_and_store_closure() { + let fixture = Fixture::new().await; + fixture.begin().await; + let prepared = part(&fixture, 111).await; + let evidence = prepared.evidence().clone(); + let replay = prepared.clone(); + let started = Arc::new(Notify::new()); + let (release, receive) = std::sync::mpsc::channel(); + let held = { + let handle = fixture.handle.clone(); + let started = started.clone(); + tokio::spawn(async move { + handle + .query(1, 1, move |_| { + started.notify_one(); + receive.recv().unwrap(); + Ok(vec![1]) + }) + .await + }) + }; + started.notified().await; + let caller = tokio::spawn(async move { prepared.execute().await }); + let accepted = tokio::time::timeout(Duration::from_millis(500), async { + while fixture + .artifacts + .lifecycle_observation() + .unwrap() + .accepted_jobs() + != 1 + { + tokio::task::yield_now().await; + } + }) + .await + .is_ok(); + fixture.artifacts.close(); + caller.abort(); + let caller_cancelled = caller.await.is_err_and(|error| error.is_cancelled()); + let pending = fixture + .artifacts + .lifecycle_observation() + .unwrap() + .accepted_jobs(); + let (polled, pending_join) = tokio::sync::oneshot::channel(); + let close = { + let store = fixture.artifacts.clone(); + tokio::spawn(async move { + let join = store.close_and_join(); + tokio::pin!(join); + let mut polled = Some(polled); + std::future::poll_fn(|context| { + let progress = std::future::Future::poll(join.as_mut(), context); + if progress.is_pending() + && let Some(polled) = polled.take() + { + let _ = polled.send(()); + } + progress + }) + .await + }) + }; + let close_polled = tokio::time::timeout(Duration::from_millis(500), pending_join) + .await + .is_ok_and(|result| result.is_ok()); + close.abort(); + let close_cancelled = close.await.is_err_and(|error| error.is_cancelled()); + // Always release and join the real SQL worker before ownership assertions. + release.send(()).unwrap(); + held.await.unwrap().unwrap(); + let observation = joined(&fixture.artifacts).await; + let resolution = fixture.client.resolve(&evidence).await.unwrap(); + let refused = matches!( + replay.execute().await, + Err(InvocationError::NotStarted( + cellule_runtime::Error::CellDraining + )) + ); + fixture.finish().await; + assert!(accepted && caller_cancelled && close_polled && close_cancelled && refused); + assert_eq!(pending, 1); + assert!(observation.locally_joined() && observation.first_failure().is_none()); + assert!( + matches!(resolution, cellule_runtime::Resolution::Committed(outcome) + if outcome.commit_sequence() == 2) + ); +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn open_prepared_blob_replay_preserves_exact_result_without_uploading_again() { + let fixture = Fixture::new().await; + fixture.begin().await; + fixture.provider.gate.kind.store(1, Ordering::Release); + fixture.provider.gate.release(); + let prepared = part(&fixture, 112).await; + let evidence = prepared.evidence().clone(); + let replay = fixture + .client + .restore_command::>( + prepared.snapshot(), + prepared.input_bytes().to_vec(), + ) + .unwrap(); + let committed = prepared.execute().await.unwrap(); + assert!(matches!( + fixture.client.resolve(&evidence).await.unwrap(), + cellule_runtime::Resolution::Committed(_) + )); + let duplicate = replay.execute().await.unwrap(); + assert_eq!(duplicate.receipt, committed.receipt); + assert_eq!(duplicate.output, committed.output); + assert_eq!(committed.receipt.commit_sequence, 2); + assert_eq!(fixture.provider.gate.entered.load(Ordering::Acquire), 1); + assert!(joined(&fixture.artifacts).await.first_failure().is_none()); + fixture.finish().await; +} diff --git a/crates/cellule-runtime/tests/primitives/blob_cron/maintenance/mod.rs b/crates/cellule-runtime/tests/primitives/blob_cron/maintenance/mod.rs index bc624469..cdd3aa28 100644 --- a/crates/cellule-runtime/tests/primitives/blob_cron/maintenance/mod.rs +++ b/crates/cellule-runtime/tests/primitives/blob_cron/maintenance/mod.rs @@ -15,7 +15,7 @@ use cellule_runtime::primitives::effects::{EffectClaimRequest, EffectLeaseOutcom mod peer; -fn node_runtime(session: SessionId) -> CellRuntime { +pub(super) fn node_runtime(session: SessionId) -> CellRuntime { // These are independent nodes. The convenience default Host shares a // process-wide disk budget, including reservations held by other runtimes. CellRuntime::new_with_replica_host( diff --git a/crates/cellule-runtime/tests/primitives/blob_cron/mod.rs b/crates/cellule-runtime/tests/primitives/blob_cron/mod.rs index 27321649..2f7e72a8 100644 --- a/crates/cellule-runtime/tests/primitives/blob_cron/mod.rs +++ b/crates/cellule-runtime/tests/primitives/blob_cron/mod.rs @@ -637,4 +637,5 @@ const fn operation(id: u32, input_limit: u32, output_limit: u32) -> OperationDes } } +mod lifecycle; mod maintenance; diff --git a/crates/cellule-runtime/tests/runtime/lifecycle/maintenance.rs b/crates/cellule-runtime/tests/runtime/lifecycle/maintenance.rs index b98a4360..41b59bc0 100644 --- a/crates/cellule-runtime/tests/runtime/lifecycle/maintenance.rs +++ b/crates/cellule-runtime/tests/runtime/lifecycle/maintenance.rs @@ -2,6 +2,215 @@ use super::*; +fn isolated_replica_host() -> cellule_ltx::Host { + let host = cellule_ltx::Host::default(); + let budget = cellule_ltx::DiskBudget::new(host.local_disk_capacity()); + host.with_local_disk_budget(budget) +} + +#[tokio::test(flavor = "multi_thread", worker_threads = 2)] +async fn publication_borrow_cannot_starve_exact_maintenance_quiescence_or_release() { + use cellule_runtime::cell::actor::{CellInventoryEntry, MaintenanceCellRelease}; + // Model other live tests or runtimes using the process-wide default budget. + // This reservation must survive this runtime's complete shutdown. + let foreign_disk = cellule_ltx::Host::default() + .local_disk_budget() + .try_reserve(4096) + .unwrap(); + for release in [false, true] { + let store = Arc::new(PausingStore::new(Arc::new(InMemory::new()))); + let fixture = fixture_with_limits_and_store( + b"maintenance-during-publication", + Limits::default(), + Store::new(store.clone()), + ); + // Default hosts intentionally share a process-wide disk budget. Give + // each simulated node its own budget so zero verifies its cleanup, + // even while unrelated nodes retain their admitted artifacts. + let session = SessionId::from_bytes([4; 16]); + let runtime = CellRuntime::new_with_replica_host( + SqlWorkerPool::new(2, 10).unwrap(), + 16 << 20, + session, + isolated_replica_host(), + ) + .unwrap(); + let handle = bootstrap_on(&runtime, &fixture, session).await; + let owner = inventory::stable_owner(&runtime).await; + let epoch = handle.owner_fence().epoch; + let identity = mutation_identity_window(96, 10, 10_000); + let digest = Digest::from_bytes([96; 32]); + store.arm_next_update(); + let executing = { + let handle = handle.clone(); + tokio::spawn(async move { + handle + .execute(identity, digest, 20, 64, 64, |transaction| { + transaction.execute("UPDATE counter SET value = value + 1", [])?; + Ok(HandlerOutcome::Success(b"accepted-publication".to_vec())) + }) + .await + }) + }; + tokio::time::timeout( + std::time::Duration::from_secs(3), + store.wait_until_blocked(), + ) + .await + .unwrap(); + let page = runtime.fleet_cells_page(None, 128).await.unwrap(); + let publisher_borrowed = page.entries().iter().any(|entry| matches!(entry, + CellInventoryEntry::Owned(current) if current.target.cell_id() == handle.cell_id() && current.position.is_none() && current.owner_fence == handle.owner_fence())); + drop(page); + let wrong_epoch = runtime + .quiesce_cell_at( + handle.cell_id(), + session, + owner.generation, + owner.incarnation, + epoch + 1, + ) + .await; + let wrong_generation = runtime + .quiesce_cell_at( + handle.cell_id(), + session, + owner.generation + 1, + owner.incarnation, + epoch, + ) + .await; + let releasing = if release { + let runtime = runtime.clone(); + let cell = handle.cell_id(); + Some(tokio::spawn(async move { + runtime + .release_maintenance_cell_at( + cell, + session, + owner.generation, + owner.incarnation, + epoch, + tokio::time::Instant::now() + std::time::Duration::from_secs(10), + ) + .await + })) + } else { + None + }; + let quiesced = if release { + None + } else { + Some( + runtime + .quiesce_cell_at( + handle.cell_id(), + session, + owner.generation, + owner.incarnation, + epoch, + ) + .await, + ) + }; + let closed = tokio::time::timeout(std::time::Duration::from_secs(3), async { + loop { + let page = runtime.fleet_cells_page(None, 128).await.unwrap(); + if page.entries().iter().any(|entry| matches!(entry, + CellInventoryEntry::Owned(current) if current.target.cell_id() == handle.cell_id() && current.quiescing)) { + break; + } + tokio::task::yield_now().await; + } + }).await; + let refused = tokio::time::timeout( + std::time::Duration::from_millis(100), + handle.query(1, 1, |_| Ok(vec![1])), + ) + .await; + let premature_response = executing.is_finished(); + let premature_release = releasing + .as_ref() + .is_some_and(tokio::task::JoinHandle::is_finished); + // Always resume and join native publication before assertions. A failed + // maintenance barrier cannot strand the paused original CAS or SQLite. + store.release(); + let outcome = executing.await.unwrap().unwrap(); + let released = match releasing { + Some(task) => Some(task.await.unwrap().unwrap()), + None => None, + }; + assert!(publisher_borrowed); + assert!(matches!(wrong_epoch, Err(cellule_runtime::Error::Fenced))); + assert!(matches!( + wrong_generation, + Err(cellule_runtime::Error::Fenced) + )); + assert!(closed.is_ok()); + assert!(quiesced.is_none_or(|result| result.is_ok())); + assert!(matches!( + refused, + Ok(Err(cellule_runtime::Error::CellDraining)) + )); + assert!(!premature_response && !premature_release); + let resolving = match released { + Some(MaintenanceCellRelease::Released(position)) => { + assert_eq!(position.root.commit_sequence, outcome.commit_sequence()); + let authority = CellAuthority::new(fixture.layout.clone()); + let idle = authority.load(handle.cell_id()).await.unwrap().unwrap(); + assert_eq!(idle.value().state, ControlState::Idle); + assert_eq!(idle.value().root.as_ref(), Some(&position.root)); + let successor_session = SessionId::from_bytes([97; 16]); + let successor = CellRuntime::new_with_replica_host( + SqlWorkerPool::new(1, 10).unwrap(), + 16 << 20, + successor_session, + isolated_replica_host(), + ) + .unwrap(); + let restored = successor + .acquire_idle_restored( + handle.catalog().clone(), + fixture.replica.clone(), + authority, + idle, + fixture + ._directory + .path() + .join("publication-successor.sqlite"), + Owner { + session: successor_session, + endpoint: "https://successor.internal:8081".into(), + }, + ) + .await + .unwrap(); + assert_eq!( + restored.resolve(identity, digest, 21, 64).await.unwrap(), + Resolution::Committed(outcome.clone()) + ); + restored.drain().await.unwrap(); + successor.shutdown().await.unwrap(); + assert_eq!(successor.stats().retained_bytes(), 0); + assert_eq!(successor.stats().local_disk_reserved_bytes(), 0); + None + } + Some(other) => panic!("publication maintenance refused: {other:?}"), + None => Some(handle.resolve(identity, digest, 21, 64).await.unwrap()), + }; + if let Some(resolved) = resolving { + assert_eq!(resolved, Resolution::Committed(outcome)); + handle.drain().await.unwrap(); + } + runtime.shutdown().await.unwrap(); + assert_eq!(runtime.stats().active_cells(), 0); + assert_eq!(runtime.stats().retained_bytes(), 0); + assert_eq!(runtime.stats().worker_jobs(), 0); + assert_eq!(runtime.stats().local_disk_reserved_bytes(), 0); + assert_eq!(foreign_disk.bytes(), 4096); + } +} + #[tokio::test(flavor = "multi_thread", worker_threads = 2)] async fn maintenance_quiescence_keeps_accepted_work_and_original_resolution() { let fixture = fixture(); diff --git a/docs/fleet-operations-plan.md b/docs/fleet-operations-plan.md index 7b2d72c8..846fb445 100644 --- a/docs/fleet-operations-plan.md +++ b/docs/fleet-operations-plan.md @@ -4,7 +4,7 @@ Status: design and execution plan; partial foundations exist, fleet execution and qualification remain incomplete. Source baseline: `e07670e2348231ed401cc7280a47e3ab97596ffe`. Prepared: September 30, 2026, America/Vancouver. -Updated: October 3, 2026, America/Vancouver. +Updated: October 4, 2026, America/Vancouver. Build one reusable fleet reconciliation path for automatic ownership balancing, sustained pressure relief, and planned node maintenance. Reuse Cellule's fenced @@ -37,7 +37,7 @@ same action and evidence contracts. | Which external facilities are required? | A linearizable journal, complete enrollment registry, authenticated management transport, and trusted local Cell inputs. The reference example supplies these behind the same contracts. | | What keeps operations simple? | One operation ID, one status contract, one application loop, bounded defaults, and one canonical host drain lane. | | What proves success? | Receipt-preserving actor activation on another node. Maintenance additionally requires settled role obligations, host Stopped, and session withdrawal. | -| What is the next concrete change? | Qualify the current head, including both routing profiles. Complete durable provider/process retention and remaining failed-owner role succession. Native source reader policy now composes exact joining/final-root verification with an actual managed writer and current reader policy. Exact reader removal now guards the original request under the native activation lane and retains its joined prefix; the local source collector supplies current policy while durable process evidence remains separate. Fresh nonexecution confirmation now consumes independently joined original witnesses; terminal rows alone remain unknown. Native donor policy lookup discovers latest history through the existing authenticated reader/follower verifiers. Canonical recovered-follower retirement plus exact failed-boot closure now closes the matching source-side follower request, but a recovered log alone stays blocked. Complete observation around original-writer successors, failed-boot closures and durable unknown work. Connect reader/follower replacement policy and accepted-work barriers to SettleRoles/Finalize. Keep cross-session recovery, process/provider and remaining failure/inspection gates in scope. | +| What is the next concrete change? | Qualify routed receiver recovery with inherited overlays and stale generations. Real-node cases now cover interrupted basis/evidence writes and replies, cancellation before/after takeover, safe Idle resumption, an ordinary acquisition winner, immutable reconstruction and refusal with missing/corrupt/substituted native history. Complete the maintenance role/fault matrix, production process and transport adapters, W9 load and mixed-version qualification, and W10 rollout/runbooks. | Start with the [first slice commit sequence](#first-slice-commit-sequence). Every increment must expose a reviewable public behavior and retain its test @@ -102,12 +102,12 @@ package, reuse matching work, and run its gates before marking it complete. | Package | Implemented foundation | Remaining exit evidence | | --- | --- | --- | | W1 | Pure controller/maintenance/movement transitions; bounded versioned journal codecs; unknown-outcome permit retention; action envelopes, physical-node intent, progress pages, immutable accepted-action records, and distinct recovery basis/evidence/outcomes with complete replay comparisons. Registry records carry boot-bound enrollment, retained intent pages, a bootstrap revision, stop/resume state, and transactional allocation gates. The embedding example supplies one SQLite transaction domain for all three journal contracts, with independent-client races, lost replies and reconstruction tests. Fresh inspection requests/observations bind nonce, complete action, registry revision, endpoint and original capture interval. | Complete fleet/role observation envelopes and enrollment producer wiring; consume fresh inspections and the durable adapter through the public reconciler. Process/fault/provider qualification remains required. | -| W2 | Schema 3 operational decoding/signing; schema 2 bridge serializer and signing checks; canonical byte comparison after one decoder verification pass; monotonic sample checks; shared sticky cordon/reversible pressure gate; writer/read/follower/recovery candidate filtering; host follower gate installation; public pressure/cordon lifecycle test. | Complete signed-snapshot production, broader public node admission scenarios, and the mixed-binary rollout campaign. | +| W2 | Schema 3 operational decoding/signing; schema 2 bridge serializer and signing checks; canonical byte comparison after one decoder verification pass; fresh directory reads and create/heartbeat emission reuse that exact immutable authentication; monotonic sample checks; shared sticky cordon/reversible pressure gate; writer/read/follower/recovery candidate filtering; host follower gate installation; public pressure/cordon lifecycle test. | Complete signed-snapshot production, broader public node admission scenarios, and the mixed-binary rollout campaign. | | W3 | Bounded actor ownership pages; worker measurements and conservative costs; residence and stable samples; mutation revision checks; persisted follower-lane pages; expired/fenced log-reference discovery; host managed-reader pages with canonical accepted-work lifetimes, local closure observations and complete bounded native-state continuation fingerprints. | Complete the observer's membership/enrollment barrier and host action consumption, including accepted queries and retained peer views. Run all W3 gates against the current diff before marking the package complete. | | W4 | Runtime receiver preparation reserves actual Cell, memory, descriptor, affine SQL-job, and scoped LTX disk credit. The host journals source/receiver effects, confirms acquisition/recovery input before CAS and recovery position before actor admission, checks current serving, and owns work across dropped waiters. CleaningReceiver preserves release evidence and charged permits; unused credit can retire before ordinary admitted activation resumes. Public local tests cover receipt preservation, refusal, duplicates, basis faults, post-release cleanup, source loss, and ordinary-acquisition races. A separate fresh inspection path bypasses historical Inspect results and cannot start recovery; its finite jobs share the existing action bound and drain. Canonical runtime acquisition retains immutable exact claim input/materialization before admission. Native prepared-root lineage proves an exact per-movement released/recovered prefix across publication and compaction; complete bounded origin verification and fresh actor/authority checks gate serving evidence. Full original-writer/suffix aggregation remains required. | Complete recovery across receiver sessions and refusal/unknown reconciliation, complete observer consumption, maintenance role actions, durable production adapters, and all process/provider and W4 exit assertions. | | W5 | Public caller-driven reconciler and observation/transport contracts; production calls to the existing planner and reducer; phase CAS before dispatch, fresh serving checks, cooldown/post-batch journal reads, and permit projections. SQLite-backed sequencing tests exercise simulated effects, lost replies, deadlines and competing drivers. The overload executable moves two real Cells across three leased runtimes; controller-restart changes claimant after actual lease expiry and adopts lost releases. Maintenance now dispatches exact journal-bound Cordon, commits Cordoned then BeginEvacuation, and applies retained intent before donor selection. Partial fresh observations can evacuate settled normal-pressure donors; operation deadlines bound new moves. | Complete producer/observer wiring and the full W5 failure/concurrency evidence; source/receiver failure adoption, cross-session recovery, busy maintenance and all W5 exit evidence. | | W6 | Node-owned monotonic Cordon closes the shared role gate. Driver retries/adopts lost results and dispatches explicit busy maintenance release using the peak receiver envelope. Canonical quiescence retains accepted foreground and native primitive completion. Public SQL, Queue, Effect, Activity and Workflow cases cover selected receipt/lease/expiry/waiter faults. Configured host startup holds all new roles until atomic Established boot/current intent confirmation and required probes; retained drain exposes management without serving. The example wires actual signed canonical boot enrollment and joined withdrawal/retirement. The existing reader loop now repairs retained producer requests and fenced views without another hint; metadata collection survives pre-lease startup and local fencing while native admission remains lease checked. | Complete primitive/fault matrix, Cron and Blob external owners, failed-owner producer reconciliation; complete reader replacement/failed-process evidence, ongoing intent supervision, sustained traffic, and all W6 assertions. | -| W7–W10 | Prepared follower ensembles retain signed boots before their original-token CAS; conditional refusal competes with that same write, while absence remains unknown. Confirmed member retirement retains original fences across ambiguous authority closure and shares canonical object coverage and authority closure. Epoch-bound host requests wake the existing supervisor, retain strict retries, reject stale/foreign replacement bindings and preserve accepted cleanup across cancellation and host deadlines. Bounded weak inspection handles expose local completion without certifying fleet settlement. The managed follower producer accepts every original member Pending before its one CAS, retains unknown results in the existing supervisor, and publishes original establishment/confirmed retirement events before releasing inventory. Atomic nonexecution exclusions fence delayed reader/follower acceptance. | Complete role observation, replacement-policy and failed-owner evidence, role evacuation/finalization, complete maintenance and failure examples, fault qualification, and runbooks. | +| W7–W10 | Prepared follower ensembles retain signed boots before their original-token CAS; conditional refusal competes with that same write, while absence remains unknown. Confirmed member retirement retains original fences across ambiguous authority closure and shares canonical object coverage and authority closure. Epoch-bound host requests wake the existing supervisor, retain strict retries, reject stale/foreign replacement bindings and preserve accepted cleanup across cancellation and host deadlines. Bounded weak inspection handles expose local completion without certifying fleet settlement. The managed follower producer accepts every original member Pending before its one CAS, retains unknown results in the existing supervisor, and publishes original establishment/confirmed retirement events before releasing inventory. Atomic nonexecution exclusions fence delayed reader/follower acceptance. Fleet Finalize closes admission, joins accepted action work through the canonical drain, and publishes Stopped only after node shutdown and exact boot withdrawal. Closing-phase reconciliation reloads after boot retirement advances the registry and revalidates the operation/evidence before completion. Minion scenarios now cover writer-only, managed-reader, live-follower and failed-owner follower-only maintenance through the public reconciler; dead-owner SettleRoles and Finalize require exact failed-boot/recovered-follower closures and never contact the retired endpoint. | Complete controller-restart maintenance scenarios; expand role and fault combinations, including cancellation/provider/process failure; qualify broad provider/process behavior and both routing profiles; finish W9 load/mixed-version qualification and W10 rollout/runbooks. | This planning pass does not certify the Rust implementation or provider/process behavior. Record verification against the exact revision and diff that ran; @@ -120,12 +120,16 @@ its source fingerprint, selected commands, and limits; it does not mark an entire work package complete. The [fleet journal example](../crates/cellule-host/minion/README.md) -currently supports `inspect-journal `, `overload`, and -`controller-restart`. The first reopens the local durable reference journal. -The latter commands exercise real bounded movement, including a new controller -epoch after lost source replies and actual lease expiry. Maintenance and -receiver-loss scenarios, complete production observations, and the full W8 -exit requirements remain deliverables. +currently supports `inspect-journal `, `overload`, +`controller-restart`, `receiver-loss`, count `balance`, writer-only +`maintenance`, and reader-only `maintenance-reader`. The first reopens the local durable reference journal. Movement +commands exercise real bounded placement, +including a new controller epoch after lost source replies and actual lease +expiry. Live-reader maintenance is runnable through the CLI; follower and failed-owner +maintenance have focused minion scenarios. Receiver loss joins the failed receiver, publishes exact boot +closure, routes the original release, loses and replays activation, and rebuilds +the controller from retained state before checking all twelve receipts. Complete +production observations and the full W8 exit requirements remain deliverables. | Start here | Contents | | --- | --- | @@ -290,11 +294,11 @@ present API; the pseudocode below specifies the remaining orchestration. | `FollowerEvacuationRecord` / `FleetFollowerEvacuationJournal` / `FleetFollowerEvacuationVerifier` | Present bounded durable live-owner ensemble history, revisioned application policy, atomic latest pointers and canonical/native revalidation/refresh. | Bind every original Retired member and replacement request, confirm current canonical ensemble plus the actual source/member native owners through `FleetSnapshotTransport`. Refresh after policy/epoch/operation changes without repeating original retirement. Complete failed-owner cases outside the recovered-log path, full observer and SettleRoles/Finalize integration; this per-owner evidence cannot finish maintenance. | | `FleetObservation::with_role_coverage` | Present retained graph attachment and planner-input binding. | Preserve the original interval and reject replacement or scope/registry mismatches. The reconciler rechecks the graph's exact head/registry and full roster digest; unchanged registry alone cannot admit an earlier controller head. Attachment never upgrades partial adapter coverage. | | `MaintenanceEnrollmentInventory`, `FleetJournal::maintenance_enrollments` and `FleetMaintenanceEnrollments::collect` | Present bounded original Pending/Established reader/follower pages frozen with the first BeginEvacuation CAS, at either physical endpoint including earlier boots. | Retain exact pretransition head, registry, operation and acceptance times. Traverse every page, match original acceptances against the fresh full roster, and attach with `FleetObservation::with_maintenance_enrollments`. Missing history remains unknown; zero requires a committed empty manifest. New source replacement work after capture remains in the current graph. All original and current policy/work obligations remain required before settlement. | -| `FleetObservation::with_role_evacuations` | Present retention of freshly rechecked durable reader/follower policy history, current authority/prefixes and native source/member evidence. Planner digest v14 binds complete original checks and the immutable maintenance enrollment manifest. | Require one exact full head, registry and roster, with every original fresh interval inside the outer observation. Compare signed replacement boots with the existing producer identities and refuse duplicate obligations. Public minion reconciliation consumes actual ready readers and rotated ensembles without upgrading partial coverage. Complete authenticated policy coverage and accepted-work/finalization barriers remain required. | -| `FleetMaintenanceNonexecution` / `FleetEnrollmentNonexecution` | Present independent read-only confirmation of retained original nonexecution, exact terminal witnesses and the complete original/current request set. Whole-capture clocks/deadlines and full head/registry rechecks refuse missing or changed evidence; planner digest v14 binds checked results. | Applications authenticate original endpoints, earlier native/external joining and durable exclusion evidence. Actual reader/follower cases consume joined original witnesses; production provider/process retention, installed-role succession, other accepted work and finalization remain required. | -| `FleetObservation::check_maintenance_policies` / `FleetReconcileReport::maintenance_policy` | Present bounded matching of every immutable original request and every currently unresolved related reader/follower against all supplied native policy checks. Exact current rows and Established donor original digests are required. Planner digest v14 binds source inputs and every explicit result. An exact canonical recovered-follower closure joined with its failed-boot process closure closes only the corresponding failed-owner source requests. | Attach original capture and all policy checks before matching. Missing history refuses matching; missing policy, Pending/Established, source succession and unproven nonexecution remain explicit gaps. Advisory counts retain their own full head/registry/interval; absent coverage stays unknown. Complete accepted-work and finalization evidence remain required. | +| `FleetObservation::with_role_evacuations` | Present retention of freshly rechecked durable reader/follower policy history, current authority/prefixes and native source/member evidence. Planner digest v15 binds complete original checks and the immutable maintenance enrollment manifest. | Require one exact full head, registry and roster, with every original fresh interval inside the outer observation. Compare signed replacement boots with the existing producer identities and refuse duplicate obligations. Public minion reconciliation consumes actual ready readers and rotated ensembles without upgrading partial coverage. Complete authenticated policy coverage and accepted-work/finalization barriers remain required. | +| `FleetMaintenanceNonexecution` / `FleetEnrollmentNonexecution` | Present independent read-only confirmation of retained original nonexecution, exact terminal witnesses and the complete original/current request set. Whole-capture clocks/deadlines and full head/registry rechecks refuse missing or changed evidence; planner digest v15 binds checked results. | Applications authenticate original endpoints, earlier native/external joining and durable exclusion evidence. Actual reader/follower cases consume joined original witnesses; production provider/process retention, installed-role succession, other accepted work and finalization remain required. | +| `FleetObservation::check_maintenance_policies` / `FleetReconcileReport::maintenance_policy` | Present bounded matching of every immutable original request and every currently unresolved related reader/follower against all supplied native policy checks. Exact current rows and Established donor original digests are required. Planner digest v15 binds source inputs and every explicit result. An exact canonical recovered-follower closure joined with its failed-boot process closure closes only the corresponding failed-owner source requests. | Attach original capture and all policy checks before matching. Missing history refuses matching; missing policy, Pending/Established, source succession and unproven nonexecution remain explicit gaps. Advisory counts retain their own full head/registry/interval; absent coverage stays unknown. Complete accepted-work and finalization evidence remain required. | | `FleetReaderEvacuationVerifier::collect_maintenance` / `FleetFollowerEvacuationVerifier::collect_maintenance` | Present complete eligible donor history lookup from the canonical original/current request set, followed by existing authenticated native/current-policy confirmation. Exact original digest/current row and full head/registry/roster checks are retained. | Collect original history first and construct the outer observation after both role collections. Missing latest history supplies no check and stays MissingPolicy in the matcher. One monotonic 30-second capture and caller deadline bound collection; changed heads/policies during native capture fail closed. Source/failed-owner succession, nonexecution, unknown native/external work and finalization remain required. | -| `FleetReaderEvacuationVerifier::collect_source_readers` / `FleetSourceReaderPolicies` / `FleetObservation::with_source_reader_policies` | Present exact native reader joins or failed-process reader closures composed with an actual managed native writer successor, minimum-prefix/origin verification, current reader policy and ready enrolled replacement prefixes. The complete original/current request set, full head/registry/roster, repeated provider allocation/presence and monotonic capture are retained in planner identity v14. | Authenticate both endpoints and canonical backend mappings. Retain the original native capsule or durable process closure through the application's existing accepted-work owner. Missing evidence stays SourceSuccessor; checked exact source requests become SourceReader. Local SQLite client reconstruction/public reconciliation cannot qualify OS/process restart exclusion or durable provider retention. Other accepted work and SettleRoles/Finalize remain required. | +| `FleetReaderEvacuationVerifier::collect_source_readers` / `FleetSourceReaderPolicies` / `FleetObservation::with_source_reader_policies` | Present exact native reader joins or failed-process reader closures composed with an actual managed native writer successor, minimum-prefix/origin verification, current reader policy and ready enrolled replacement prefixes. The complete original/current request set, full head/registry/roster, repeated provider allocation/presence and monotonic capture are retained in planner identity v15. | Authenticate both endpoints and canonical backend mappings. Retain the original native capsule or durable process closure through the application's existing accepted-work owner. Missing evidence stays SourceSuccessor; checked exact source requests become SourceReader. Local SQLite client reconstruction/public reconciliation cannot qualify OS/process restart exclusion or durable provider retention. Other accepted work and SettleRoles/Finalize remain required. | | `CellNode::follower_evacuation` | Present per-live-owner replacement and retirement check after the original requested rotation. | Require the declared member minimum, donor exclusion, complete Established replacement rows, pinned signed boots, current authority and full journal rechecks. Retain the original completion/error history and publish/revalidate its interval. This does not settle failed owners, Pending producers or the physical node's other obligations. | | Recovered-log authorization, member retirement and canonical Retired CAS | Present runtime tail closure after canonical recovery pinning. | Use `RecoveredNodeLogTransport`, receiver-side `authorize_recovered_log_retire` and `FollowerStore::retire_recovered`; confirm every original member before `retire_recovered_log`. Adopt exact committed closure with `retired_recovered_log` before repeating effects. Publish original enrollments through the host capsule, then complete failed-process barriers and maintenance orchestration; native closure alone cannot finish W7. | | `RecoveryManifestStore::load_manifest` | Complete digest-verified recovered-suffix metadata across every original application, retained after successor materialization. | Read the identity from canonical sealed-log authority, then verify each bundle through its application's ordinary store. Retain the complete original writer set separately, including object-covered Cells without a suffix; check current successor authority, acknowledged prefixes and serving. This metadata read cannot settle roles or finalize a node. | @@ -746,6 +750,84 @@ ledger, expire credit immediately after confirmed release, and drop a cleanup reply. Each test must reach current serving plus confirmed resource settlement, or retain an inspectable charged blocker without claiming cancellation. +### Receiver session loss after release + +The original `MoveAttemptSpec` remains immutable: it records the preferred +receiver boot that admitted the original reservation. A successor receiver is +a separate, journaled continuation of the same charged attempt, never a rewrite +of that spec or a second movement permit. + +1. Reconfirm the exact retired receiver boot against the full roster and durable + process provider. Lease expiry, a missing endpoint, a reboot, or an empty + local ledger is not a cleanup proof. +2. Bind a bounded receiver handoff to the exact previous and next node/session, + the original attempt, the previous boot's closure digest, and the current + registry barrier. Commit it before dispatch. Keep the original release and + all prior accepted action results. +3. Select only an Established, live, Active-intent receiver with fresh capacity + and compatible Cell support. The target runtime performs ordinary resource + admission before any authority CAS; the fleet permit remains charged during + this continuation. +4. Read the current authority before resuming. Idle at or beyond a clean release + uses exact-root acquisition. Serving or Recovering under a closed prior + receiver requires a matching takeover proof and retains the new recovery + basis before successor publication. These outcomes remain distinct. +5. Give the continuation a stable target-bound action identity. A lost reply + replays only at that target. A controller restart recovers the accepted + target/closure chain from durable journal records. A second receiver failure + may continue only after its own fresh process closure; bound the chain and +leave an inspectable charged blocker when the bound is exhausted. + +An interrupted evidence write after native materialization can roll the target +back to a safe Idle root. Replay reconfirms or reconstructs the original receiver +evidence only from the exact canonical acquisition record. It verifies both the +original recovery prefix and clean release prefix before ordinary admitted +acquisition. The retained acceptance, original recovery input/materialization, +capture times and charged permit remain unchanged. A same-session ordinary +acquisition winner must pass those same current-serving checks. Missing or +substituted history remains Unknown; inspection cannot repair records or acquire. + +Failed-source Recover now applies this continuation to its original retained +`RecoveryBasis` and exact native materialization as well. Missing metadata is +reconstructed only from the original immutable acquisition input/result, never +from the current Idle root or a newer owner. For sealed suffixes, the explicit +Idle verifier checks the complete unowned selected control before and after +the bounded original-history/origin walk. Ordinary acquisition winners still +require original native history and fresh actor proof. Public host cases cover +lost replies, repeated write failure, cancelled repair waiters, missing/corrupt +or substituted history, and real sealed suffixes with receipt readback. The full +routed inherited-overlay and controller/provider fault campaign remains open. + +Three selected routed reference cases now inherit an actual sealed follower tail +through an interrupted earlier native claim. They combine evidence-write failure, +safe Idle reacquisition, historical manifest absence/corruption, cancelled replay +waiters, and real controller expiry with an independently reconstructed journal. +They retain the manifest epoch, exact native materialization and original command +receipts, and join original receiver credit before retiring the attempt. Original +follower requests are Pending before canonical enrollment; complete native +retirement and typed publication precede receiver process closure. These are +in-process fault models. Exhaustive provider/backend/lineage faults, successive +receiver-boot failures and measured process/provider qualification remain required. + +An interrupted target may instead retain its own Recovering claim and attached +overlay. Effect replay reconfirms the original fenced input and uses the shared +native takeover continuation without another ownership CAS. The current control +must be the exact original claim or canonical materialized successor. Pinned +suffixes retain their manifest's original Cell epoch, even after takeover +advances the current epoch. Selected public host cases cover repeated manifest +failure, replay waiter cancellation, committed materialization before admission, +changed-input refusal, receipt readback and later historical manifest corruption. +The routed integration/fault campaign still must combine these overlay cases +with boot and controller succession. + +Tests must kill the original receiver after confirmed release, preserve the +acknowledged root, recover on a different live receiver, and verify current +actor/authority, exact action replay, one writer, one charged attempt, resource +settlement, and no dispatch to either retired session. Add pre- and post-CAS +failures, a lost continuation reply, controller restart during handoff, an +ineligible/full receiver, and repeated receiver failure. Provider errors and +cancellation must retain the original source and blocker. + ## Durable controller state and bounded execution Use an application-owned strongly consistent CAS backend. Its logical records @@ -1255,6 +1337,7 @@ The current additions preserve previous layouts and discriminants: | Recover remote action | 7; Retire 6 remains journal-local. | | Recovered outcome | 11 | | RecoveryBasis and RecoveryEvidence record kinds | 10 and 11 | +| ReceiverRecoveryBasis and ReceiverRecoveryEvidence record kinds | 30 and 31; separate from original source recovery. | | RegistryVersion, IntentPage, EnrollmentRecord, EnrollmentPage, and EnrollmentSpec record kinds | 12 through 16 | Only the new Recovered phase carries its recovery evidence extension. Old strict @@ -1461,9 +1544,29 @@ and pure action contracts have focused coverage in the [execution evidence](fleet-operations-progress.md). The driver now accepts busy maintenance demand with a separate configured peak envelope under exact Evacuating intent and fresh source/receiver evidence; ordinary pressure/count -moves still require settled samples. Complete the remaining host fault cases, -every primitive's acceptance matrix, Blob owners and process/provider -qualification before claiming this work package complete. +moves still require settled samples. Canonical minion `maintenance-busy` now +runs two continuous SQL lanes through the public reconciler: original owner-fence +inventory survives publisher borrowing, root changes invalidate count coverage, +receiver preparation precedes native quiescence, and every acknowledged outcome +and exact audit row survives restoration. Publication-paused runtime cases cover +exact epoch/generation refusal and no premature release. + +Blob artifacts now have one shared original-operation admission and close/join +lifecycle. Public mutations retain staging through command response; range reads +retain metadata and all parts, including after caller cancellation or closure. +The existing host facility drain owns the installed store and joins accepted +namespace/GC work before runtime shutdown. GC retains its complete supplied +reference set with shared ownership. Focused tests cover cancelled provider jobs, +capacity, original failure/panic sources, metadata-paused ranges and manifest +publication after upload waiter loss. This supplies local lifetime coverage; +Prepared command dispatch through the configured client now acquires that same +original admission, including clones and restored snapshots; closure refuses +new dispatch without changing request identity or snapshot bytes. Cell-scoped +stream/upload/pin coverage, global retention and unknown remote effects remain +separate obligations. `BlobInventory` +still blocks release. Complete the remaining host fault cases, every primitive's +acceptance matrix, Blob owner/pin barriers and process/provider qualification +before claiming this work package complete. Dependencies: W1, W3, W4. @@ -1484,6 +1587,20 @@ blockers and prevent declaring general maintenance support complete. Dependencies: W2, W3, W5, W6. +Controller reconstruction now refreshes a committed native `SettleRoles` +receipt after a lost reply at the new request's head, retaining the original +acceptance and action key. The minion case checks Closing/Completed, exact boot +withdrawal, canonical-root and command-result readback, and joined resources. +Two focused cases renew the same claimant and replace it after actual lease +expiry. The latter renews the original signed boots while waiting, checks the +new controller epoch, and requires the old controller to be fenced without +changing the journal. Before-write and after-commit result failures now require +the old proof's Conflict, original join/retirement and fresh complete evidence +before Closing. Cancelled-controller cases retain the actual paused native +publication owner and refuse replacement while it is running. Original journal +acceptance and historical receipts remain unchanged. External process/provider +faults and the complete role/primitive matrix still require qualification. + Extend host durability/read-replica orchestration and runtime follower inventory/retirement adapters to evacuate foreign obligations. Add an explicit requested rotation trigger to the existing durability supervisor so maintenance @@ -1509,8 +1626,26 @@ for terminal action handoff. Managed boots can bind their original authenticated directory version and Established registry row before readiness. The native closing task checks canonical withdrawal and committed boot retirement before Stopped, retaining its original evidence across a deadline or ambiguous reply. -This is wired into the reference example. Complete role settlement, the original -action's join before handoff, and committed operation completion remain required. +The Finalize handoff is wired through the native node boundary and reference +transport. It requires the committed Closing evidence and a bound managed-boot +withdrawal, accepts outside the ordinary finite action bank, closes that bank, +joins its accepted work through the canonical drain owner, and publishes Stopped +only after the retained drain confirms shutdown and withdrawal. Retirement +advances the enrollment registry, so Closing-phase reconciliation reloads the +snapshot and revalidates the same operation/evidence before committing +Completed. The reference SQLite journal now refuses ReadyToClose unless the +exact operation/session has a committed SettleRoles result bound to the current +head revision and RegistryVersion. A head or registry change after settlement +fences the transition, and +the legacy inventory-only result remains readable but cannot authorize Closing. +The host SettleRoles executor, its complete inventory/replacement-policy proof, +and the reconciler path that supplies this evidence are now wired. Real reader- +only, live-follower-only and combined reader/follower scenarios verify durable +replacement policy, settle roles, finalize the source boot, and check service +evidence afterward. The combined scenario requires both policies and refuses +reader closure while a selected replacement lacks Established enrollment. +Dead-owner follower maintenance still requires external process/provider +qualification. Exit: maintenance with zero local writers but uncovered foreign follower tails remains blocked; live-owner rotation and dead-owner recovery both clear @@ -1712,13 +1847,21 @@ node crates/cellule-runtime/docs/validate.mjs ``` The following example commands are deliverables of W8. The current tree supports -`overload` and `controller-restart` with the evidence limits recorded in -[execution evidence](fleet-operations-progress.md); `maintenance` and -`receiver-loss` remain unimplemented: +`overload`, `controller-restart`, `receiver-loss`, writer-only `maintenance`, +reader-only `maintenance-reader`, live-follower `maintenance-follower`, and +combined reader/follower `maintenance-roles` with +the evidence limits recorded in [execution evidence](fleet-operations-progress.md). +Reader-only, live-follower-only and combined reader/follower maintenance each +have focused end-to-end minion scenarios, as does failed-owner follower +maintenance. A role-complete maintenance CLI and the full fault/provider +qualification remain required: ```sh cargo run -p cellule-host --example fleet_operations --locked -- overload cargo run -p cellule-host --example fleet_operations --locked -- maintenance +cargo run -p cellule-host --example fleet_operations --locked -- maintenance-reader +cargo run -p cellule-host --example fleet_operations --locked -- maintenance-follower +cargo run -p cellule-host --example fleet_operations --locked -- maintenance-roles cargo run -p cellule-host --example fleet_operations --locked -- controller-restart cargo run -p cellule-host --example fleet_operations --locked -- receiver-loss ``` diff --git a/docs/fleet-operations-progress.md b/docs/fleet-operations-progress.md index 610956b5..fdbaca3a 100644 --- a/docs/fleet-operations-progress.md +++ b/docs/fleet-operations-progress.md @@ -4,6 +4,83 @@ The [implementation plan](fleet-operations-plan.md) remains the full scope. This page records focused checkpoints; it does not establish complete fleet balancing, maintenance, or deployment qualification. +## October 4 2026 managed-reader maintenance checkpoint + +The reference observer now accepts the full enrollment roster, including +non-node roles, and retains the installed native role graph instead of rejecting +readers or follower logs. It collects current reader and follower evacuation +evidence separately from the maintenance replacement-policy check, so missing +policy proof remains a blocker. + +A new minion end-to-end scenario starts with a managed reader on the maintenance +node, activates a replacement on another managed node, publishes and verifies +the durable evacuation policy, then drives SettleRoles and Finalize through the +public reconciler and native node path. It confirms the source boot is stopped, +withdrawn and retired, the old reader no longer serves, and the replacement still +reads the expected value. The focused case passes; the full minion scenario +suite passes (343 passed, 0 failed). + +The same reference observer now handles the four-boot follower fixture. Its +global role coverage check validates an Established replacement lane that has +not appended yet against the original producer and physical follower references; +the generic per-node enrollment check alone rejected that valid empty lane. A +new live-follower-only minion scenario verifies the two replacement members, +reconciles SettleRoles and Finalize, confirms the donor boot is stopped, +withdrawn and retired, and checks the new epoch's exact membership. Its focused +case and formatting pass. At this checkpoint, dead-owner process closure was +still open; the next checkpoint records that path. + +## October 4 2026 failed-owner maintenance checkpoint + +The reconciler now handles a maintenance target whose original boot has already +been retired. `FleetRoleSettlement` binds settlement to the exact failed-boot +closure, canonical fence, full journal head and registry. A dedicated transport +path accepts and publishes the exact `RolesSettledAt` result without sending +SettleRoles to the absent endpoint. Closing uses a matching closed-boot Finalize +path; its default transport fails closed, while the reference SQLite adapter +publishes Stopped only from the current retirement closure and exact maintenance +evidence. Matching durable results are adopted idempotently. + +The new end-to-end case captures the failed process and recovered-follower +closures, collects the original maintenance enrollment set and rechecks native +role inventories plus physical follower references. It reconciles through +Closing to Completed with an empty node endpoint list, proving the retired +session is never contacted. The fixture uses a joined process-lifetime test +stand-in. It also drops the Finalize reply after the exact `Stopped` result is +durably published, waits for the real 2.5-second journal lease to expire, then +starts a different controller session. The replacement replays the same +accepted action from the retained SQLite result and completes the operation; +provider and production process qualification remain open. + +| Command | Observed result | +| --- | --- | +| `cargo test -p cellule-host --example fleet_operations --locked failed_owner_maintenance_settles_and_finalizes_only_from_fresh_process_closure` | 1 passed. | +| `cargo test -p cellule-host --example fleet_operations --locked` | 344 passed; 0 failed; 68.26 seconds. | +| `cargo check --workspace --all-targets --all-features --locked` | Passed. | +| `cargo clippy --workspace --all-targets --all-features --locked -- -D warnings` | Passed. | +| `RUSTDOCFLAGS='-D warnings' cargo doc --workspace --all-features --no-deps --locked` | Passed. | +| Format, boundary, module-layout, Rust-fence, Markdown-link, and SQL/peer-contract checks | Passed; 137 Rust snippets, 1307 local links, 28 protocol assertions and 570 validator links. | + +After adding the lost-Finalize/controller-restart assertion, the focused case, +all 344 minion cases, formatting, and `git diff --check` passed again. The +workspace-wide checks above predate this scenario-only change. + +GitHub reports PR #37 merged and PR #56 clean, mergeable, with all listed +checks passing. The three newly reported conflict files are unchanged in this +checkout and contain no conflict markers. These local changes remain uncommitted +and are not included in PR #56. + +### Highest remaining work + +1. Qualify failed-owner and live role maintenance under cancellation, provider + failure, controller restart and real process loss; add missing role/fault + combinations. +2. Complete receiver-session loss/recovery/adoption and add the canonical + receiver-loss executable scenario. +3. Finish W9 physical process/provider faults, mixed-version and load/soak + qualification, then exercise W10 runbooks and staged rollout/rollback. +4. Re-run the complete qualified source and hosted CI on the eventual PR head. + ## October 4 2026 source reader policy checkpoint The canonical executable remains `crates/cellule-host/minion`, Cargo target @@ -6561,3 +6638,1347 @@ assertions with 570 validator links. The remaining maintenance inspection gap is the full role, accepted-work, facility, Stopped, and withdrawal barrier. This local check is only a Cordon recovery step and cannot move the operation into Closing or Completed. + +## October 4 2026 writer-only end-to-end maintenance checkpoint + +The `maintenance` command now drives a real three-node reference fleet through +the public reconciler and native node actions. It cordons node 0, moves all 12 +SQL Cells to nodes 1 and 2, observes the complete writer-only role inventory, +commits `SettleRoles`, finalizes the drain, withdraws and retires the exact boot, +and verifies each command receipt and final placement. The observer derives +active node sessions from the unresolved Established enrollment rows, so the +post-finalization check can still prove the stopped node is absent while +checking follower references across all physical nodes. + +| Command | Observed result | +| --- | --- | +| `cargo test -p cellule-host --example fleet_operations --locked maintenance_moves_every_cell_then_settles_roles_and_withdraws_the_node -- --nocapture` | 1 passed. | +| `cargo test -p cellule-host --example fleet_operations --locked` | 340 passed; 0 failed; 68.34 seconds. | +| `cargo run -p cellule-host --example fleet_operations --locked -- maintenance` | 12 released, activated and retired; 12 receipt checks; two receivers; final counts `[0, 6, 6]`; maintenance completed and boot withdrawn. | + +The workspace all-target/all-feature `cargo check`, warning-denied Clippy, and +warning-denied API docs passed. Clippy prompted sharing the large settlement +proof through `Arc`; the focused maintenance test passed again after that +change. Format, boundaries, module layout, Rust fences, Markdown links, +SQL/peer contracts, and `git diff --check` also passed. The full 340-test +example run preceded only that proof-storage change. The PR #56 routing rerun +and hosted checks for the current local diff have not completed. + +This establishes one runnable, real-node writer-only maintenance path. Its +empty reader/follower inventory comes from the scenario's closed writer-only +constructor. It does not qualify reader/follower-enabled maintenance, foreign +live- or failed-owner obligations, receiver loss, controller restart during +maintenance, provider/process failures or rollout operations. PR #56's hosted +routing rerun still targets its earlier committed head and was in progress when +this checkpoint was recorded; it does not cover these local changes. + +## October 4 2026 closed-boot role-settlement checkpoint + +Role settlement now distinguishes a live maintenance target from one whose +original boot has already been retired. The opaque `FleetRoleSettlement` +retains the exact failed-boot closure digest when the target process is closed, +and verifies that the retired boot, canonical fence, closure digest, full head, +and registry all match the settlement barrier. Missing or ambiguous live/closed +target evidence refuses settlement. + +`FleetReconciler` sends that proof through a dedicated closed-boot transport +method instead of dispatching SettleRoles to a stopped `CellNode`. The default +transport refuses this case. The reference SQLite adapter accepts the exact +maintenance action and publishes the checked `RolesSettledAt` result through +the shared journal after validating both opaque proofs. The minion now exercises +the full failed-owner path: complete follower-only role settlement, terminal +`Finalize` from the same fresh failed-boot closure, and no transport call to the +retired `CellNode`. The scenario drops the committed `Stopped` reply, waits +beyond actual controller lease expiry, and completes under a different +controller session replaying the same SQLite journal record. + +| Command | Observed result | +| --- | --- | +| `cargo check -p cellule-host --example fleet_operations --locked` | Passed. | +| `cargo test -p cellule-host --example fleet_operations --locked` | 344 passed; 0 failed at the prior full-suite checkpoint. | +| `cargo test -p cellule-host --example fleet_operations --locked failed_owner_maintenance_settles_and_finalizes_only_from_fresh_process_closure` | 1 passed; 0 failed in 4.49s on this checkpoint. | +| `cargo test -p cellule-host --lib role_settlement::tests --locked` | 3 passed; 0 failed. | +| `cargo clippy -p cellule-host --example fleet_operations --locked -- -D warnings` | Passed. | +| `cargo fmt --all --check` | Passed. | + +### Remaining W7 work after closed-boot settlement and finalization + +Exercise this transport path with the full role matrix, recovered follower +closure and joined process evidence. Closed-boot terminal Finalize is now +covered for a follower-only failed owner, including a lost durable reply and +controller restart; extend that evidence to the remaining role and fault +combinations. Then qualify provider and process faults. W4 +receiver-session-loss reconciliation, the W8 receiver-loss executable, and W9–W10 +qualification/rollout remain high-priority work. These local changes are +uncommitted; PR #56 remains clean at its prior remote head and does not include +this checkpoint. + +## October 4 2026 receiver-route acceptance checkpoint + +W4 now has a bounded post-release receiver-route record. Each hop names the +previous and next exact node/session, a nonzero closed-process digest, and a +bootstrapped registry version. The route limits retries to two handoffs, +rejects source/repeated nodes and nonmonotonic registry revisions, and binds the +ordered route into a separate routed-action key and action family. Ordinary +movement action keys and encoded bytes remain unchanged. + +First acceptance checks the current head, endpoint and registry together. The +reference SQLite journal rejects first acceptance through ordinary dispatch; +its routed acceptance entry point requires the opaque `FleetFailedBootClosure` +and compares its exact endpoint, digest, and full head/registry snapshot before +publishing. Route history is bounded and rejects competing branches. The journal +can enumerate accepted endpoints after a controller restart, and reconciliation +chooses the unique deepest retained route. + +| Command | Observed result | +| --- | --- | +| `cargo check -p cellule-runtime --lib --locked` | Passed. | +| `cargo check -p cellule-host --example fleet_operations --locked` | Passed. | +| `cargo test -p cellule-runtime --lib fleet::operations::tests::contracts::receiver_continuation_is_registry_bound_endpoint_bound_and_versioned --locked` | 1 passed; 0 failed. | +| `cargo test -p cellule-host --example fleet_operations --locked routed_receiver_acceptance_requires_typed_closed_boot_proof` | 1 passed; 0 failed. | +| `cargo clippy -p cellule-runtime --lib --locked -- -D warnings` | Passed. | +| `cargo clippy -p cellule-host --example fleet_operations --locked -- -D warnings` | Passed. | +| `cargo fmt --all --check`, `git diff --check`, boundary and documentation checks | Passed; 137 Rust snippets and 1,307 Markdown links/anchors checked. | + +This checkpoint establishes only the durable route contract and acceptance +boundary. The reconciler still selects the original receiver, and the executor +does not yet prepare or recover on a replacement boot. W4 still needs fresh +eligible-node selection, closed-boot route creation before dispatch, new-session +resource admission, takeover of a prior receiver that reached `Serving` or +`Recovering`, and end-to-end unknown/refusal/restart tests. These local changes +remain uncommitted and are not included in PR #56. + +## October 4 2026 routed receiver execution checkpoint + +The reconciler now consumes a fresh roster and placement capture before it +routes `Activate` or `Cancel`. It requires the exact previous receiver's +retired-process closure, confirms that current advertisements map to established +active boots, excludes source and already visited nodes, projects other charged +attempts, and checks signed headroom before choosing an activation target. It +then first-accepts the routed action through the typed closed-boot journal API +before transport dispatch. The reference minion sends the action to the route's +exact node/session, and its journal validates that endpoint's current active +intent. + +A replacement receiver does not inherit the old process's local reservation. +Its routed activation uses ordinary runtime admission after reading canonical +control. It proceeds only from an unowned `Idle` control with the exact released +root and a valid acquisition basis; a same-boot serving result is rechecked +through the actor. A routed retry rereads control under the same accepted route. +Routed cleanup can settle the original receiver's lost credit only after the +closed-process proof is recorded and the routed receiver has no unresolved +local activation. + +The path remains fail-closed when the adapter omits process closures, the +roster is incomplete, the replacement lacks advertised capacity, or a previous +receiver may have reached `Serving` or `Recovering`. The latter still needs the +separate canonical failed-owner recovery proof. The normal movement observer +does not synthesize failed-receiver closures; the dedicated test scenario below +uses an explicit test-only process provider. Controller reconstruction, the +Serving/Recovering refusal boundary, app-specific Cell contracts, and network +transport adapters remain to be qualified. + +| Check | Observed result | +| --- | --- | +| `cargo check -p cellule-host --example fleet_operations --locked` | Passed after routed dispatch, minion endpoint validation, and new-session activation changes. | +| `cargo clippy -p cellule-host --example fleet_operations --locked -- -D warnings` | Passed after the same changes. | +| `cargo fmt --all --check` | Passed. | +| Process-failure routed execution scenario | Added below; focused test passed. | + +These local changes remain uncommitted and are not included in PR #56. The +highest remaining W4 work is controller-reconstruction and refusal/fault +coverage, followed by previous-receiver `Serving`/`Recovering` recovery, +app-contract selection, and qualification of routed remote transports. + +## October 4 2026 routed receiver-loss scenario checkpoint + +The executable minion fixture now covers receiver loss after durable source +release. It starts three real `CellNode`s, reserves the original receiver, +releases the source, joins the original receiver shutdown, withdraws its exact +canonical session, and publishes a typed failed-boot closure. The observer +recaptures that closure against each current head. Reconciliation then accepts +both `Activate` and receiver `Cancel` at the exact replacement node/session, +restores the released root at the next epoch, retires the attempt, and reads +back the acknowledged SQL receipt and committed value. It replays the exact +accepted `Activate` after the reply is lost and confirms the receiver still has +one active Cell. The test then lets the first controller lease expire and +resumes with a new controller session that adopts the retained action and +completes cleanup. It shuts down all nodes and verifies runtime and +disk-reservation resources are released. + +The process provider is a test-only in-process stand-in that treats joined +`CellNode` shutdown as retained evidence. It exercises the provider contract +and routing barrier but is not process-supervision or multi-process +qualification. + +| Command | Observed result | +| --- | --- | +| `cargo test -p cellule-host --example fleet_operations driver_routes_activation_and_cleanup_after_receiver_boot_closure --locked -- --nocapture` | 1 passed; 0 failed. | +| `cargo test -p cellule-host --example fleet_operations successor_tests::driver_ --locked -- --nocapture` | 2 passed; 0 failed. | +| `cargo clippy -p cellule-host --example fleet_operations --all-targets --locked -- -D warnings` | Passed. | +| `cargo fmt --all --check`, `git diff --check`, Rust-fence and Markdown-link checks | Passed; 137 snippets and 1,307 links/anchors checked. | + +W4 still needs an explicit routed stale-generation check and refusal when the +closed receiver's control has advanced to `Serving` or `Recovering`. Existing +public cases cover pre-release capacity refusal, duplicate preparation, +activation and cleanup, exact idle-release generation checks, and shutdown +resource joins. App-specific Cell contract selection and routed network +adapters also remain unqualified. PR #56 and PR #37 are merged; this follow-up +work remains local until separately reviewed. + +## October 4 2026 accepted receiver activation continuation checkpoint + +An accepted activation whose receiver stopped before claiming authority no +longer stalls at failed endpoint inspection. The reconciler can consult fresh +typed process closure, durably accept the bounded replacement route, and invoke +the ordinary receiver executor. A failed inspection or routing attempt keeps +the first original error in the report, bounded to one error per attempt. +Controller deadlines do not authorize this continuation. + +The shared real-node fixture now checks both unaccepted and already accepted +activation from the unchanged Idle root, including lost replacement replies, +duplicate replay, controller lease expiry, receipt readback and joined cleanup. +Two further cases stop the original receiver after real authority transitions +to Recovering or Serving but before actor/result publication. Routed execution +leaves that control unchanged, records Unknown without an acquisition basis, +installs no replacement writer, and retains the full fleet charge. Explicit +receiver cancellation also stays unresolved rather than erasing that claim. + +| Command | Observed result | +| --- | --- | +| `cargo test -p cellule-host --example fleet_operations successor_tests:: --locked -- --nocapture` | 5 passed; 0 failed. | +| `cargo clippy -p cellule-host --example fleet_operations --all-targets --locked -- -D warnings` | Passed for the continuation and shared fixtures. | +| Module ownership, crate boundaries, whitespace, Rust fences and Markdown links | Passed; 137 snippets and 1,307 links/anchors checked. | + +These are in-process scenarios with an explicit test-only stopped-process +provider. The next implementation priorities remain canonical recovery of a +previous receiver's failed ownership claim, routed stale-generation and +provider-fault cases, complete maintenance/controller-restart combinations, +app-specific contracts and production transport/process adapters. W9 measured +provider and mixed-version qualification and W10 rollout evidence remain +required. Narrow tests do not establish completion of W4–W10. + +## October 4 2026 canonical failed-receiver recovery checkpoint + +Routed activation can now recover a failed receiver that already claimed the +Cell. It reads the exact failed control, requires a canonical NodeTakeoverProof +for a process-closed boot in the accepted route, verifies the release prefix, +then calls the existing observed runtime takeover. ReceiverRecoveryBasis is +durably confirmed before the ownership CAS; ReceiverRecoveryEvidence is +confirmed before actor admission. These bounded records use separate kinds 30 +and 31 and retain the original routed acceptance. Source recovery and clean +source release keep their existing semantics and bytes. + +The SQLite reference adapter retains receiver records in a separate table, +compares immutable inputs on replay, forbids mixing them with an Idle +acquisition basis, and checks the exact retained basis/evidence before publishing +Activated. Fresh serving binds the entire input and recovered control to native +acquisition history, verifies any original pinned overlay suffix and the source +release prefix, and rechecks the actor. Missing canonical proof returns Unknown +without another ownership CAS; missing journal support refuses acquisition. + +Two new real-node cases stop the receiver after canonical Serving/Recovering +transitions but before actor/result publication. They recover at the replacement +node, lose and replay the committed reply, wait for controller lease expiry, +reopen an independent SQLite journal client, then retire after exact cleanup. +They preserve the release root, advance through both receiver epochs, read the +original command receipt/value, and join all node resources. The original +refusal tests still prove that an unconfigured recovery provider preserves the +foreign claim and full fleet charge. These are in-process lifecycle cases, not +actual OS crashes or provider-backed qualification. + +Hosted workspace CI on PR #57's prior head caught a regression in the earlier +source RecoveryBasis validation: an arbitrary nonzero action key was accepted. +This change restores the exact source Recover key and preferred endpoint checks. +The existing mutation test passes unchanged. The local warning-denied lint also +passes after replacing a newly deprecated atomic update in the reply-loss +fixture with a bounded-count CAS loop supported by the declared Rust minimum. + +| Command | Observed result | +| --- | --- | +| `cargo test -p cellule-runtime --lib fleet::operations::tests:: --locked` | 75 passed; 0 failed, including the CI regression and both new codec cases. | +| `cargo test -p cellule-host --example fleet_operations successor_tests:: --locked -- --nocapture` | 7 passed; 0 failed. | +| `cargo test -p cellule-host --test node fleet_receivers:: --locked -- --test-threads=2` | 29 passed; 0 failed. | +| `cargo clippy -p cellule-host --example fleet_operations --all-targets --locked -- -D warnings` | Passed after the final prefix/publication changes. | +| Boundaries, module ownership, whitespace, Rust fences and Markdown links | Passed; 137 Rust snippets and 1,348 local links/anchors checked. | + +Highest next priorities are the standalone receiver-loss executable; interrupted +receiver basis/evidence writes and inherited-overlay cases; routed generation +and provider-fault coverage; and the remaining maintenance role/fault matrix. +Production process/transport adapters, W9 measured provider and mixed-version +qualification, and W10 rollout/runbook evidence remain required. PR #57 remains +a draft while hosted CI verifies the updated implementation. + +## October 4 2026 standalone receiver-loss checkpoint + +Source: parent `b92b038`; implementation/CI diff SHA-256 `a3efc37e199cfd54b227cf322358848ac720c1dc19cab843bdeb470ef14b4954`. +Reproduce this fingerprint from the checkpoint commit with +`git diff --binary b92b038 HEAD -- .github/workflows/rust.yml crates/cellule-host/minion ':(exclude)*.md' | shasum -a 256`. + +Minion now exposes `receiver-loss` through the same finite owner used by the +other commands. It reserves one native receiver, proves clean source release, +joins the failed receiver, and publishes typed closure of the exact enrolled +boot. Shared reference adapters supply that closure and lose one committed +routed activation reply. The exported reconciler selects the replacement and +uses ordinary host execution. The command replays the exact retained action, +waits for controller lease expiry, reopens an independent SQLite client, and +settles the unchanged release with controller epoch 2. The former controller +returns Fenced and cannot change the successor journal. + +The command checks all twelve original receipts and SQL values, the moved +source handle's fencing, one destination actor, final placement `[11, 0, 1]`, +and unchanged release root/epoch progression. Both journal clients close on +every exit; the outer owner joins all nodes, retires all boots and checks empty +resource ledgers. The closure is joined in-process evidence, not OS crash or +external-job supervision qualification. Other focused successor cases retain +the separate failed Serving/Recovering receiver takeover coverage. + +The normal-stack focused test exposed excessive stack use when CLI future +construction was nested with complete receiver observation. Scenario and +continuation factories now construct heap-owned futures before polling them. +The executable and ordinary two-worker tests pass without increasing the stack +or changing deadlines. Hosted workspace CI on the preceding `b92b038` head also +aborted during the overload command and failed an outdated observation-count +assertion. That model now asserts the exact planning, activation and cleanup +capture counts at each phase while retaining all movement assertions. CI now +prints assertion details immediately so a later abort cannot hide them. + +| Command | Observed result | +| --- | --- | +| `cargo run -p cellule-host --example fleet_operations --all-features --locked -- receiver-loss` | Exit 0; one release/activation/retirement, twelve receipt checks, epoch 2, one lost activation reply and replay, three joined nodes/retired boots. | +| Same command with `overload` | Exit 0; two releases/activations/retirements, two receipt checks, three joined nodes/retired boots. | +| Same command with `controller-restart` | Exit 0; two lost release replies, epoch 2, two receiver-credit cleanups, two receipt checks, three joined nodes/retired boots. | +| Same command with `maintenance` | Exit 0; twelve releases/activations/retirements and receipt checks, final counts `[0, 6, 6]`, Completed and exact withdrawal, three joined nodes/retired boots. | +| `cargo test -p cellule-host --example fleet_operations receiver_loss_command_preserves_every_receipt_and_joins_all_owners --locked -- --nocapture` | 1 passed; 0 failed. | +| Same test command with `reconciler_tests:: --all-features` | 27 passed; 0 failed. | +| Same test command with `successor_tests:: --all-features` | 7 passed; 0 failed. | +| Same test command with `scenario::tests:: --all-features` | 3 passed; 0 failed. | +| Host example/all-target/all-feature warning-denied Clippy, format, boundaries, module ownership, whitespace, documented Rust and local Markdown links | Passed; 137 snippets and 1,348 links/anchors checked. | + +The preceding hosted follower-only maintenance case failed before the abort; +its two focused local reruns pass without changing its deadline, pass limit or +required Completed/withdrawal/follower assertions. Its diagnostic assertion now +retains the complete reconcile report. The broader hosted result remains +unverified until CI runs this updated head. Focused success does not establish +W4–W10 completion. + +Highest remaining priorities: explain any current-head CI failures; interrupted +receiver basis/evidence persistence, inherited overlays and stale generations; +the maintenance role/controller/fault matrix; production process/transport +adapters and measured W9 fault/mixed-version qualification; and W10 rollout and +runbook evidence. The four required CLI scenarios are available, while W8's +complete application integration and qualification obligations remain open. + +## October 4 2026 interrupted receiver recovery checkpoint + +Source: parent `c61fc55`; implementation/CI diff SHA-256 +`e333a421e2ad6a98bfe10c459df6a6be0072366c85b7bc8284fd80eb7b6ae33a`. +Reproduce from this checkpoint commit with +`git diff --binary c61fc55 HEAD -- .github/workflows/rust.yml crates/cellule-host/src/fleet/movement crates/cellule-host/minion ':(exclude)*.md' | shasum -a 256`. + +Actual transaction-boundary faults exposed a liveness gap: an interrupted +receiver evidence write left a safe native rollback root, but replay could not +publish its result after ordinary acquisition because the recovery evidence was +missing. The host now reconstructs that record only from the exact original +canonical acquisition input/materialization. A retained record keeps its +original capture time. Missing, corrupt or valid-but-substituted history cannot +be replaced by current owner, root equality or epoch counters. + +Replay also resumes an Idle rollback itself through ordinary admitted native +acquisition. It confirms the original recovery evidence and verifies its +materialized prefix and the clean release prefix before another ownership CAS. +The original routed acceptance, basis and movement charge remain unchanged; +there is no second movement or replacement Idle basis. A same-session ordinary +acquisition winner passes the same current-serving/history checks. Fresh +inspection performs neither evidence repair nor acquisition. + +Nine new real-node cases pause SQLite basis/evidence writes before commit or +after commit before reply. They cover both basis boundaries, both evidence +boundaries, cancelled waiters before and after takeover, another interrupted +reconstruction write, automatic Idle resumption and an ordinary acquisition +winner. Missing/corrupt/substituted native records leave authority unchanged and +retain the full charge; an already retained journal recovery record cannot +replace missing native history. The substituted record comes from a different +fully joined canonical acquisition and decodes successfully for the same +Cell/incarnation/epoch. Each successful case reopens an independent journal +client, compares immutable records, settles through the public reconciler, +resolves the original command receipt/value and joins all node resources. +While paused, active-cell credit is distinguished from actual Owned actor +inventory. Original injected I/O errors remain in the source chain. + +| Command | Observed result | +| --- | --- | +| `cargo test -p cellule-host --example fleet_operations successor_tests:: --all-features --locked -- --nocapture --test-threads=2` | 16 passed; 0 failed, including the 9 new fault cases. | +| `cargo test -p cellule-host --test node fleet_receivers:: --all-features --locked -- --test-threads=2` | 29 passed; 0 failed. | +| `cargo test -p cellule-host --example fleet_operations reference_observer_reconciles_follower_only_maintenance_to_completion --all-features --locked -- --nocapture --test-threads=1` | 1 passed; 0 failed with unchanged deadlines and completion assertions. | +| `cargo run -p cellule-host --example fleet_operations --all-features --locked -- receiver-loss` | Exit 0; 12 receipt checks, one release/activation/retirement, controller epoch 2, final placement `[11, 0, 1]`, 3 joined nodes/retired boots and lost activation reply replay. | +| Host example/all-target/all-feature Clippy with `-D warnings`, format, boundaries, module ownership, whitespace, documented Rust and local Markdown links | Passed; 137 Rust snippets and 1,348 local links/anchors checked. | + +Hosted Rust workspace CI on `c61fc55` ran all 352 minion cases: 351 passed and +one follower-only maintenance collection hit its finite observer deadline. +There was no stack abort or observation-count failure. The new CI scheduling +isolates unrelated fixtures with one harness test thread; every case retains +its native worker concurrency, races, deadlines and required evidence. This is +a scheduling change, not a measured diagnosis or proof that the hosted timeout +is resolved. Hosted current-head results must confirm it. Object-proof, +follower-proof, MSRV, contract and quality checks passed on `c61fc55`; those +results do not qualify this newer source. + +Highest next priorities: + +1. Verify current-head CI; qualify routed recovery with inherited overlays, + historical suffix loss/substitution and stale generation/provider faults. +2. Complete maintenance controller loss during evacuation/closing and the + remaining reader/follower/primitive fault combinations, including external + Cron/Blob owners and accepted-work/process closure. +3. Deliver production process/transport integration and the versioned W9 fleet + qualification profile/runner with actual provider, load, resource and + mixed-binary evidence; exercise W10 rollout, rollback and runbooks. + +This remains a focused W4/W5 increment. In-process process-closure providers +and selected test success do not establish completion of W4–W10. + +## October 4 2026 interrupted native takeover continuation checkpoint + +Source: parent `8e11aca`; implementation diff SHA-256 +`18f546f8f5c515376c4d3873610556912bf7b7c5be665e98e21170cd1e536f28`. +Reproduce from this checkpoint commit with +`git diff --binary 8e11aca HEAD -- crates/cellule-runtime/src/cell/actor crates/cellule-host/src/fleet/movement crates/cellule-host/tests/node/fleet_receivers crates/cellule-host/minion ':(exclude)*.md' | shasum -a 256`. + +A failed overlay materialization can retain the target's own Recovering control +and inherited suffix. Asking for another failed-owner takeover cannot resume +that same accepted claim. The native runtime now exposes +`resume_takeover_restored_observed`: validate the original fenced input and +current claim, reconfirm the original recorder, then finish without another +ownership CAS. A current materialized claim must equal the complete canonical +result derived from the original pinned manifest. Changed scope, epoch, owner +or publication is refused. It retains immutable native acquisition input/result +before the recorder's pre-admission confirmation. Initial takeover and resume +share one materialization/activation/rollback owner and the existing resource +admission path. No persisted record format or action key changes. + +Failed-source Recover and routed receiver Activate now use this native +continuation for their own interrupted claims. Routed resumption verifies the +clean release prefix before native effects and obtains the failed predecessor's +proof from the original basis. Fresh inspection remains read-only. + +Two new public host cases use an actual sealed follower suffix. One removes the +manifest, dispatches and replays failed recovery, and confirms the exact owned +Recovering claim/overlay with no actor or acquisition metadata. Restoring the +manifest allows a cancelled replay waiter to join the same accepted native work. +The other pauses a direct native embedding owner after real materialization and +durable recording, interrupts that owner, refuses a changed original control, +then resumes through the host. Both retain the original epoch in the suffix, +activate one writer at the already claimed epoch, and resolve an acknowledged +source command receipt. A historical manifest substitution remains a blocker +while the live actor can still read its acknowledged state. + +The added minion routed case models the post-CAS interruption window with a +real canonical authority transition after durable basis confirmation. Replay +creates actual native acquisition/materialization history and activates without +another epoch; independent journal readback and resource joining use the existing +fixture. This model does not represent an OS crash or a complete routed overlay +and controller succession campaign. + +| Command | Observed result | +| --- | --- | +| `cargo test -p cellule-host --test node fleet_receivers:: --all-features --locked -- --test-threads=2` | 31 passed; 0 failed, including both new public overlay cases. | +| `cargo test -p cellule-host --example fleet_operations successor_tests:: --all-features --locked -- --nocapture --test-threads=2` | 17 passed; 0 failed, including the new exact routed-claim replay. | +| `cargo test -p cellule-runtime --test runtime lifecycle::ownership::recovery:: --all-features --locked -- --test-threads=2` | 8 passed; 0 failed; 1 provider case remains explicitly ignored without its documented isolated RustFS environment. | +| `cargo check -p cellule-host --all-targets --all-features --locked` | Passed. | +| Runtime/host all-target/all-feature Clippy with `-D warnings`, format, boundaries, module ownership, whitespace, Rust fences and Markdown links | Passed; 137 snippets and 1,348 local links/anchors checked. | + +The previous checkpoint's hosted Rust run `37242017429` completed with failure: +360 of 361 minion cases passed, and the follower-only maintenance observer again +timed out in `fleet-follower-evacuation-deadline`. Serial harness scheduling did +not resolve that failure. MSRV, workspace feature/target checks, API +documentation, workspace unit/integration tests and the balanced three-process +smoke steps passed on `8e11aca`. Those results do not qualify this newer native +implementation. Reproduce and diagnose the observer timeout without changing +its deadlines or required completion evidence; this checkpoint also requires +its own broad CI after publication. + +Highest next priorities: + +1. Verify hosted CI and combine inherited overlays with routed boot/controller + succession, historical suffix/lineage faults and provider outages. Complete + missing-evidence continuation of safely rolled-back failed-source recovery. +2. Complete maintenance controller/role/primitive fault combinations and actual + accepted external-job/process closure, including Cron/Blob obligations. +3. Deliver the W9 versioned measured fleet profile/runner and provider/process, + mixed-binary and load/resource evidence; exercise W10 staged operations, + rollback and runbooks. These remain required for full plan completion. + + +## 2026-10-05 — Failed-source evidence repair and bound Idle suffix proof + +Parent: `15b4f2923a8ea58cab3fcf0cf09cab5e546d3637`. The code-only diff from +that parent, excluding Markdown, has SHA-256 +`fafc62ff7fdb11f4b27c9dcd0e46894fa443ecf856af90590a8f87913fde711b`. +This checkpoint advances W4 recovery; it does not complete W4–W10. + +Accepted failed-source Recover replay now repairs a missing evidence write only +from the exact original retained basis and immutable native acquisition +input/materialization. A committed evidence reply lost in transport retains the +original record and time. A safe Idle rollback can resume ordinary admitted +acquisition after prefix/origin verification; an ordinary acquisition winner +retains its actor and must satisfy the same original-history checks. Fresh +inspection writes no metadata and starts no acquisition. Failed evidence writes +keep the full attempt charged and preserve their original I/O source error. + +Actual sealed-suffix replay exposed a missing pre-acquisition contract: the +Serving suffix verifier correctly rejects Idle. The new explicit +`verify_recovered_idle_prefix` requires a complete unowned Idle control with +cleared overlay, shares the bounded canonical-history/origin walk and runtime +I/O/memory admission, and rechecks the complete control before and after I/O. +It supplies no ownership, actor, retention, role or settlement rights. The +existing Serving verifier retains its strict state requirement. + +Ten new public host cases cover lost evidence replies, failed and repeated +writes, cancelled repair waiters, ordinary acquisition winners, missing/corrupt +original history in both Idle and Serving states, valid substituted history, +missing materialized origin, and actual sealed-suffix manifest faults. Restoring +exact original bytes permits replay without another movement permit or rewritten +basis. Completion checks current writer, original canonical evidence, repeated +action identity, acknowledged command resolution and joined resources. + +| Command | Observed result | +| --- | --- | +| `cargo test -p cellule-host --test node fleet_receivers:: --all-features --locked -- --test-threads=2` | 41 passed; 0 failed, including all ten new cases. | +| `cargo test -p cellule-runtime --lib control::authority::acquisition::tests:: --all-features --locked -- --test-threads=2` | 11 passed; 0 failed. Idle refusal by the Serving API, zero bound, missing acquisition and stale complete control covered. | +| `cargo test -p cellule-host --example fleet_operations successor_tests:: --all-features --locked -- --nocapture --test-threads=2` | 17 passed; 0 failed after the temporary deadline probes were compiled. | +| Exact minion `reference_observer_reconciles_follower_only_maintenance_to_completion` | 1 passed; 0 failed, 2.55 seconds locally. This does not establish a hosted CI fix. | +| Runtime/host all-target/all-feature Clippy with `-D warnings`; API docs with `RUSTDOCFLAGS='-D warnings'` | Passed. | +| Format, whitespace, boundaries, module ownership, Rust fences, Markdown links and runtime SQL/peer validator | Passed: 137 snippets, 1,348 Markdown links/anchors, 28 protocol/schema assertions and 571 validator links. | + +Authoritative hosted results for parent `15b4f29`: Rust run `37244070031` +finished with workspace failure, 343 of 362 minion cases passing and 19 failing. +Most failures were observation deadlines; the previously failing follower-only +observer failed again. Workspace all-feature/target checks, workspace tests, +API docs, MSRV and the balanced three-process smoke passed before that step. +Follower and object proof checks also passed on the parent. These results do +not qualify the new code. Contract run `37244070048` failed the unchanged +paused-clock LTX compaction test: 500 ms observed versus the required 400 ms. +Neither failure is resolved by this checkpoint. + +Temporary test-only `[DEBUG-fleet-57]` probes report failed/cancelled collection +stages, native snapshot waits, policy verification and original elapsed budgets +in the next hosted run. Normal successful observations print no probe output. +No deadlines, qualification profiles or assertions were changed. Remove the +probes after the hosted cause is confirmed and fixed. The isolated ARM64 Linux +observer and LTX repetitions passed earlier; AMD64 emulation failed in native +compiler/linker setup before tests and supplies no AMD64 test evidence. + +Highest next priorities: + +1. Diagnose the hosted observer and LTX failures with original deadlines and + evidence intact; verify this checkpoint in CI and keep PR #57 mergeable. +2. Complete the routed inherited-overlay/boot/controller succession campaign, + historical suffix/lineage and provider faults, and the remaining maintenance + role/primitive fault combinations and accepted external-job/process closure. +3. Deliver the W9 versioned measured fleet profile/runner, actual provider/process + and mixed-binary/load/resource evidence, then exercise W10 staged operations, + rollout/rollback and runbooks. These remain required for full plan completion. + + +## 2026-10-05 — Routed inherited suffix and controller fault campaign + +Parent: `a509e2c1d087f39086c93802837060f0453ea81e`. The code-only diff, excluding +Markdown, has SHA-256 +`5127a30bbfbe190133f197722bbde06029f85b14e51e69604ca0b1cc7c4cc7e5`. +This checkpoint adds three selected W4/W5 fault models; full W4–W10 delivery +remains open. Production authority, formats, action keys and profiles are unchanged. + +The routed fixture now acquires the released Cell through an actual earlier +native owner, captures a SQLite tail and fsyncs it to every selected follower. +Each original member request is durably Pending before canonical enrollment. +Activation follows all first-append acknowledgments. Native recovery seals and +pins the complete tail; complete member retirement and typed fleet publication +retain every original request before the later receiver's process closure. + +Removing that original manifest makes a real preferred-receiver takeover commit +its claim, then fail materialization. The current Recovering control inherits +the exact earlier overlay; no successful acquisition record is fabricated. +Routed recovery retains that control as its original basis while preserving the +manifest's earlier Cell epoch. The additional runtime and all fleet runtimes +join their jobs, actors, descriptors and admission credits before private paths +are dropped. Process evidence remains the documented in-process reference. + +The new cases prove: + +- Evidence-write failure rolls back safely to Idle; replay confirms the original + canonical materialization before ordinary admitted reacquisition. The explicit + native suffix proof identifies the earlier manifest epoch and exact acquisition + that materialized it, despite a missing interrupted intermediate record. +- Missing/corrupt historical manifests keep the attempt charged with original + storage/manifest errors and no writer or additional ownership claim. Restoring + original bytes resumes the same evidence. A replacement controller waits for + real lease expiry, opens an independent SQLite client, adopts the original route + and checked result, joins receiver credit and retires the attempt. The old + controller is fenced and leaves the journal unchanged. +- A manifest outage after a routed ownership CAS preserves that exact Recovering + claim. Restored bytes permit native resumption without another epoch. A replay + waiter cancelled at the actual pre-admission evidence write cannot cancel + materialization or actor admission; the owned action retains exact native + evidence and the replacement controller joins its completion. + +Every completion resolves the original acknowledged command and reads the +actual materialized suffix value. Retention and complete process/provider +qualification remain separate. + +Sharing the larger fixture initially caused a reproducible stack overflow in +an existing receiver evidence test, both alone and in the combined suite. +Boxing the shared finite constructor and completion futures fixes that seam +without increasing stack limits or changing test profiles. The same failing +case and the full successor suite then pass. The existing member transport +moved into one focused module and is reused by both fixture families. + +| Command | Observed result | +| --- | --- | +| `cargo test -p cellule-host --example fleet_operations successor_tests:: --all-features --locked -- --nocapture --test-threads=2` | Final code: 20 passed; 0 failed, 54.13 seconds. An intermediate repeated run hit controller Fenced in an existing history fixture while its journal write was paused; that timing debt remains recorded rather than weakening its expectation. | +| `cargo test -p cellule-host --example fleet_operations successor_tests:: --all-features --locked -- --nocapture --test-threads=1` | Final code: 20 passed; 0 failed, 131.63 seconds using the serial CI harness. | +| `cargo test -p cellule-host --example fleet_operations scenario::recovered_followers::tests:: --all-features --locked -- --nocapture --test-threads=2` | 7 passed; 0 failed after the shared transport move. | +| Host all-target/all-feature Clippy with `-D warnings`; format, whitespace, architecture and module ownership | Passed. | + +Hosted parent Rust run `37246672617` is terminal Failure: 343 of 362 minion +cases passed, 19 failed. Its test-only probes localize repeated routed failures +to failed-boot closure with about 1.05 seconds of remaining budget versus about +1.10 seconds elapsed. The follower-only case exhausted its roughly 2.50-second +share during native collection and roster confirmation. No native snapshot hang +is shown by those traces. The exact cause and correction for the consumed +controller partitions remain open; do not extend deadlines or suppress failures. +Contract run `37246672563`, follower/object proof, MSRV and cookbook quality +checks passed on that parent. The intermittent earlier LTX timer failure is +not established as fixed by this one green run. New-head CI remains required. + +Highest next priorities: + +1. Reproduce and fix the hosted collection/controller-budget failures with + original profiles and proof checks intact; remove temporary probes only after + confirming the cause and regression. Diagnose the observed fixture lease-expiry + race and retain broad CI evidence for each published head. +2. Extend routed provider/backend/lineage faults and successive boot failures, + complete remaining maintenance role/primitive combinations, and qualify + actual accepted external-work/process providers, including Cron/Blob closure. +3. Deliver W9's committed measured fleet profile/runner and actual provider/process, + mixed-version/load/resource campaign; exercise W10 rollout, rollback and + runbooks. Keep PR #57 synchronized with main and mergeable throughout. + + +## 2026-10-05 — Native deadline reproduction and canonical signature boundaries + +Parent: `ed16bf9cade53c831242f34f3524c4084854c978`. The code-only diff, +excluding Markdown, has SHA-256 +`503e4e80ad3ded04d265a36ee71651b51a27a748699cb636a0c475052e1dff2e`. +This checkpoint fixes redundant authentication at specific directory boundaries; +it does not claim green fleet CI or completion of W4–W10. + +An isolated native Ubuntu 24.04 workflow reproduced all three selected hosted +failures with their original debug harness, deadlines, controller profile and +assertions. Stage probes measured routed collection at 922 ms: foreign reference +collection/rechecks consumed 631 ms, while journal confirmation used less than +one millisecond. Follower collection took 2.69 seconds, including 1.61 seconds +in those reference scans. This rules out journal confirmation as the principal +cost in these captures; it does not establish a general provider latency bound. + +Canonical decoding had already authenticated identity and understood placement +signatures before exact loads, live/advertised scans and follower inventories +verified the same immutable bytes again. Those consumers now retain every fresh +read, canonical-byte/path check and their existing scope/time policy, while +using the decoder's original authentication. Create and heartbeat refresh still +validate signatures before any CAS, then use the one canonical serializer on +the unchanged candidate. Unverified producers retain the checked encode path. +There is no signature cache across reads, changed bytes or directory scans; +formats, signature inputs, error sources, action keys and budgets are unchanged. + +Seven regression cases cover all three advertisement forms, verification counts, +canonical emission/readback, changed valid bytes, forged identity/placement +signatures, expiry, future issue time, foreign scope and misplaced paths. +Invalid producers retain their original signature errors and cannot create or +change a canonical record. Three read/scan tests and the producer test failed +with two verification passes before their respective fixes. + +Native diagnostics, each selecting one exact case per command: + +| Snapshot / run / job | Observed result | +| --- | --- | +| `573214e`; [run 37249579053 / job 111574292502](https://github.com/crabbuild/cellule/actions/runs/37249579053/job/111574292502) | Original routed claim, interrupted evidence write and follower-only maintenance each failed. Stage timings above come from this run. | +| `80691d4`; [run 37250078568 / job 111575760319](https://github.com/crabbuild/cellule/actions/runs/37250078568/job/111575760319) | Both routed cases passed after the read-boundary change. Follower collection improved to 1.81 seconds but policy verification still exhausted its unchanged 2.50-second share. | +| `8a73df3`; [run 37250615027 / job 111577285575](https://github.com/crabbuild/cellule/actions/runs/37250615027/job/111577285575) | Both routed cases passed after producer emission also stopped repeating verification: 6.14 and 4.51 seconds for the complete cases. Follower-only maintenance reached a different failure: native Host snapshot reported the original `fleet-action-journal` Conflict after 329 ms, with 2.17 seconds of observation budget remaining. This is a changed-barrier refusal, not a proven complete CI fix. | + +The diagnostic snapshots add only a temporary workflow and stage probes to the +recorded code. They are not protected qualification receipts and do not replace +current-head broad CI. Raw logs are retained in the linked workflow runs. +Temporary probes remain because the follower conflict and broad failures still +need diagnosis; no deadline, pass partition, assertion or qualification profile +was relaxed. + +| Command | Observed result | +| --- | --- | +| `cargo test -p cellule-runtime --lib node:: --all-features --locked -- --test-threads=2` | 119 passed; 0 failed, including all seven new cases. | +| Exact minion follower-only maintenance selector | 1 passed locally, 2.30 seconds; native changed-barrier failure above remains open. | +| Serial minion successor suite after the read-boundary change | 20 passed; 0 failed, 127.62 seconds. | +| Serial minion successor suite with the final producer change and probes | 20 passed; 0 failed, 97.32 seconds. | +| Runtime/host all-target/all-feature Clippy and warning-denied API docs | Passed for the final producer change and test additions. | +| Format, whitespace, boundaries, module ownership, document and SQL/peer gates | Passed: 137 snippets, 1,348 local Markdown links, 28 protocol/schema assertions and 571 validator links. | + +Highest next priorities: + +1. Diagnose the native follower registry conflict and preserve its original + changed-barrier refusal while making safe reconciliation resumable. Verify + the published head's complete CI, the observed controller-expiry timing debt + and the intermittent LTX timer case; remove probes after confirmed regression. +2. Extend routed provider/backend/lineage and successive-boot faults, complete + remaining maintenance role/primitive combinations, and qualify actual external + work/process providers, including Cron/Blob closure. +3. Deliver W9's measured profile/runner, provider/process and mixed-version/load + evidence; exercise W10 staged rollout, rollback and runbooks. Keep PR #57 + synchronized with main and mergeable. + + +## 2026-10-05 — Native role settlement after controller reconstruction + +Parent: `f42df6a6f0e7c0bb2e495e06f4c626d81d1b8aa0`. The Rust-only diff +has SHA-256 `da4b6e40b2896bf2dc75cdbf7935304bedeed39d1cd2427585e95a96851a6199`. + +The new real-node regression drops the reply after the donor commits +`RolesSettledAt`, then reconstructs the public driver over an independent SQLite +client. It failed before the fix: the executor validated the new proof against +the original accepted envelope's earlier head, then failed publication with +`fleet-action-journal` Conflict and retained Evacuating. The executor now checks +stable replay identity against the original acceptance, validates the refreshed +opaque proof against the current request, and publishes through the existing +atomic current-head/registry check. Movement still uses its original accepted +inputs. No persisted format, action key, original deadline or qualification +profile changed. + +The regression proves that acceptance stays byte-identical while the receipt +covers the renewed head, then reaches Closing/Completed and checks Stopped, +permanent withdrawal and unchanged follower policy history. Rotation drained +the original actor, so final service evidence restores the authority-pinned +canonical root and compares both the value and original `sys_requests` result +with the acknowledged 29. Joined cleanup checks runtime/disk/enrollment ledgers. +This is same-claimant reconstruction, not replacement after lease expiry or +external process qualification. The new case has its own ten-second pass bound; +existing five-second follower qualification and its partition remain unchanged. + +Earlier full-head evidence now has a terminal result: + +| Source / job | Observed result | +| --- | --- | +| `f42df6a`; [Rust workspace job 111578870302](https://github.com/crabbuild/cellule/actions/runs/37251146273/job/111578870302) | All 365 minion cases passed, 0 failed, 731.52 seconds; workspace tests, process smoke, Axum provider checks and local/replica LTX suites passed. The job later failed at warning-denied Clippy on one deprecated `fetch_update` in the public host fault fixture. | +| `f42df6a`; [follower-proof job 111578923743](https://github.com/crabbuild/cellule/actions/runs/37251146389/job/111578923743) | Success. Object proof, MSRV, qualification contract, website and cookbook quality/scenario jobs also completed successfully. This does not qualify the new source checkpoint. | +| `ffc7688`; [native diagnostic job 111580309609](https://github.com/crabbuild/cellule/actions/runs/37251638607/job/111580309609) | Both routed cases passed. Follower-only maintenance exceeded its original 2.50-second share: 1.675 seconds in collection and 0.824 seconds in policy checks. No journal-conflict probe fired in this reproduction. | +| `0bf64cc`; [native repeated diagnostic job 111581716096](https://github.com/crabbuild/cellule/actions/runs/37252117116/job/111581716096) | Both routed cases and three independent executions of the unchanged follower-only case passed. Each command selected one exact test. This run adds collector logging and repetitions, not a deadline/profile change. | +| `35e1b0c`; [native settlement diagnostic job 111582701148](https://github.com/crabbuild/cellule/actions/runs/37252452852/job/111582701148) | The new reconstructed-controller settlement/restore case passed, 1 selected test, 14.48 seconds. Both routed cases passed. All three unchanged follower-only repetitions failed at their original observation deadline; the job is Failure. This snapshot contains the settlement fix/test, but predates the separate public-fixture lint rename. | + +The CI lint failure used the newer stable compiler; local Rust is 1.97. The last +`fetch_update` was changed to the already-supported `try_update` with identical +ordering and closure. No lint suppression was added. + +Local verification of the final source: + +| Check | Result | +| --- | --- | +| New native restart/restore regression | 1 passed; before-fix failure was observed at the real node/journal seam. | +| Opaque role-settlement binding/refusal unit cases | 3 passed. | +| Public host maintenance actions | 6 passed. | +| Public host action acceptance, cancellation and publication cases | 7 passed. | +| Exact journal Closing reconstruction and enrollment-during-capture refusal | 1 passed each; changed-registry Conflict is still required before allocation. | +| Host all-target/all-feature warning-denied Clippy | Passed on Rust 1.97 after the rename. Latest stable verification belongs to current-head CI. | + +Temporary test-only capture probes remain because the observed native +changed-barrier refusal and timing variation require repeated campaign evidence. +The strict snapshot refusal is unchanged. A green run alone does not establish +that all timing or controller-expiry failure patterns are eliminated. + +Highest next priorities: + +1. Verify this checkpoint's complete CI and remove probes after confirmed native + regressions. Qualify actual new-claimant maintenance restart, unknown result + publication and changed-head/registry races without weakening refusal gates. +2. Complete the role/primitive fault matrix, successive boot/provider/lineage + cases and actual external process, Cron and Blob ownership barriers. +3. Deliver W9 measured load/profile/provider and mixed-version evidence, then + W10 staged rollout/rollback and exercised operator runbooks. W4–W10 remain + incomplete; keep PR #57 synchronized and mergeable. + + +## October 4 2026 shared inventory and controller-expiry checkpoint + +Parent: `eba3b6c34c363ab40a707425a510ccb4ff6ed6b9`. The Rust-only diff +has SHA-256 `92b66d2924d9c269dd488ddc713af1296981884e51d1d7cbdde9d654752dc99d`. + +Delivered: + +- `NodeDirectory::follower_logs_pages` supplies independent follower windows + from one fresh authenticated directory traversal. Requests are unique and + bounded; combined page rows stay at most 128, plus one lookahead per member. + The scalar API uses the same implementation. Existing topology domains, + 48-byte cursors, exact rows and expired/fenced obligations are unchanged. +- `FleetFollowerReferences::collect_all` and `recheck_all` preserve the complete + per-member inventories and original full roster. A failed recheck invalidates + every previous confirmation without changing any original row or interval. + Coverage remains invalid until every member is freshly confirmed. Minion uses + these APIs at its existing collection/recheck boundaries. +- Immutable enrollment fingerprints serialize their already-authenticated + original private inputs through the canonical serializer. The regression was + red with four repeated signature passes, then green with zero and identical + original digest bytes for attempt, enrollment and refusal. Every fresh + canonical directory read still independently authenticates each record. +- Maintenance restart now also waits for actual controller lease expiry while + the original boot owners publish real signed heartbeats. A new claimant gains + epoch 2, refreshes the same accepted settlement at its current head, and reaches + Completed/Stopped/withdrawn. The old claimant is fenced with an unchanged + journal; canonical-root value and original request-result readback remain 29. + +Native evidence: + +| Source / job | Result | +| --- | --- | +| Published parent `eba3b6c`; [workspace job 111583565455](https://github.com/crabbuild/cellule/actions/runs/37252756090/job/111583565455) | Failure: 364 minion cases passed, 2 observer cases failed, 1002.01 seconds. The failures exhausted the unchanged controller/observation deadlines. Follower proof, object proof, MSRV, contract, website and cookbook checks passed. | +| Shared traversal snapshot `1c6f6c2`; [diagnostic job 111585942801](https://github.com/crabbuild/cellule/actions/runs/37253571241/job/111585942801) | Failure: all three original follower-only repetitions still exhausted their original share; policy/graph composition also refused changed accepted work during native capture. The strict refusal remains unchanged. | +| Shared traversal plus immutable fingerprint snapshot `5d162c4`; [diagnostic job 111588516964](https://github.com/crabbuild/cellule/actions/runs/37254474648/job/111588516964) | Success: both routed cases, three original follower-only repetitions, same-claimant restart, policy/graph composition, 2 reference batch cases, 8 cursor cases, 8 canonical authentication cases, actual-expiry restart and 13 enrollment cases. Each command selected its expected cases. The actual-expiry case took 36.28 seconds. | + +In the final native reproduction, original follower collection took about +0.66–0.67 seconds instead of 1.44 seconds in the preceding shared-scan-only +snapshot. Repeated producer-page signature work was the remaining cost. Original +five-second follower deadlines, their observation partition and qualification +profiles were preserved. These are diagnostic observations, not a measured W9 +production capacity profile. The successful diagnostic snapshot includes tagged +probes; they have now been removed from the source checkpoint. + +Local verification: 124 runtime node cases, 3 original observer cases, 2 shared +reference cases and both controller restart cases passed. The 3 original observer +cases also passed after probe removal. Host/runtime all-target/all-feature +warning-denied Clippy and warning-denied API docs passed. Formatting, boundaries, +module layout, 137 Rust snippets, 1348 Markdown links and 28 SQL/peer assertions +passed; the runtime validator checked 571 links. Full cleaned-source workspace +and process evidence belongs to the new PR head's CI. + +PR #37 is merged. The continuation PR #57 includes the latest fetched main and +was mergeable when this checkpoint was prepared; none of the three reported +conflict files has conflict markers. + +Highest next priorities: + +1. Complete the cleaned checkpoint's full PR CI and preserve strict changed-head, + changed-registry and accepted-work capture refusals. +2. Qualify uncommitted SettleRoles publication across controller replacement, + successive boot/lineage cases and the remaining W4–W7 role/primitive faults, + including external Cron/Blob owners. +3. Complete W9 recorded process/provider, load/soak and mixed-version campaigns, + then exercise W10 operator rollout, rollback and recovery runbooks. W4–W10 + remain incomplete; these focused cases do not establish full fleet readiness. + + +## October 5 2026 role-result publication and cancelled-owner checkpoint + +Parent: `ab5eb0bfe4a8dedcc62a7cb44528e250210c9512`. The Rust-only diff +has SHA-256 `70aa7b017943121bed73faa6f9e1c4316a99ad55611dec6763002a372d3fdfd0`. + +A real-node regression reproduced an unpublished native `SettleRoles` proof +that prevented every later complete observation. The joined executor retried +the original proof after controller renewal; its old-head publication correctly +returned Conflict, but its retained local job prevented fresh capture forever. +The executor now returns that original publication error and retires only this +joined, read-only local proof after its failed retry. Original durable acceptance +and historical results remain unchanged. A later pass must collect complete fresh +native and foreign-role evidence and commit through the existing atomic +head/registry check before Closing. Physical-effect receipts remain retained +across publication failures and are never reexecuted. + +Four new restart cases fault the actual journal write before commit or its reply +after commit, under both same-claimant renewal and replacement after actual lease +expiry. Two additional cases cancel the public controller while the original +native publication owner is paused at those boundaries. Exact duplicates join +that owner; competing fresh proofs are refused while it runs. After its original +completion and failed stale-proof retry, fresh evidence advances Closing and +Completed. All cases preserve original acceptance, errors and historical results, +check canonical-root value and original request-result readback of 29, and join +all runtime, disk and enrollment ledgers. No existing qualification profile, +five-second deadline or expected evidence changed. These new restart/cancellation +cases use their own ten-second pass bound. + +| Source / job | Observed result | +| --- | --- | +| Published parent `ab5eb0b`; [workspace job 111590038595](https://github.com/crabbuild/cellule/actions/runs/37254990107/job/111590038595) | Success: all 369 minion cases, 0 failures, 864.60 seconds; workspace, provider/process smoke, local/replica LTX, warning-denied lint and contract gates passed. Environment-specific ignored cloud suites are not qualified by this result. | +| Snapshot `4b847c4`; [native publication job 111594331530](https://github.com/crabbuild/cellule/actions/runs/37256405859/job/111594331530) | Success on Rust 1.99: 6 restart cases (128.49 seconds), 2 cancelled owners (17.81 seconds), 3 unchanged observer contracts, 7 physical-action retention cases, 3 opaque-proof refusals, and host all-target/all-feature warning-denied Clippy. | + +The native snapshot contains the exact staged Rust source and no probes. Its +21 selected tests establish the focused regression scope; the new published +head still requires its full workspace and provider CI. Local verification also +passed the six restart and two cancellation cases, seven physical-action cases, +six public maintenance cases, three opaque-proof cases, host Clippy/API docs, +format, boundaries/layout and document/SQL-peer gates. + +Highest next priorities: + +1. Verify this source checkpoint's full CI and preserve current main ancestry + and PR mergeability. Extend the maintenance executable beyond writer-only + coverage using the existing public reader/follower orchestration. +2. Complete successive boot/lineage and remaining W4–W7 role/primitive faults, + including actual external Cron/Blob and failed-process ownership barriers. +3. Deliver W9 recorded provider/process, load/soak and mixed-version campaigns, + then exercise W10 rollout, rollback and recovery runbooks. W4–W10 remain + incomplete; focused publication regressions do not establish fleet readiness. + + +## October 5 2026 executable reader maintenance checkpoint + +Parent: `9257b18c825118db00518fd8b494a7bd9a75792b`. The Rust-only diff +has SHA-256 `5045af8d1faba496434d2414631c62eb294e6f492e9efd15976d9c0ad9802aa5`. + +Delivered `maintenance-reader` through the canonical minion executable. Three +managed boots retain Pending/Established before readiness, one real writer +acknowledges a mutation from 17 to 29, and the public driver cordons the reader's +physical node. A complete observation without replacement policy remains +Evacuating, preserving the original Established reader and its usable value. +The command opens a selected native replacement, performs canonical evacuation, +publishes immutable policy evidence and lets fresh complete observation +authorize SettleRoles and Finalize. It checks Completed, Stopped, exact boot +withdrawal, retired/fenced original reader, receipt-bound replacement value 29, +the original writer's stored mutation result and unchanged writer ownership. +The shared exit path joins three nodes, both reader enrollments and boot rows, +then checks every existing runtime resource ledger. + +The signed native peer verifier/dispatcher adapter moved from the private reader +test tree into a shared production minion module. Existing cancellation/probe +controls remain test-only; the executable and tests use the same routing. No +framework transport, second scheduler or application authentication policy was +added to the framework. + +The first executable regression failed because replacement activation still saw +the original donor in the reader directory's bounded membership cache. It now +waits within its deadline for canonical selection to observe the signed cordon +and choose the spare, preserving the original reader during that interval. +The reference heartbeat also waits for a newer actual classifier sample that +matches the current local admission mode before signing. No mode, sample time +or sequence is fabricated, and the canonical reader cache remains unchanged. +All temporary diagnosis probes were removed. + +| Source / check | Observed result | +| --- | --- | +| Exact Rust snapshot `0409277`; [native reader job 111598959724](https://github.com/crabbuild/cellule/actions/runs/37257960935/job/111598959724) | Success on Rust 1.99: executable regression (1 test, 8.26 seconds), production `maintenance-reader` command, 10 unchanged reader fault cases, 2 existing public reader observer cases, 3 unchanged follower observer cases, and host all-target/all-feature warning-denied Clippy. The snapshot matches all nine changed Rust paths, including the old adapter deletion. | +| Local executable and regression | Success: zero writer movement/permits, two receipt checks, one receiving reader node, three joined nodes and boot retirements, final writer counts `[1, 0, 0]`, maintenance Completed and exact withdrawal. | +| Local existing reader faults | 10 passed, 0 failed, 32.70 seconds. Original profiles, deadlines and assertions remain unchanged. | +| Local quality gates | Host all-target/all-feature warning-denied Clippy and API docs, format, boundaries/layout, 137 Rust snippets, 1348 Markdown links, and 28 SQL/peer assertions with 571 validator links passed. | +| Parent `9257b18`; [Compose smoke job 111597096334](https://github.com/crabbuild/cellule/actions/runs/37257164322/job/111597096334) | Failure in the unchanged sustained mixed-reader load at three constrained nodes: `ReplicaUnavailable` at `process_scaling.rs:548`. Earlier smoke/fault cases and initial reader readiness passed. Captured containers show no OOM kill or unexpected node exit. Cause remains unproven; do not treat focused reader qualification as a full CI repair. | + +Raw Compose driver, node, provider, arrival/receipt and container evidence is in +the failed run's `cell-reference-compose-37257164322-1` artifact. The parent Rust +workspace job was still live at this checkpoint. Complete new-head workspace and +provider CI remain required. + +Highest next priorities: + +1. Reproduce and fix the constrained mixed-reader `ReplicaUnavailable` failure + using the actual process/provider campaign and original availability gates. + Finish full current-head CI and keep PR #57 synchronized and mergeable. +2. Deliver the live-follower maintenance CLI using the existing native supervisor, + then complete successive boot/lineage and the remaining W4–W7 role/primitive + faults, including external Cron/Blob and failed-process ownership barriers. +3. Complete W9 measured profiles, provider/process, load/soak and mixed-version + campaigns, then exercise W10 rollout, rollback and recovery runbooks. W4–W10 + remain incomplete; this finite reader executable is one W8 deliverable. + + +## October 5 2026 executable live-follower maintenance checkpoint + +Parent: `f9a476e0be2bab8d076ac07745596f5ac0cf6e86`. The Rust-only diff +has SHA-256 `b9b239290215d2a532e346cf69fac2f22bbad31ff89f2a9b0bb534490d72d333`. + +Delivered `maintenance-follower` through the canonical minion executable. +Four managed boots retain enrollment before readiness: a live writer, two +original follower stores and one spare. The donor has zero local writers but +retains a foreign lane. A complete observation without replacement policy stays +Evacuating and preserves its original Established row. The command enables the +spare through a real signed heartbeat, drains the original writer to cover its +acknowledged tail, and requests rotation through the existing durability +supervisor. It checks nonzero original coverage, both original Retired members, +exact replacement epoch 2 and two eligible replacement members. Immutable policy +publication and fresh full native/foreign observation authorize SettleRoles and +Finalize; completion requires Stopped and permanent exact-boot withdrawal. + +Canonical Idle acquisition resumes the writer on its original physical node, +resolves the original outcome and value 29, acknowledges a new command, and +resolves that new outcome. Canonical-root restoration must cover both command +sequences and preserve the original request digest, expiry, sequence and exact +stored outcome. Final writer counts include the spare: `[1, 0, 0, 0]`. The shared +exit path joins all four nodes, retires all four follower-member rows across the +two ensembles and all four boot rows, closes the journal, and checks every +existing runtime resource ledger. Fleet writer release/activation/retirement and +movement permits stay zero; five service/receipt checks and two replacement +members are separate evidence. + +An initial scenario incorrectly expected the drained source to remain a live +writer. The corrected command performs canonical acquisition before service +checks. A subsequent added check incorrectly required a new command to activate +the follower log: native object proof can win before that activation CAS. The +command now checks the real installed ensemble independently through canonical +native evacuation and requires a published root covering both acknowledgements. +No runtime durability gate, original deadline or qualification profile changed. +The finite embedding adapter retains exactly two original preparations and +checked close barriers, authorizes live append/retire through the canonical +directory and supplies no failed-owner seal/tail capability. It establishes no +OS-crash, external-provider restart or mixed-role/primitive qualification. + +Current evidence: + +| Source / check | Observed result | +| --- | --- | +| Published parent `f9a476e`; [Rust workspace job 111600595261](https://github.com/crabbuild/cellule/actions/runs/37258527941/job/111600595261) | Success: 376 minion cases, 0 failures, 975.39 seconds; workspace tests, three-process smoke, provider checks, local/replica LTX and warning-denied lints passed. Environment-specific ignored cloud suites remain unqualified. Follower/object proof, MSRV, contract, website and cookbook quality also passed. | +| Local final source | New follower regression and production command passed; 12 original native observation cases, original reader CLI and count-balance case passed. Host all-target/all-feature warning-denied Clippy/API docs, format, boundaries/layout, 137 Rust snippets, 1348 Markdown links and 28 SQL/peer assertions with 571 validator links passed. | +| Exact Rust snapshot `fb5badf`; [native follower job 111605650581](https://github.com/crabbuild/cellule/actions/runs/37260209483/job/111605650581) | Success on Rust 1.99: 18 selected regressions (new follower CLI, existing reader CLI, count balance, 12 complete observation cases and 3 follower observer cases), production `maintenance-follower`, and host all-target/all-feature warning-denied Clippy. All 11 changed Rust paths match the staged source. The initial snapshot harness incorrectly expected 7 observation cases although all 12 passed; only that selector-count expectation was corrected. No test assertion or profile changed. | +| Snapshot `03261c6`; [mixed-reader diagnostic](https://github.com/crabbuild/cellule/actions/runs/37258791194) | Both independent attempts succeeded with original 3/5/10/20-node scaling, resource limits, assertions and reader SIGKILL. Each scale's original 60-second mixed window committed all 300 scheduled writes with zero misses. Attempt 1 recorded 106808/104358/79189/42047 successful reads; attempt 2 recorded 128368/129112/102646/60227. The isolated snapshot contains tagged router/server/lease probes; no probe was added to the published source. | + +The earlier `9257b18` Compose failure remains unexplained. Two passing diagnostic +attempts do not establish a repair. Raw driver/node/container/provider evidence, +arrival/receipt TSVs and binary hashes are retained in the run's +`mixed-reader-diagnosis-37258791194-1` and `-2` artifacts. These are existing +reader-scaling profile measurements, not W9 fleet-movement qualification. The +published parent's Compose smoke and 3/5/10/20 reader-scaling steps also passed; +its entity/routing measurements were still live. The complete Compose result and +the new follower head's complete CI remain required. PR #57 contains current fetched main and is mergeable. + +Highest next priorities: + +1. Finish native follower and full-head CI, investigate any recurrence of the + unexplained constrained-reader availability failure, and preserve main + ancestry and PR mergeability. +2. Complete combined-role maintenance and remaining W4–W7 successive-boot, + lineage and primitive faults, including actual external Cron/Blob and + failed-process ownership barriers. +3. Deliver W9's committed fleet profile/runner, provider/process, load/soak and + actual mixed-binary evidence, then exercise W10 rollout, rollback and recovery + runbooks. W4–W10 remain incomplete; this command is one further W8 deliverable. + + +## October 5 2026 combined reader/follower maintenance checkpoint + +Parent: `fe8a96cee5464f9ef5002fc1d79ec5bc3cadaa5f`. The six changed Rust +paths match snapshot `8f64d2f274ed4b7ce5c70385f03278d2c27c12f1` exactly; +the Rust delta SHA-256 is +`a336914718a77127c797a00ef91cc7e62c99e9ff19b14c2100082e1cb8d94dd8`. + +Delivered canonical minion `maintenance-roles`, sharing the existing four-boot +follower setup, supervisor, native peer dispatcher, observer, public reconciler +and cleanup path. Managed readers are installed and enrolled before readiness. +The donor has zero local writers but holds both a reader and a foreign follower +tail. Confirmed original tail coverage, both original member retirements and a +durable follower policy alone cannot authorize Finalize. The operation remains +Evacuating with its original reader Established and open. Native reader +evacuation refuses the exact missing-Established replacement while one selected +reader is absent, and another public pass must still retain the donor. Both +eligible replacements must then establish native readers before original view +closure, joined retirement and immutable reader-policy publication. Complete +fresh observation must prove both roles before SettleRoles and Finalize. + +Readback checks both replacements against the captured minimum receipt, original +reader join/retirement, both writer acknowledgements and the original stored +request digest, expiry, sequence and outcome. Final counts are `[1, 0, 0, 0]`; +eight service/receipt checks remain separate from zero fleet writer moves. +Cleanup joins all four nodes, all eleven boot/reader/follower enrollment rows, +the journal and every existing runtime resource ledger. + +The new composed async path initially overflowed the default test stack. +Temporary boundary probes located nested complete reconciliation; separating +reader preparation/completion and boxing reconciliation at the main scenario +boundary fixed the original regression. Both final tests and production commands +passed with normal stack limits and no probes. No profile, deadline, resource +limit or assertion was relaxed. + +| Source / check | Observed result | +| --- | --- | +| Local final Rust 1.97 source | 18 selected regressions passed: 2 follower/combined CLI, 1 reader CLI, 12 complete observations, 3 follower observations. Production combined command, host all-target/all-feature warning-denied Clippy/API docs, format, boundaries/layout, 1348 Markdown links, 137 Rust snippets and 28 SQL/peer assertions with 571 validator links passed. | +| Exact snapshot `8f64d2f`; [native job 111610477201](https://github.com/crabbuild/cellule/actions/runs/37261830490/job/111610477201) | Success on Rust 1.99: all 19 selected regressions, including count balance, passed with exact nonzero selector counts. Both production follower and combined commands exited zero; warning-denied host Clippy passed. The artifact `follower-maintenance-37261830490-1` retains source, binary hash, native environment, resource limits and raw regression/command logs. | +| Earlier parent `f9a476e`; [complete Compose campaign](https://github.com/crabbuild/cellule/actions/runs/37258527934) | Success: smoke, original constrained 3/5/10/20 reader scaling, leased and object-only routing measurements and routing gate. This does not establish a cause or repair for the earlier intermittent `9257b18` failure. | +| Published parent `fe8a96c`; [full Rust workspace](https://github.com/crabbuild/cellule/actions/runs/37260710441) | Success: workspace and MSRV. Follower/object capacity, contract, website and cookbook quality also passed. Other workflows were still live; full current-head CI remains required. | + +This is finite in-process reader/follower maintenance evidence, not completion of +W4–W10. Busy primitive combinations, external Cron/Blob owners, successive boots, +actual process/provider failures, measured fleet movement and mixed binaries +remain unqualified. Highest next priorities: + +1. Finish current-head CI, investigate any recurrence of constrained-reader + availability failure, and keep current main ancestry and PR mergeability. +2. Complete busy primitive/external-owner maintenance and successive-boot role, + lineage and unknown-result fault coverage through the existing public paths. +3. Deliver W9's committed profile/runner and measured process/provider/load/mixed- + binary campaigns, then exercise W10 rollout, rollback and recovery runbooks. + + +## October 5 2026 busy SQL maintenance checkpoint + +Parent: `ba50f650ae85be31e616338a52ab84d6b1139eb5`. The sixteen changed +Rust paths match native snapshot `e3f6dac05e12500ca1f3b36eee895e3368d9b929` +exactly. The Rust delta SHA-256 is +`39d002eb02c35f888ad6e79e43e400ebc63672cf372aa0eda2a4071f4434878c`. + +Two reproducible gaps prevented controlled maintenance under continuous writes: +publication temporarily borrowed the publisher needed by exact quiescence, and +the reference observer discarded unchanged writer demand when the root advanced. +A real root-CAS pause reproduced the first gap; an actual acknowledged mutation +reproduced the second. Both regressions were observed failing before the fix. + +The actor's immutable admission fence now admits exact quiescence and busy +release during publication borrowing. The returned native inventory retains the +same owner fence independently of its optional published position. Host planning +binds it in digest v15 and permits it only under exact Evacuating maintenance and +the existing peak receiver envelope. The reference observer independently checks +canonical owner/epoch, boot, native generation, executable contract and envelope. +Root changes still invalidate complete counts and ordinary movement demand. +Release still prepares the receiver first, joins accepted work/publication, +refreshes readiness and obtains the exact final canonical root; role settlement +and finalization keep their existing full barriers. + +Canonical minion `maintenance-busy` runs two continuously offering SQL command +lanes on one original donor Cell through that same public driver. Each must +receive native admission refusal before the finite 512-command bound. Every +accepted request retains its digest, expiry, sequence and stored result, resolves +exactly on its canonical successor, and contributes one exact audit row. All +twelve original receipts/values survive movement. All clients join before shared +node and journal cleanup on success or failure. A startup-refusal regression +preserves the native source error and checks joined sibling tasks/resources. + +| Check | Observed result | +| --- | --- | +| Local Rust 1.97 source | 29 selected tests passed: 2 busy command/startup, 13 native observations, 6 host inventory/digest contracts, 3 runtime maintenance cases and 5 planner/quiet-maintenance cases. The earlier busy repeat and final production command also passed. | +| Production busy command | 88 accepted commands, 2 admission refusals, 88 exact audit rows; 12 released/activated/retired Cells, 100 receipt checks, `[0, 6, 6]` ownership, 3 joined nodes and 3 retired boots; Completed and exact withdrawal confirmed. | +| Static/API gates | Host/runtime all-target/all-feature warning-denied Clippy and API docs, format, architecture boundaries/layout, 1348 Markdown links, 137 Rust snippets and 28 SQL/peer assertions with 571 validator links passed. | +| Native exact snapshot | [Job 111619387502](https://github.com/crabbuild/cellule/actions/runs/37264858825/job/111619387502) succeeded on Rust 1.99: all 36 nonzero-count regressions, three production commands and warning-denied host/runtime lint. The busy regression preserved 379 commands; production preserved 384 plus the twelve original receipts, with `[0, 6, 6]` ownership and joined/retired boots. Artifact `busy-maintenance-37264858825-1` retains exact source, binary digest, CPU/memory/limits and raw logs. | +| Published parent `ba50f65` | Complete Rust workspace/MSRV, follower/object capacity, contract, website and cookbook campaigns succeeded. Compose run 37262366756 is live; smoke reader scaling and both routing measurements are running. The earlier `fe8a96c` Compose campaign completed successfully. | + +No profiles, deadlines, native stack limits or expected evidence were weakened. +The earlier unexplained constrained-reader availability failure is still not +claimed repaired by passing campaigns. Full W4–W10 completion remains unproven. +Highest next priorities: + +1. Finish exact native and current-head CI, preserve source hashes, push the + qualified checkpoint and keep current main ancestry/PR mergeability. +2. Complete the primitive/external-owner fault matrix, Blob stream/upload/pin + barriers, successive-boot lineage and original accepted-work ownership under + controller/owner/process failures. +3. Deliver W9's versioned fleet profile/runner and measured provider/process, + sustained-load/soak and mixed-binary evidence, then exercise W10 rollout, + rollback, stuck-drain and recovery runbooks. + + +## October 5 2026 maintenance test disk-budget isolation checkpoint + +Parent: `7dad25d0881527ba5a65d9b82cd7ff1aa3978e31`. PR #37 merged on +October 4; continuation PR #57 is mergeable, and fetched main +`80c4fd99c10fc7e91ab34c628ea61fcd488e0cea` is an ancestor. There are no +unmerged paths. Canonical minion remains `crates/cellule-host/minion`. + +The full Rust workspace run +[37265455102](https://github.com/crabbuild/cellule/actions/runs/37265455102) +failed in the new publication/maintenance regression's final disk-zero +assertion. Its runtime used the intentionally process-wide default disk budget, +which reports other live fixtures' reservations too. Holding an unrelated 4096 +byte default-budget reservation reproduced the same assertion failure in one +selected test, without timing or parallel-load dependence. + +The donor and successor test nodes now each use a separate disk budget of the +unchanged default capacity. The unrelated reservation remains live throughout +both native quiescence and release variants. Donor disk, worker, retained-memory +and active-Cell zero assertions remain required; successor disk-zero is also +asserted after exact receipt recovery and native shutdown. The unrelated +reservation must retain all 4096 bytes. Production budget sharing, shutdown and +publication behavior are unchanged. + +Local Rust 1.97: all three public runtime maintenance regressions passed with +parallel selected execution. Format, module layout and architecture boundaries +passed. Exact native snapshot `a3cfd497438614af1bba7f75bd0b570bde7bdfd7` +contains only this test change over the published parent and isolated workflow. +[Run 37267463895](https://github.com/crabbuild/cellule/actions/runs/37267463895) +is running both complete parallel runtime integration suites and warning-denied +runtime lint; success is not yet claimed. The snapshot deliberately excludes +unpublished Blob lifecycle work. + +The published parent's follower/object capacity, contract, website and cookbook +quality checks passed. The earlier Compose run +[37262366756](https://github.com/crabbuild/cellule/actions/runs/37262366756) +finished: reader smoke and object-only routing passed, leased routing failed its +local expired-burst query p99 gate. Retained frozen-binary evidence reports +candidate/baseline p99 ratio 2.13677 (3.059677 ms / 1.431917 ms), p95 ratio 1.06777, +and unchanged paced throughput. Its cause remains open; no gate is weakened or +performance fix claimed. Complete current-head checks are still required. + +Highest next priorities: + +1. Complete exact native cleanup and full PR CI; diagnose the retained leased + routing tail-latency failure and constrained-reader availability gap. +2. Finish and qualify unpublished Blob accepted-I/O ownership and integrate + logical upload/range/pin lifetimes into the existing native role barriers. +3. Complete remaining W4–W10 fault, boot-lineage, process/provider, load, + mixed-version and rollout/rollback/recovery evidence. The plan remains open. + + +## October 5 2026 original Blob lifetime checkpoint + +Parent: `973e0eeaddb51158cdc7d075e06c2feb84795cd5`. PR #37 is merged; +continuation #57 is mergeable and includes current main +`80c4fd99c10fc7e91ab34c628ea61fcd488e0cea`. No unresolved index entries or +conflict markers remain in the three reported files. + +The original Blob store now owns at most 64 accepted operations through shared, +irreversible admission. A retained supervisor joins the original native future +and preserves its first source-bearing failure, including provider panic +JoinError. Caller loss or cancellation of a close waiter leaves original work +running. Forced Tokio runtime loss records an unjoined original supervisor, +closes admission and refuses local closure even after the provider eventually +finishes. Known accepted counts alone cannot prove joining. The regression uses +an actual paused spawn_blocking provider operation and observes its bytes +published after runtime loss; removing the unjoined guard reproduced the false +closure before restoring the passing implementation. + +Public namespace operations share that owner. Convenience mutation retains +part staging through its normal Cell command response. Preparation retains +staging/preparation and returns the existing caller-owned PreparedCommand; +closure cannot revoke its later execution. Range queries retain metadata lookup +and every bounded part read without readmission between parts. Head/list and +non-part mutations use the same admission. Codec, ID, part digest/path, SQLite +manifest and durable response contracts remain unchanged. GC accepts a complete +Arc> retained through original listing and deletion, without +truncation or unbounded copying. + +`CellNode::install_blob_artifact_store` installs the configured provider as an +existing owned facility before readiness, after the task group. The returned +clone configures the existing client. Canonical reverse facility drain closes +and joins that original store before runtime shutdown. Duplicate, closed-store +and late installation are refused. Lost original joining returns a drain error; +cancelled waiters cannot fabricate Stopped or close an unrelated store. + +Public regressions hold actual SQLite work ahead of range metadata, close +admission, cancel callers/close waiters and finish all original part reads under +one lifetime. Removing whole-query ownership reproduced the admission failure +after cleanup. Cancelled upload coverage resolves original prepared evidence as +Committed before replay; sequence 2 and one upload are preserved. Host coverage +pauses actual GC listing, cancels its caller/shutdown waiter, retains complete +references and finishes one original sweep with joined drain and zero resources. +Five provider cases cover original put/read/GC, failure and panic sources, all +64 retained jobs, invalid inputs and forced runtime loss. + +The published parent's full Rust job passed workspace tests but failed one +minion corruption test: it unwrapped construction of an inconsistent native +identity that the strengthened constructor correctly refuses. The exact local +case reproduced that failure. The fixture now accepts construction refusal and +checks each exact error; corruption cases expand from eight to fourteen, +including coherent wrong fences that must fail original-writer attachment. +No production validation, profiles, deadlines or assertions are weakened. + +Verification: + +- Local Rust 1.97: 25 Blob unit/codec cases, five public Blob/Cron cases, two + public host Blob cases, ten successor observation cases and the existing typed + application consumer passed. Warning-denied host/runtime lint and API docs, + format, boundaries/layout and document gates passed. +- Final Blob snapshot `7fd0a2e5172618b7ab90cd7918af42227ab4e1b8` passed + [run 37271127229](https://github.com/crabbuild/cellule/actions/runs/37271127229): + 25 Blob cases, all 52 parallel public primitive cases and all 129 parallel + public node cases (206 total), warning-denied lint/API docs and static gates + on Rust 1.99. All 22 recorded snapshot paths matched current bytes before + this progress update; delta SHA-256 + `966568654b6fb252f2fd6635265c14864034f67c0e607d5009a0513a17291398`. +- Minion snapshot `9dad2e91d80bc9852bc2e638e663b5d58ba24eb2` passed + [run 37270034524](https://github.com/crabbuild/cellule/actions/runs/37270034524): + ten exact observation contracts, all 382 serial minion cases in 804.49 seconds, + the typed application consumer and warning-denied lint. This snapshot precedes + the forced-runtime-loss guard; it proves the minion fix at that source. + Delta SHA-256 + `14c5313674464a9fed5e48a512ccfb42beac17f3a20f06e0fc5d948317c4f8dc`. +- Earlier parallel maintenance cleanup passed two complete runtime suites, each + 231 passed/4 ignored, in + [run 37267463895](https://github.com/crabbuild/cellule/actions/runs/37267463895). + Full CI on the newly published head remains required. The published parent + also passed follower/object capacity, contract, website, all 21 cookbook + scenarios, Compose smoke and leased routing. Object-only routing was still + running at the final observation; no outcome is inferred from that state. + Passing leased routing does not explain the retained earlier p99 failure. + +The sole dependency change adds already-resolved async-trait as a host test +dependency for the real ObjectStore decorator; resolved versions are unchanged. +This checkpoint supplies original local operation ownership. Returned prepared +commands, Cell-scoped upload/stream/read/backup pins, migration, cross-Cell GC +references and unknown remote outcomes still require canonical owner coverage. +BlobInventory and missing Blob maintenance cost remain blocking. W4–W10 remain +incomplete. + +Highest next priorities: + +1. Finish current-head CI and keep the continuation mergeable; diagnose the + retained leased routing p99 and constrained-reader availability failures. +2. Add Cell-scoped Blob logical/pin and global retention proof through the + original owner, then qualify maintenance handoff and native role settlement. +3. Complete remaining fault/boot-lineage, W9 process/provider/load and + mixed-version campaigns, then exercised W10 rollout/rollback/recovery. + + +## October 5 2026 prepared Blob admission checkpoint + +Parent: `b8b99002143b15013069444f31000da5aae255ab`. +Prepared command execution for a compiled Blob namespace with a configured +artifact store now uses that same original admission and retained native owner. +This applies to ordinary client commands, returned namespace commands, clones +and restored snapshots without adding a wrapper, field or persisted codec. +Closure refuses new dispatch; already accepted execution completes after caller +loss. Convenience mutation calls the same native dispatch under its existing +whole-operation owner, preserving accepted staging/publication across closure +and using one slot. Original command identity, digest, body, snapshot bytes, +request resolution and durability are unchanged. + +Three new public regressions prove closed originals/clones/restores cannot +publish a manifest, cancelled accepted execution survives a held real SQLite +worker and joins after closure, and open replay preserves exact receipt/output +at sequence 2 with only one part upload. The cancelled close waiter is confirmed +polled and Pending before it is aborted. Previously committed results remain +resolvable after closure. The existing upload-cancellation case now reads exact +original result bytes and sequence through resolution; open replay separately +retains its deduplication assertions. Before production changes, the closed +prepared-dispatch regression failed after complete fixture cleanup. + +Full parent Rust CI failed its upload job-count assertion with 2 rather than 1: +[run 37271901314](https://github.com/crabbuild/cellule/actions/runs/37271901314). +The fixture assumed the prior reply also meant its supervisor had joined. +A temporary delayed-cleanup probe reproduced that exact result. Fixture setup +now waits for prior original jobs to join before measuring the next paused job; +the same probe passed with the exact one-job assertion. The probe is removed. +Production ownership ordering and all qualification profiles/deadlines remain +unchanged. + +Local Rust 1.97: all eight public Blob/Cron cases, all five command-snapshot +contracts and the typed application consumer passed, along with warning-denied +host/runtime lint/API docs, format, architecture/layout, document links/fences +and SQL/peer contracts. Exact native snapshot +`f817476440399c00302fc23227be8d8b656a8cfd` passed +[run 37273384545](https://github.com/crabbuild/cellule/actions/runs/37273384545) +on Rust 1.99: all 25 Blob cases, all 55 parallel primitive cases, all 129 parallel +node cases, five snapshot contracts and one application consumer (215 total), +plus warning-denied lint/API docs and static gates. All ten qualification paths +matched current bytes before this final progress update; delta SHA-256 +`363fd210b356ce16ca81ec769e292a9cb0141068009162910b646b66e8515050`. +The first qualification attempt passed 214 native cases before naming an +unavailable application test target; its terminal failure was retained, the +command was corrected to the existing integration target, and the final run +passed. No live run was cancelled or profile weakened. + +The parent's follower/object capacity and Compose smoke passed; its leased and +object-only routing jobs were still live at the final observation. Full CI on +the newly published head remains required. Current main remains an ancestor and +PR #57 is mergeable; no unresolved index entries exist. + +These are local dispatch barriers. Cell-scoped upload/stream/read/backup pins, +complete cross-Cell and pinned-root retention, migration, unknown remote effects, +and external client/provider capabilities still need original owner coverage. +BlobInventory and missing Blob maintenance cost remain blocking. The full +W4–W10 goal remains open. + +Highest next priorities: + +1. Complete exact native qualification and new-head CI; preserve main ancestry + and PR mergeability, and diagnose retained routing/availability failures. +2. Bind Blob original-operation and pin inventories to exact Cell/boot scope, + then integrate complete retention and maintenance handoff barriers. +3. Complete remaining primitive/role faults, successive boot lineage, W9 + process/provider/load/mixed-version campaigns and W10 rollout/runbooks.