Skip to content

Require epoch-bound guest readiness acknowledgment - #117

Merged
jiashuoz merged 2 commits into
feat/guest-reconnect-transportfrom
feat/guest-reconnect-readiness
Oct 7, 2026
Merged

jiashuoz merged 2 commits into
feat/guest-reconnect-transportfrom
feat/guest-reconnect-readiness

Conversation

@jiashuoz

@jiashuoz jiashuoz commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

The reconnecting guest currently returns its stream immediately after applying configuration, leaving the host without a readiness boundary before ordinary relay traffic can interleave. Add strict guest_reconnect_ready / guest_reconnect_ready_ack frames tied to the accepted epoch. The guest stays offline until a matching acknowledgment arrives within five seconds (also bounded by delivery TTL and caller cancellation).

Lost, malformed, wrong-epoch or late acknowledgments fail closed. The consumed epoch/token require a fresh signed attempt. The executable TCP probe drops an acknowledgment and reconnects with a fresh epoch while retaining the same PTY shell process. Existing fresh boot remains unchanged.

Depends on #116; base is feat/guest-reconnect-transport. This is still a disabled protocol prerequisite: no shipping listener enablement, fresh host configuration resolver, guarded relay takeover or deployment. The design document records required host ownership checks and acknowledgment/publication ordering.

Validation: full Linux CI passed (run 36816477521): make verify including Docker tests, non-root jail ownership, CLI/client race tests, and fleet syntax. Sessiond, relay and shared protocol race suites also pass locally. The built guest executable TCP/PTY probe passed after the local suite. Local make verify passes with Docker explicitly unavailable; the initial Docker attempt encountered the existing credential-helper/registry problem, and an existing latency cleanup timing failure passed in isolation. Independent and adversarial reviews passed with no required findings. Additional adversarial race-enabled probes passed the actual five-second blocked-write/read deadlines and pipelined-byte continuity.

…118)

* Wire authenticated guest handoff and guarded recovery admission

* Fence delayed RPC delivery and reject ambiguous bootstrap preambles

* Require exclusive state ownership for every microVM runner

* Fence cold resume placement before launch and reconcile uncertain results

* Preserve cold launch ownership across crashes and terminal races

* Retain fresh and interrupted VM launches until process exit is proven

* Use exact pending placement capacity in API explanations

* Require exact instance and original process lifetime before VM teardown

* Require host authority support to negotiate guest reconnect

* fix: revalidate VM lifetime before shutdown escalation

* feat: compose durable standalone guest recovery

* fix(sessiond): launch the command supplied by microVM boot configuration

* test(controld): await asynchronous placement dispatch before asserting

* fix(driver): recognize original VM exit before reaping

* fix(driver): require whole process exit before teardown

* fix(microvm): wait for remaining recovered VMM threads

* fix(microvm): resolve recovered cold resume configuration
@jiashuoz
jiashuoz marked this pull request as ready for review October 7, 2026 16:44
@jiashuoz
jiashuoz merged commit d3f135d into feat/guest-reconnect-transport Oct 7, 2026
1 check passed
@jiashuoz
jiashuoz deleted the feat/guest-reconnect-readiness branch October 7, 2026 16:44
jiashuoz added a commit that referenced this pull request Oct 7, 2026
* feat(driver): bound guest admission and bridge reconnect proofs

* Require epoch-bound guest readiness acknowledgment (#117)

* feat: require epoch-bound guest readiness acknowledgment

* Integrate authenticated guest reconnect and guarded runner recovery (#118)

* Wire authenticated guest handoff and guarded recovery admission

* Fence delayed RPC delivery and reject ambiguous bootstrap preambles

* Require exclusive state ownership for every microVM runner

* Fence cold resume placement before launch and reconcile uncertain results

* Preserve cold launch ownership across crashes and terminal races

* Retain fresh and interrupted VM launches until process exit is proven

* Use exact pending placement capacity in API explanations

* Require exact instance and original process lifetime before VM teardown

* Require host authority support to negotiate guest reconnect

* fix: revalidate VM lifetime before shutdown escalation

* feat: compose durable standalone guest recovery

* fix(sessiond): launch the command supplied by microVM boot configuration

* test(controld): await asynchronous placement dispatch before asserting

* fix(driver): recognize original VM exit before reaping

* fix(driver): require whole process exit before teardown

* fix(microvm): wait for remaining recovered VMM threads

* fix(microvm): resolve recovered cold resume configuration
jiashuoz added a commit that referenced this pull request Oct 7, 2026
* feat(sessiond): authenticate guest reconnect before configuration refresh

* fix(sessiond): refresh exec environments before reconnect readiness

* docs: clarify configuration state after reconnect expiry

* feat(driver): bound guest admission and bridge reconnect proofs (#116)

* feat(driver): bound guest admission and bridge reconnect proofs

* Require epoch-bound guest readiness acknowledgment (#117)

* feat: require epoch-bound guest readiness acknowledgment

* Integrate authenticated guest reconnect and guarded runner recovery (#118)

* Wire authenticated guest handoff and guarded recovery admission

* Fence delayed RPC delivery and reject ambiguous bootstrap preambles

* Require exclusive state ownership for every microVM runner

* Fence cold resume placement before launch and reconcile uncertain results

* Preserve cold launch ownership across crashes and terminal races

* Retain fresh and interrupted VM launches until process exit is proven

* Use exact pending placement capacity in API explanations

* Require exact instance and original process lifetime before VM teardown

* Require host authority support to negotiate guest reconnect

* fix: revalidate VM lifetime before shutdown escalation

* feat: compose durable standalone guest recovery

* fix(sessiond): launch the command supplied by microVM boot configuration

* test(controld): await asynchronous placement dispatch before asserting

* fix(driver): recognize original VM exit before reaping

* fix(driver): require whole process exit before teardown

* fix(microvm): wait for remaining recovered VMM threads

* fix(microvm): resolve recovered cold resume configuration
jiashuoz added a commit that referenced this pull request Oct 7, 2026
…114)

* Authorize guest reconnect through one runner control connection

* feat(sessiond): retain guest identity and authenticate reconnect (#115)

* feat(sessiond): authenticate guest reconnect before configuration refresh

* fix(sessiond): refresh exec environments before reconnect readiness

* docs: clarify configuration state after reconnect expiry

* feat(driver): bound guest admission and bridge reconnect proofs (#116)

* feat(driver): bound guest admission and bridge reconnect proofs

* Require epoch-bound guest readiness acknowledgment (#117)

* feat: require epoch-bound guest readiness acknowledgment

* Integrate authenticated guest reconnect and guarded runner recovery (#118)

* Wire authenticated guest handoff and guarded recovery admission

* Fence delayed RPC delivery and reject ambiguous bootstrap preambles

* Require exclusive state ownership for every microVM runner

* Fence cold resume placement before launch and reconcile uncertain results

* Preserve cold launch ownership across crashes and terminal races

* Retain fresh and interrupted VM launches until process exit is proven

* Use exact pending placement capacity in API explanations

* Require exact instance and original process lifetime before VM teardown

* Require host authority support to negotiate guest reconnect

* fix: revalidate VM lifetime before shutdown escalation

* feat: compose durable standalone guest recovery

* fix(sessiond): launch the command supplied by microVM boot configuration

* test(controld): await asynchronous placement dispatch before asserting

* fix(driver): recognize original VM exit before reaping

* fix(driver): require whole process exit before teardown

* fix(microvm): wait for remaining recovered VMM threads

* fix(microvm): resolve recovered cold resume configuration
jiashuoz added a commit that referenced this pull request Oct 7, 2026
* Define bounded guest reconnect RPC messages and decoders

* Bind guest reconnect authorization to the runner control connection (#114)

* Authorize guest reconnect through one runner control connection

* feat(sessiond): retain guest identity and authenticate reconnect (#115)

* feat(sessiond): authenticate guest reconnect before configuration refresh

* fix(sessiond): refresh exec environments before reconnect readiness

* docs: clarify configuration state after reconnect expiry

* feat(driver): bound guest admission and bridge reconnect proofs (#116)

* feat(driver): bound guest admission and bridge reconnect proofs

* Require epoch-bound guest readiness acknowledgment (#117)

* feat: require epoch-bound guest readiness acknowledgment

* Integrate authenticated guest reconnect and guarded runner recovery (#118)

* Wire authenticated guest handoff and guarded recovery admission

* Fence delayed RPC delivery and reject ambiguous bootstrap preambles

* Require exclusive state ownership for every microVM runner

* Fence cold resume placement before launch and reconcile uncertain results

* Preserve cold launch ownership across crashes and terminal races

* Retain fresh and interrupted VM launches until process exit is proven

* Use exact pending placement capacity in API explanations

* Require exact instance and original process lifetime before VM teardown

* Require host authority support to negotiate guest reconnect

* fix: revalidate VM lifetime before shutdown escalation

* feat: compose durable standalone guest recovery

* fix(sessiond): launch the command supplied by microVM boot configuration

* test(controld): await asynchronous placement dispatch before asserting

* fix(driver): recognize original VM exit before reaping

* fix(driver): require whole process exit before teardown

* fix(microvm): wait for remaining recovered VMM threads

* fix(microvm): resolve recovered cold resume configuration
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant