Document module lifecycle recovery operations

This commit is contained in:
2026-08-03 07:02:23 +02:00
parent 0b4e601719
commit cdcb477b55
2 changed files with 6 additions and 1 deletions
+5
View File
@@ -66,3 +66,8 @@ the Workflow handoff remains the reconciliation surface. Verify the provider,
then record **Effect confirmed** to continue without replay or **Effect absent** then record **Effect confirmed** to continue without replay or **Effect absent**
to enable a deliberate retry. Instance, trigger-delivery, and timer leases show to enable a deliberate retry. Instance, trigger-delivery, and timer leases show
which runtime currently owns transition authority across hosts. which runtime currently owns transition authority across hosts.
For Core module lifecycle, the installer run record supplies the operation id.
Treat `recovery_required` and `outcome_unknown` as a deployment-wide stop: verify
the hashed package/database evidence, complete the declared rollback or forward
repair, and reconcile the operation before another install or live graph change.
The local installer lock is not a substitute for this database fence.
+1 -1
View File
@@ -139,7 +139,7 @@ manifest = ModuleManifest(
id="ops.runtime-coordination-and-recovery", id="ops.runtime-coordination-and-recovery",
title="Drain runtime nodes and inspect recovery evidence", title="Drain runtime nodes and inspect recovery evidence",
summary="Ops projects shared runtime heartbeats, replica gaps, drain controls, and recovery states that require operator attention.", summary="Ops projects shared runtime heartbeats, replica gaps, drain controls, and recovery states that require operator attention.",
body="Use the runtime table to identify stale or composition-skewed API and worker replicas. Drain before replacement so API readiness closes and workers stop taking new queue work; cancellation is available while the node is still draining. The recovery table reports durable Core recovery operations. A rejected operation is a verified provider rejection and needs no recovery; outcome-unknown and recovery-required operations still require reconciliation through the owning module. Mail SMTP and IMAP APPEND entries use stable attempt identifiers and digest-only evidence: reconcile the Mail command from provider evidence, never by replaying the original effect from Ops. Dataflow database-only runs are atomic, while published-output runs use forward recovery: reconcile the recorded output digest and sink idempotency key before allowing another publication. Backup status separately projects only the sanitized deployment verification receipt: a verified status identifies a coordinated recovery point and isolated restore drill, while absent, expired, or invalid evidence blocks a release-changing migration.", body="Use the runtime table to identify stale or composition-skewed API and worker replicas. Drain before replacement so API readiness closes and workers stop taking new queue work; cancellation is available while the node is still draining. The recovery table reports durable Core recovery operations. A rejected operation is a verified provider rejection and needs no recovery; outcome-unknown and recovery-required operations still require reconciliation through the owning module. Core module-lifecycle entries block every later install or live graph change: use the installer run id to verify package, backup, migration, and health evidence before rollback or forward repair. Mail SMTP and IMAP APPEND entries use stable attempt identifiers and digest-only evidence: reconcile the Mail command from provider evidence, never by replaying the original effect from Ops. Dataflow database-only runs are atomic, while published-output runs use forward recovery: reconcile the recorded output digest and sink idempotency key before allowing another publication. Backup status separately projects only the sanitized deployment verification receipt: a verified status identifies a coordinated recovery point and isolated restore drill, while absent, expired, or invalid evidence blocks a release-changing migration.",
documentation_types=("admin", "user"), documentation_types=("admin", "user"),
audience=("operator", "system_admin"), audience=("operator", "system_admin"),
conditions=( conditions=(