# Recovery And Rollback Guarantees ## Principle GovOPlaN must prove recovery claims with durable state recorded before and after side effects. A failed operation is not automatically rolled back merely because the previous application image still exists. Database schema and external effects may make release rollback unsafe. Core therefore distinguishes five recovery modes: | Mode | Meaning | | --- | --- | | `atomic` | One database transaction either commits or rolls back. No external effect is claimed. | | `compensation` | Durable evidence identifies explicit inverse actions for completed effects. | | `snapshot_restore` | A separately verified backup reference and restore procedure exist. | | `forward_recovery` | Repair or resume the current version; reverting code/configuration is not claimed safe. | | `irreversible` | No automated recovery is claimed and an approval reference is mandatory. | An operation plan must include verification steps. Compensation requires named compensation steps, snapshot restore requires a verified backup reference, forward recovery requires repair steps, and irreversible work requires explicit approval. ## Core Recovery Ledger Core stores recovery operations and append-only, hash-chained checkpoints in PostgreSQL. The contract provides: - installation/module/resource identity; - an idempotency key bound to a canonical request hash; - recovery mode, preconditions, verification steps, and references; - optional runtime lease holder and fencing token; - explicit planned, prepared, running, recovery-required, recovering, succeeded, recovered, failed, outcome-unknown, and manual-intervention states; - an evidence-chain head and sequence count; - rejection of plaintext secrets in metadata or evidence. Preparation cannot succeed without durable precondition evidence. A non-atomic operation cannot hide a partial effect by transitioning directly from running to failed. Success and recovery require explicit verification evidence with at least one check. The ledger verifies its hash chain before evidence is trusted. This is a platform contract, not an assertion that every existing module operation has adopted it. Module operations with external or multi-resource effects must be migrated to the ledger before claiming these guarantees. ## Deployment Journal Every `govoplan-deploy apply` begins an operation journal before it pulls images or mutates runtime state. The private installation directory records: ```text operations//operation.json operations//before/ applied-state/ ``` Each stage is hash-chained. The previous applied bundle is copied with per-file SHA-256 evidence. Applied state is replaced atomically after health verification; an interrupted replacement restores its previous directory. New journals also bind the complete desired deployment plan, snapshot availability, failure summary, recovery mode, and terminal status to the evidence chain. Recovery verifies every snapshot entry and checksum before it changes any live bundle file, then replaces each live file atomically. A crash between file replacements is recoverable by rerunning the same idempotent recovery command under the deployment lock. Inspect operations: ```sh python tools/deployment/govoplan-deploy.py operations \ --directory /srv/govoplan/default ``` Recover the latest failed operation, or provide its identifier: ```sh python tools/deployment/govoplan-deploy.py recover \ --directory /srv/govoplan/default \ --operation-id 20260801T120000Z-1234abcd ``` Add `--apply` only after reviewing the reported action. ## Migration Boundary Before database migration starts, a failed deployment with a verified prior applied snapshot may restore its prior release/configuration bundle and reconcile that desired state. As soon as migration starts, the journal permanently changes to `forward_recovery`. It will not restore old application configuration because old code may not understand the new schema. Recovery then means one of: 1. fix and re-run the current release; 2. deploy a newer compatible repair release; 3. restore a separately verified, coordinated database/object/key backup and then deploy the matching release. The deployment tool does not create or validate that database backup. A `backup-required` annotation on the Kubernetes migration Job is an operator gate, not backup evidence. Production automation must provide a backup hook or external backup controller whose artifact, timestamp, scope, encryption key, and restore test can be referenced from the recovery record. ## Scaled Nodes Recovery actions must be safe across replicas: - drain affected API and worker nodes before incompatible changes; - use the deployment-wide PostgreSQL advisory lock for schema migration; - use distributed leases and fencing tokens for singleton or externally visible effects; - use idempotency keys for retried commands and jobs; - retain shared object keys and database references until deletion succeeds; - classify uncertain external outcomes instead of retrying blindly; - verify the exact software/module composition after replacement. Campaign generated-message objects now follow this model: object writes are compensated when a build fails before database commit, workers verify stored size and digest before delivery, and retention keeps the database reference when storage deletion fails. A hard process loss between object creation and database commit can still leave an orphan object; an inventory reconciler is a separate operational slice and must use the build-specific object prefix. ## Required Drills Record evidence for at least these scenarios before production acceptance: 1. Kill an API replica and verify traffic continues without session loss. 2. Drain and replace a worker while work is queued and while one job is active. 3. Start two migration jobs and verify only one mutates schema. 4. Kill the fenced scheduler and verify one replacement acquires a higher fencing token. 5. Fail deployment before migration and restore the prior applied bundle. 6. Fail deployment after migration and verify old configuration is not restored. 7. Restore PostgreSQL, object storage, and encryption keys to one coordinated recovery point and verify representative object hashes. 8. Interrupt object storage during Campaign build and retention and verify compensation/reference-preservation behavior. 9. Tamper with a deployment or Core recovery checkpoint and verify chain validation rejects it. No runbook, status badge, or green health endpoint substitutes for a dated, repeatable restore drill against the actual deployment topology.