DOCUMENTATION

Checkpoints & recovery

Checkpoint 1, restore evidence, recovery hierarchy and control-plane resilience.

Docs menu

Checkpoint 1 is the human-attested, immutable historical baseline for a scope. It combines state facts with recovery evidence so that a service does not receive agent write authority on the strength of an untested backup claim.

Checkpoint 1

The manifest cannot be edited or deleted by an agent. Evidence may later become stale, but the original record remains intact. A checkpoint is not automatically the correct rollback destination after data has changed; recovery must still evaluate compatibility, RPO/RTO and external effects.

Minimum evidence manifest

Evidence groupExamples
Application and infrastructureImage/release digest, config hash, IaC revision, provider resource IDs.
Database and restoreSchema digest, backup ID, PITR coverage, restore job and startup/health/data checks.
Network and dependenciesRoutes, DNS, load balancer policy, service graph, owners and external limits.
Operational baselineHealth, performance/saturation, cost and evidence timestamps.

Restore verification

  1. Choose a consistency boundary and recovery point.
  2. Validate backup integrity and the required key/log chain.
  3. Restore in an isolated environment with production DNS and side effects blocked.
  4. Start a compatible application version with test-safe secrets.
  5. Run health, read/write synthetic and data-consistency checks.
  6. Record actual RTO, RPO/window and redacted machine-readable evidence.
  7. Require human review and signature; a FAIL or UNKNOWN result cannot become VERIFIED.

Recovery hierarchy

The least disruptive safe recovery is selected by compatibility, authority, data-loss window and blast radius—not by a fixed sequence. Candidates may include restarting a process, rolling back an image/configuration, recovering infrastructure, or a database restore/PITR with data-owner authority. External payments, consumed queues and new data may need compensation or reconciliation rather than time travel.

DESIGN LIMIT

The MVP proposes a PostgreSQL restore-test fixture and imported customer backup evidence. Automatic restoration of a production database remains outside MVP scope.

Control-plane recovery

After control-plane recovery, the environment remains frozen. The audit chain is checked against an independent witness, old permits are revoked, in-flight work is reconciled and provider logs are compared before a human requests return to normal. Missing evidence is not treated as evidence that nothing happened.