Recovery

Recovery for a tamper-evident chain is restore, then verify, never rewrite. A restore is not done when the data is back; it is done when the chain re-verifies. This runbook pairs with the backup runbook.

Restore the database

scripts/restore.sh, from the agledger-ai/install repository, takes a backup tarball, stops the application containers, drops and recreates the database, runs pg_restore --no-owner, and brings the stack back up:

./scripts/restore.sh /tmp/backup/backup-compose-2026-06-09-155745.tar.gz
./scripts/restore.sh --non-interactive <tarball>   # skip the confirmation prompt
./scripts/restore.sh --force <tarball>             # see "What it refuses to do" below
./scripts/restore.sh --keep-version <tarball>      # stay on the installed release

backup.sh names each archive backup-<compose project>-<timestamp>.tar.gz, so the project segment is whatever COMPOSE_PROJECT_NAME the install runs under; the chart's backup Job writes backup-<timestamp>.tar.gz with no project segment.

A backup records the release it was taken from, and by default the restore returns the install to that release before starting it, because the dump and the schema have to agree: starting a newer release over an older dump re-applies that release's migrations on top of it. --keep-version keeps the release installed now and says what that means. An archive with no version record (every archive the chart's CronJob writes) leaves the install where it is.

Privileges come back with the data. pg_restore applies the dump's own grants: the DML the engine writes with, the ALTER DEFAULT PRIVILEGES that keep later migrations reachable, and the append-only REVOKEs on the audit chain. A least-privilege runtime role can therefore serve the restored database without any grant being reapplied by hand. Restoring onto a server that does not carry the same roles (a cross-server DR) reports the grants it could not apply and carries on with the rows intact.

It prints the restore sequence and the container orchestration as services come back, and ends with a banner between two rules: Restore complete., then a line saying which release the install is on and whether that is the release the backup came from (This install is on <version>, the version this backup came from.). The checks and reconciliation notes that follow the banner are the rest of this runbook.

For an external database the script needs psql on the host, and connects as DATABASE_URL_MIGRATE when that is set (the owner role, for the DDL) and DATABASE_URL otherwise. The database it drops and restores into is the one DATABASE_URL names. POSTGRES_DB configures the bundled PostgreSQL container and is ignored here. The role needs CREATEDB and must own the database it drops, and no other session may be connected to it.

On Kubernetes

Restore is still a restore.sh run. The chart ships no restore Job, on purpose: a restore has questions a template cannot answer for you, and every one of them is asked before anything is dropped. Which archive. Whether the connection has the rights the audit event trigger needs. Whether the database it is about to drop is actually this Server's. A Job would answer them by assuming.

What the chart does ship is the backup half (backup.cronJob), which writes the archive shape this script reads, and the steps that follow a restore (postRestore, below), which need the release's configuration. There is no restore chart value; restoring a cluster is the recipe below, driven by hand.

Two flags make the cluster case explicit. --target external names the database to drop: a DATABASE_URL passed in from the environment that names localhost looks exactly like the bundled container, and the script refuses to guess which one to drop rather than pick. --no-start leaves the Compose stack down when the restore finishes, so the run does not bring up a local API and Worker beside the cluster's own workloads. Without it the script's last step would be docker compose up -d --wait, a second writer on the restored chain, so on a host with no compose/.env the script refuses to run without it.

The host also has to reach the database from the one-off containers the script runs, not only from its own shell: the revocation read, the migration and the privilege check run in the api image. A kubectl port-forward on a laptop is not enough on its own, because a container cannot reach the laptop's localhost. Run it from a bastion or a pod that reaches the database by a routable address, or forward with --address 0.0.0.0 and name the host's own address in DATABASE_URL.

Two more things come from the release rather than from a checkout, and the script refuses before it stops anything when either is missing or wrong.

Workloads are addressed by label rather than by name, because the chart's fullname is <release>-agledger-chart and a fullnameOverride changes it. The recipe reads the image, the Secret and the app version off the release itself for the same reason:

NS=<namespace>; REL=<release>
SEL="app.kubernetes.io/instance=$REL,app.kubernetes.io/component in (api,worker)"

# 1. a pod that holds the backup PVC open long enough to read from. Not --rm:
#    that deletes it the moment the listing exits, and step 3 needs it alive.
kubectl run agl-backups -n "$NS" --restart=Never --image=busybox \
  --overrides='{"spec":{"containers":[{"name":"agl-backups","image":"busybox","command":["sleep","3600"],
    "volumeMounts":[{"name":"b","mountPath":"/backups"}]}],
    "volumes":[{"name":"b","persistentVolumeClaim":{"claimName":"'"$REL"'-agledger-chart-backups"}}]}}'
kubectl wait -n "$NS" --for=condition=Ready pod/agl-backups
kubectl exec -n "$NS" agl-backups -- ls -1t /backups

# 2. record the replica counts, then stop the writers. The script does this
#    for you on Compose; here it is yours.
kubectl get deploy -n "$NS" -l "$SEL" \
  -o jsonpath='{range .items[*]}{.metadata.name}={.spec.replicas}{"\n"}{end}'
kubectl scale -n "$NS" --replicas=0 deploy -l "$SEL"

# 3. copy the archive out, read the release's version, image and database URLs,
#    and restore from a host that can reach the database. <host>:<port> is the
#    address this host and its containers reach the database at.
kubectl cp "$NS"/agl-backups:/backups/backup-<timestamp>.tar.gz ./backup-<timestamp>.tar.gz
API="app.kubernetes.io/instance=$REL,app.kubernetes.io/component=api"
VERSION=$(helm list -n "$NS" -f "^$REL\$" -o json | jq -r '.[0].app_version')
IMAGE=$(kubectl get deploy -n "$NS" -l "$API" -o jsonpath='{.items[0].spec.template.spec.containers[0].image}')
SECRET=$(kubectl get deploy -n "$NS" -l "$API" \
  -o jsonpath='{.items[0].spec.template.spec.containers[0].envFrom[1].secretRef.name}')
db_url() {
  kubectl get secret -n "$NS" "$SECRET" -o jsonpath="{.data.$1}" | base64 -d \
    | sed 's#@[^/]*/#@<host>:<port>/#'
}
DATABASE_URL="$(db_url DATABASE_URL)" DATABASE_URL_MIGRATE="$(db_url DATABASE_URL_MIGRATE)" \
  AGLEDGER_IMAGE_PIN="$IMAGE" ALLOW_DB_WITHOUT_SSL=true \
  ./scripts/restore.sh --target external --no-start --version "$VERSION" ./backup-<timestamp>.tar.gz

# 4. run the post-restore steps in the cluster. restore.sh ends by printing these
#    commands with the input directory it wrote and the Job name filled in.
PR=$(kubectl get cronjob -n "$NS" -l "app.kubernetes.io/instance=$REL,app.kubernetes.io/component=post-restore" \
  -o jsonpath='{.items[0].metadata.name}')
kubectl delete secret -n "$NS" "$PR" --ignore-not-found
kubectl create secret generic -n "$NS" "$PR" --from-file=<input directory>
kubectl create job -n "$NS" --from=cronjob/"$PR" <job name>
kubectl wait -n "$NS" --for=jsonpath='{.status.conditions[0].status}'=True --timeout=15m job/<job name>
kubectl logs -n "$NS" job/<job name>
kubectl get job -n "$NS" <job name> -o jsonpath='succeeded={.status.succeeded} failed={.status.failed}{"\n"}'

# 5. bring each Deployment back to the count you recorded, and clean up
kubectl scale -n "$NS" --replicas=<recorded count> deploy/<name>
kubectl delete secret -n "$NS" "$PR"
kubectl delete pod -n "$NS" agl-backups

That is the bundled database. ALLOW_DB_WITHOUT_SSL=true is what the chart itself sets for it, since it speaks plaintext on the cluster network; drop it for an external database, whose URLs already carry their TLS parameters. An external install that never set secrets.databaseUrlMigrate has one role and no DATABASE_URL_MIGRATE in its Secret: pass DATABASE_URL alone.

Step 3 exits 0 only once the data is restored and migrated, and its banner says the restore is not complete until step 4 has run. When the migration fails, the script says so, starts nothing, and exits 1: leave the Deployments at 0, fix what the migration reported, run the migration command it prints, which does not drop anything again, and then step 4. When the migration refuses the restored database as another release line's, no re-run of this release will migrate it; restore that archive into the release that wrote it.

Step 4 is what restore.sh does on its own after the migration on Compose: re-apply the revocations made after the backup, write the restore marker the next boot turns into a RESTORE_EPOCH chain entry, and compare the external anchors against the restored database. Those need the Server's whole configuration (the signing key, API_KEY_SECRET, the external URL, the anchor settings), which only the release holds, so on a cluster they run as a Job with the api's image, env sources, volumes and security context. The chart renders it as a CronJob that never fires on its own: suspended, on a schedule naming February 31. kubectl create job --from starts one. It reads its input from a Secret of the CronJob's name, which the commands above create from the directory step 3 wrote under the checkout's backup/ (or BACKUP_DIR): the restore's marker id, when it started, the archive name, and the revocations step 3 read out of the database before dropping it. Step 3 reads none when that database could not be read or is not the one the archive was dumped from (a fresh release's own database, for one), and says which; the Job then lists what can authenticate against the restored database instead of replaying. Without that Secret the Job exits 2 and says what is missing.

The Job ends with succeeded=1 whenever the steps ran, and its log says what they found: how many revocations were re-applied, the marker, and the anchor comparison. A rewind the anchors prove refuses chain writes until you acknowledge it at POST /v1/admin/vault/rewind/acknowledge, exactly as on Compose; with anchoring off, the log says nothing was compared. The log also names the restored database's instance id, and when the release's AGLEDGER_INSTANCE_ID differs (a restore into a release other than the one that took the backup, which generated its own), says to set the chart value instanceId to the database's. helm upgrade --reuse-values --set instanceId=<id> does that and starts the Deployments at the chart's replica counts, so step 5 follows it. When the log says the anchor comparison stopped before the end of the bucket, finish it once the Deployments are back up: kubectl rollout restart -n "$NS" deploy -l "app.kubernetes.io/instance=$REL,app.kubernetes.io/component=worker" restarts the worker, whose boot reconciliation walks the bucket with a longer budget than an API call holds, and GET /v1/admin/vault/rewind reports what it found. failed=1 means the steps did not run: fix what the log names and create another Job under a new name from the same Secret, which records the same restore once. postRestore.enabled is on by default; leave it on, because turning it back on mid-restore is a helm upgrade, which resets the Deployments' replica counts and starts the Server before these steps.

The backup PVC is ReadWriteOnce, so on most storage classes the reader pod has to land on the node the CronJob last ran on. If it stays Pending, kubectl get pod agl-backups -o wide and the PVC's events name the node it wants.

Step 5 restores the counts step 2 recorded rather than a fixed number, because the chart can run either Deployment under a HorizontalPodAutoscaler (hpa.enabled for the api, worker.autoscaling.enabled for the worker). Scaling to a count the autoscaler did not choose is overridden at its next pass, and scaling an install that ran three replicas to one leaves it downsized until somebody notices.

The refusals and the privilege check below apply, with one difference: a refusal on this path comes after step 2, so the Deployments are still at 0 and the release is serving nothing. The script says so and names the kubectl scale that step 5 runs; fix what it reports and re-run step 3, or run step 5 to put the release back as it was.

The chain verification that ends the restore does not run as written on this path: vault-verify.sh and vault-dump.sh are Compose wrappers that run docker compose run against the checkout's service definition, and vault-verify.sh also needs the release's signing key, which a cluster host does not hold. Run the shipped checker in the cluster instead, as a Job cut from the post-restore CronJob with dist/scripts/verify-vault.js as its command. The CronJob's pod is the api's, so the check runs with the release's ConfigMap, Secret, extraEnv (the anchoring settings included) and volumes, and carries the post-restore component label the bundled PostgreSQL's NetworkPolicy admits. restore.sh prints these commands; PR is the CronJob from step 4:

kubectl get cronjob -n "$NS" "$PR" -o json | jq '{apiVersion: "batch/v1", kind: "Job",
    metadata: {name: "agledger-vault-verify"}, spec: (.spec.jobTemplate.spec
    | .template.spec.containers[0].command |= (map(select(. != "--restore-input" and . != "/etc/agledger/post-restore"))
    | map(if . == "dist/scripts/post-restore.js" then "dist/scripts/verify-vault.js" else . end)))}' \
  | kubectl create -n "$NS" -f -
kubectl wait -n "$NS" --for=jsonpath='{.status.conditions[0].status}'=True --timeout=30m job/agledger-vault-verify
kubectl logs -n "$NS" job/agledger-vault-verify
kubectl delete job -n "$NS" agledger-vault-verify

The dump for the offline proof needs a pod that stays up long enough to copy the dump out; the audit runbook carries that pod spec, with the emptyDir the dump needs. Reaching the bundled PostgreSQL from outside the cluster means a port-forward to the <release>-agledger-chart-postgres Service, under the reachability rule above; an external database is reachable from wherever it already was.

What it refuses to do

Every one of these refusals happens before the script stops the application containers or drops anything. On Compose the Server is still serving when you read the message. On Kubernetes the recipe's step 2 has already scaled the Deployments to 0, so the script says so and names the kubectl scale that puts them back:

After the restore, before the Server starts

On an external database the restore does two things after the rows are back and before anything is started.

It returns the pgboss schema to the runtime role. pg_restore runs as the migrate role and so owns everything it writes, which is right for the application schema and wrong for pgboss: the Server installs that for itself, creates each queue's table as a partition of pgboss.job and tunes autovacuum on those partitions at boot, all of which require ownership. If the migrate role cannot perform the transfer (it has to be a member of the runtime role), the script says so and names the grant, then carries on to the runtime-role check below, which stops the run when the role cannot use the schema.

The restored records are not at risk; the job queue is. A runtime role that cannot use pgboss cannot boot (the API and the Worker refuse on permission denied for schema pgboss and print the remedy), and a process already running when the schema is taken from its role drops every job it enqueues. A webhook push enqueued in that window is lost: it is never delivered, retried or dead-lettered, and repairing the schema does not bring it back. GET /status reads Workers privilege_missing and GET /v1/admin/system-health names the missing privilege while it lasts. The events behind those pushes are still recorded, so the recovery is on the receiving side: each receiver reads GET /v1/events with since set to the start of the outage and walks to hasMore: false, as it would to reconcile any missed delivery. The Server has no call that pushes a past event again.

Then it runs the same runtime-role check install.sh and upgrade.sh run. If the role in DATABASE_URL cannot serve, the script stops there with the missing grant named: your data is restored and intact, the API and Worker are still stopped, and finishing is cd compose && docker compose up -d --wait from the install directory once the grant is in place.

Check your credentials before you declare recovery complete. The restore replaces api_keys with the copy in the backup, so the credential situation changed underneath you in three ways:

curl -s -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $KEY" "$AGLEDGER_API_URL/v1/auth/me"

A 401 there is not a data problem. The chain and every record are intact; you have no working credential. If API_KEY_SECRET_PREVIOUS above did not apply, or the old secret is gone, mint a fresh platform key. On Compose, from the install directory:

cd compose && docker compose exec -T -e NODE_OPTIONS= agledger-api /nodejs/bin/node \
  dist/scripts/generate-api-key.js platform 00000000-0000-0000-0000-000000000000 "recovery platform key"

On Kubernetes:

kubectl exec -n "$NS" deploy/<release>-agledger-chart-api -- env NODE_OPTIONS= /nodejs/bin/node \
  dist/scripts/generate-api-key.js platform 00000000-0000-0000-0000-000000000000 "recovery platform key"

It mints that one key under the API_KEY_SECRET the container runs with, prints it once, and touches nothing else. Do not reach for init.js here: it is the first-run wizard, and on a live install it also generates a federation signing key and prints a fresh .env block, none of which belongs to this install.

After it returns, wait for the readiness gate before trusting anything:

curl -s "$AGLEDGER_API_URL/readyz"
{"status":"ready","version":"2.0.0","timestamp":"..."}

A ready Server is serving. It is not yet a verified Server.

Every restore leaves a marker, and an anchored install checks itself against the bucket

restore.sh writes one row into chain_restore_markers, and the next boot turns it into a signed RESTORE_EPOCH entry on the platform-ops chain plus a chain.restore_detected row on the SIEM stream, whether or not anchoring is configured and whether or not anything was lost. The marker alone never stops writes.

The restore also adopts the restored database's instance id, the prefix every external anchor this install wrote is under, and writes it back into compose/.env. A rebuilt host arrives with a fresh .env, and a Server whose AGLEDGER_INSTANCE_ID disagrees with the id an operator set in the database refuses to boot and names both values. On the cluster path, set the chart's AGLEDGER_INSTANCE_ID to the restored value or leave it unset.

With VAULT_ANCHOR_ENABLED=true the script also runs POST /v1/admin/vault/anchors/reconcile against the restored database, and the Server repeats it at boot and daily. It walks this Server's anchor prefix in the bucket and compares the highest anchored position per record against the restored chain head. rewound means the bucket holds a position past the database, which is a restore that lost entries an anchor points at. From then on records, completions, verdicts, schema registrations and SCITT registrations answer 409 with reason: CHAIN_REWIND_DETECTED and a recoveryHint naming POST /v1/admin/vault/rewind/acknowledge, and GET /v1/admin/vault/rewind carries the evidence. Acknowledging resumes writes and appends a second RESTORE_EPOCH entry carrying that evidence, the ranges it covers (coveredFindings) and your note, so the two histories stay tellable apart.

The acknowledgement sticks while it counts. The anchors are retention-locked and a lost record does not come back, so every later reconciliation (at worker boot, daily, or on request) finds the same evidence, reports it under acknowledged with status: acknowledged, and does not stop writes again. Whether it counts is judged by the keys the process running the reconciliation trusts, which start from its own key:

Acknowledging repairs nothing in the bucket. A record chain the restore rewound takes the lost positions again with new content; when the checkpoint sweep tries to anchor one of those positions, the create-only write finds the lost history's checkpoint there and reports the write as superseded, which stops nothing, and a checkpoint re-created over the same chain tip is a replay. Writes stop again only for evidence the acknowledgement did not cover: a position anchored past the acknowledged range, another record, an anchor written after the acknowledgement, a later restore to a backup taken before it, or an acknowledgement that stops counting (one copied back into the database after a later restore, or signed by a key you have since retired with force or distrusted; the anchoring runbook has the rules).

Nothing on this path detects a loss of entries written since the last checkpoint sweep, and an install with no anchoring configured gets the marker and nothing else; the backup runbook says what to reconcile against instead.

Verify the restored chain: the step that ends the restore

Prove the chain came back intact before you declare recovery complete. Run the connected check, then the authoritative offline verification. Both scripts ship in the agledger-ai/install repository and run the tools that already live inside the Server image, with no source checkout, Node.js, or pnpm on the host.

The in-database check runs the Server's own walk over every chain in audit_vault, the same one POST /v1/admin/vault/scan runs: hash links and positions, the signed chain claim in each envelope, each entry's signature against the key registry, each key's published window, and the payload and certificate bindings.

./scripts/vault-verify.sh
Verifying 3 chain(s): 2 record, 1 schema...

[PASS] record 01a07322-9f82-74bc-b208-2af87e683d3e (1 entries)
[PASS] record 01a07322-9fee-7de1-a7e8-e781bc8d3c2b (1 entries)
[PASS] schema:01a07321-540d-7eee-b520-af2f8b6263f8 (4 entries)

3 chain(s) checked, 0 error(s), 0 unverifiable
  record chains: 2
  schema chains: 1

audit_vault holds two kinds of chain and the run covers both. Record chains are keyed by record id, including the platform-operations chain at the all-zero sentinel id. Schema chains hold the SCHEMA_REGISTERED, SCHEMA_IMPORTED and SCHEMA_DIGEST_MISMATCH entries, which carry no record id and are chained independently per Org, so they print as schema:<orgId>. Confirm the counts match what you expect (on a fresh install, the schema chain's entries are the example contracts it seeds): a restore that brought back records but not the schema registrations shows up as a schema chain that is missing or short, and that is the reading the breakdown lines exist for.

Then the authoritative proof: produce a database-independent dump and verify it offline with @agledger/verify 2.0.0, the first release that reads a 2.0 Server's dumps, pinned on the vault key pins you hold, following the off-box verification in the audit runbook:

./scripts/vault-dump.sh ./dump
npx -y @agledger/verify@2.0.0 ./dump --trust-anchor <pin>

Pass each pin you hold as its own --trust-anchor (the installer printed the first key's). Without one, a clean dump passes as [VERIFIED, NOT ANCHORED], which is not a trusted verdict: the keys were taken on the restored database's word. What matters here is that the pinned run comes back clean across entries signed by the active key and any rotated-out retired key, because retired keys travel in the dump's vault_signing_keys.ndjson, the signed key statements that link them to your pin travel in vault_key_statements.ndjson, and the verifier resolves whichever key signed each entry. The restore is complete when that run passes, not when the containers are healthy.

When verification reports a failure

A failure is information, not a dead end. The class tells you which kind of problem you have.

A chain with no entries at all is reported as [EMPTY], not as a pass, and the run exits non-zero. Nothing about that chain was proven: every creation path writes its first chain entry in the same transaction as the row it covers, so a chain with zero entries is one whose entries are gone. That is the partial-restore signature, and it is why the check enumerates from records and vault_checkpoints as well as from audit_vault. To re-run against one chain, pass the key exactly as the report printed it: ./scripts/vault-verify.sh --chain-key schema:<orgId>, or a record id.

vault-verify.sh stops each chain at its first break and prints one line under it, for example - [chain_broken_at] pos=3: Chain integrity failed at position 3 (chain_broken_at). chain_broken_at covers every structural break: a missing position, a broken link, a first entry that does not start at genesis, or a stored hash that does not match its envelope. A signature or key finding prints its own reason instead, such as signature_invalid or key_expired. The offline verifier that reads the NDJSON dump splits the structural break into four classes, and that split is what tells you which kind of problem you have:

Offline classWhat it meansRestore went wrong, or real problem?
CHAIN_POSITION_GAPA chain position is missingUsually a partial or interrupted restore: re-restore from a complete backup
CHAIN_LINK_BROKENAn entry's previous_hash does not match the prior entryPartial restore, or a backup taken mid-write: re-restore
CHAIN_GENESIS_INVALIDThe first entry does not start at genesisTruncated restore: re-restore
CHAIN_HASH_MISMATCHA stored hash does not match sha256(cose_sign1)If the backup itself verifies clean, the restore corrupted bytes: re-restore. If the backup also fails here, the backup faithfully captured a real tamper you must investigate

The discriminator: verify the backup (its NDJSON dump) independently. If the backup verifies and the restore does not, the restore is at fault: repeat it. If the backup itself fails, recovery will not paper over it; you have a real integrity problem from before the backup was taken.

There is one more class, and it is not about your restore at all. The offline verifier calls it CHAIN_UNSUPPORTED_ALGORITHM and vault-verify.sh and the Server's own scan call it unsupported_algorithm; they are the same finding, and it means the restore succeeded onto a host whose crypto provider cannot compute the algorithm that signed the entries. The chain is intact and this host cannot check it. FIPS 140 hosts covers the case that produces it in practice and what to do about it.

Signing-key recovery

The vault_signing_keys registry and its signed key statements are part of the database dump, so they return with the restore, and the offline run above covers the active key plus any rotated-out retired keys. Records signed under a retired key still verify, because the engine resolves historical keys from the registry.

Restoring the registry is not the same as restoring the ability to sign. New records need the private key in VAULT_SIGNING_KEY, or the ARN of a KMS key in VAULT_SIGNING_KEY_KMS_ARN, and neither is in the backup. Restore the database and set it. A key the registry already trusts signs at once. A new key is registered only when something outside the database vouches for it: VAULT_SIGNING_KEY_PREVIOUS set to a key the registry trusts (it signs a succession statement with the new key), or VAULT_TRUST_ANCHORS naming the pins of the keys whose history you vouch for, after which the new key registers under a fresh genesis and you give auditors its pin. Without either the process reports signingKey.gate: "unanchored" and signs nothing. Registration retires nothing; retirement is a separate call, POST /v1/admin/vault/signing-keys/{keyId}/retire, made from a process on the new key (see signing-key rotation). Every process reports the key it holds as signingKey.keyId on GET /health, which is how you confirm a restored host is signing under the key you meant.

A restore after a rotation, or onto a fresh host. A backup taken before the current key was staged holds a registry without it, and so does any restore onto a host with a key of its own. Put VAULT_SIGNING_KEY_PREVIOUS (the key the backup's registry holds), or VAULT_TRUST_ANCHORS as above, in compose/.env (or secrets.vaultSigningKeyPrevious in the chart) before running restore.sh. The post-restore step registers the key the way the Server's boot does and only then reconciles the anchors, so with the predecessor in place it can tell whether a rewind acknowledged before the backup still covers what the bucket holds. Without it the key is unanchored, and that step does not reconcile: it prints the gate and its fix, and restore.sh reports that no comparison was completed. The worker restore.sh then starts reconciles at its boot, and under a key that is still not usable it fails closed as above, recording the acknowledged rewind open again, so chain writes stay refused until you set the predecessor, restart, and acknowledge again.

API_KEY_SECRET_PREVIOUS is different: no restore step reads it, only the Server's authentication does, so set it before restore.sh and the stack it starts accepts restored keys at once, or set it afterwards and restart, as the note on restored keys above describes.

The key registers at boot even though a rewind refuses chain writes, because the signing-key registry's own entries are the one kind of write besides the restore marker that the refusal lets through: the acknowledgement is signed by a registered key, so holding the registration back would leave no way to acknowledge. Confirm it before acknowledging:

curl -s "$AGLEDGER_API_URL/health" | jq '.signingKey'

gate should read usable and keyId the key you set. Then read the evidence at GET /v1/admin/vault/rewind and acknowledge at POST /v1/admin/vault/rewind/acknowledge; the RESTORE_EPOCH is signed under the current key. If gate reads unregistered instead, the registration failed for another reason and retries on its own; the log line Could not register VAULT_SIGNING_KEY in vault_signing_keys carries the error. unanchored means nothing links the key to the restored registry: set VAULT_SIGNING_KEY_PREVIOUS or VAULT_TRUST_ANCHORS and restart. GET /v1/admin/vault/rewind names the gate in its nextSteps when the process answering it cannot sign the acknowledgement.

What recovery cannot do

You cannot rewrite history to "fix" a chain. A record that was lost before the backup shows up as a chain gap (chain_broken_at from vault-verify.sh, CHAIN_POSITION_GAP from the offline verifier), and that gap is the honest record of what happened. There is nothing to restore it from and nothing legitimate to paper over it with. Recovery brings back what the backup holds and proves it; it does not manufacture entries.

Multi-Server recovery

Each Server owns its own database and restores its own slice independently; there is no shared store to coordinate. Cross-Server delegation chains re-link by signature reference, not by restoring a common database: once each Server is restored and verifying, the links across them resolve through the published public keys. Restore and verify Server by Server.

The peering itself does not re-form on its own. A peer is an ordinary row in federation_peers, nothing reconciles or rediscovers it, and the boot check only asks whether an active peer exists. A peer handshaken after the backup was taken is gone from the restored Server, and the operators on both sides run the handshake again (POST /federation/v1/admin/peering-tokens, then POST /federation/v1/peer) exactly as at first bootstrap. A peer that existed at backup time and was revoked after it comes back as it was, and a restore over the database the backup was dumped from replays that revocation with the others (see the credential check above).