Backup

A complete, restorable backup of a Server is three separate things. Two of them are not in the database dump, and the operator decides where each lives.

WhatWhere it livesIn the DB dump?
The PostgreSQL database - records, the audit_vault chain, the signing-key registry, provisioning statePostgresYes
The customer-held encryption keys (encrypted mode)Wherever you keep them - the Server never stores themNo
The config-as-code provisioning filesYour PROVISIONING_CONFIG_PATH directory / source controlNo

The split is the point. A database dump without the encryption keys is opaque-but-safe - an attacker who steals it cannot read encrypted payloads. The encryption keys without the database are useless. Back up each on its own schedule, to its own place.

This runbook pairs with the recovery runbook - a backup you have never restored is a hope, not a backup.

Back up the database

scripts/backup.sh - from the agledger-ai/install repository - is the supported path. It writes a timestamped, compressed tarball and prunes to the last --keep N (default 7):

./scripts/backup.sh            # keep last 7
./scripts/backup.sh --keep 30
BACKUP_DIR=/mnt/backups ./scripts/backup.sh

Archives land in backup/ inside the checkout unless BACKUP_DIR names somewhere else, and --keep takes 1 or more (--help prints the options). Each tarball contains the database dump, the key-registry export in almost every case, and a metadata file. Under the hood the script runs pg_dump in custom format against the bundled postgres (or your external DATABASE_URL for Aurora/RDS):

db.dump                  # pg_dump -Fc, the full database, custom format
vault-public-keys.csv    # the signing-key registry, PUBLIC keys only
backup-metadata          # created_at, database_token, wal_lsn, compose_project, agledger_version, agledger_image_pin

backup-metadata is a short key=value file recording the release the archive was taken from. restore.sh reads it to return the install to that release before starting it, and --keep-version restores the data onto the release the install is on now instead. --keep-version is refused for a 1.x archive on a 2.x install, whose migration would refuse the database. Each archive is named backup-<compose project>-<timestamp>.tar.gz, and --keep counts and removes only archives carrying this install's project name, so several installs can share one BACKUP_DIR.

[2026-09-22T02:00:04Z] PostgreSQL backup complete (732K), archive header verified.

The vault-public-keys.csv is public-key metadata only - fingerprints, algorithms, status, activation and retirement dates. Private signing keys are never written to the database and never appear in a backup:

key_id,public_key,algorithm,status,activated_at,retired_at
c4dd3e20388b594d,MCowBQYDK2VwAyEAo95XH8DQ...,Ed25519,active,2026-06-09 15:54:50.700979+00,

It is there so that after a restore you can confirm the registry came back intact (see the recovery runbook). A fresh install carries the one active key above; after a key rotation the CSV lists the active and retired keys, because retired keys still verify records signed before the rotation.

For an external database (Aurora, RDS, Cloud SQL), back up with your provider's snapshot mechanism instead - backup.sh detects an external DATABASE_URL and uses pg_dump directly. The three-way split above is unchanged.

pg_dump refuses to dump a server newer than itself, and distribution packages lag: Ubuntu 24.04 ships client 16 against the PostgreSQL 18 this product is validated on. backup.sh compares the two and falls back to a matching client in a container when the host's is older, so the case that needs your attention is a host with neither an adequate client nor docker. It names the PostgreSQL apt repository lines when it hits that.

A backup that cannot be read is not written. Before reporting success, backup.sh checks that db.dump begins with the archive header pg_restore expects. If it does not, the whole backup directory is removed and the script exits non-zero naming what it found, rather than leaving a plausible tarball whose only symptom appears during a restore, after the stack is down and the database is dropped. The key-registry export gets the same treatment against its header row; a bad one is dropped and the run says so, leaving a one-file tarball with the database dump intact.

On Kubernetes

backup.sh drives docker compose and does not run against a cluster. The chart carries the same job as a CronJob. Keep its settings in the values file you pass on every upgrade:

# values.yaml
backup:
  cronJob:
    enabled: true
    schedule: "0 2 * * *"   # UTC
  persistence:
    enabled: true
    storageClassName: gp3   # a class from `kubectl get storageclass`
  keep: 14
helm upgrade --install agledger oci://registry-1.docker.io/agledger/agledger-chart \
  --namespace agledger --values values.yaml

--reuse-values is the wrong shortcut here. It re-coalesces the previous release's chart defaults, which pins backup.image to whatever the chart you installed with shipped.

Name the class unless your cluster has a default one. Without storageClassName, the backup PersistentVolumeClaim takes the cluster's default StorageClass, and a cluster with none (stock EKS 1.30 and later) leaves the claim Pending. The install still succeeds, because nothing waits on the claim, and every backup Job then sits Pending behind it, so no backup ever runs. kubectl get storageclass marks the default with (default); with none marked, set backup.persistence.storageClassName. helm-install.sh refuses to install when the cluster has classes but none is the default, and names the setting. For a claim already Pending, delete it (it holds nothing), name the class, and run the upgrade again: a claim's class cannot be changed in place.

To take a backup now rather than wait for the schedule, start a Job from the CronJob. The chart names it <release>-agledger-chart-backup, so for the release agledger:

kubectl create job --from=cronjob/agledger-agledger-chart-backup agledger-backup-manual -n agledger
kubectl wait --for=condition=complete job/agledger-backup-manual -n agledger --timeout=30m
kubectl logs job/agledger-backup-manual -n agledger

A Job name can be used once, so give each manual run its own. kubectl get cronjob -n agledger shows the name if your release is called something else.

The CronJob runs the same pg_dump -Fc, exports the same public-key registry, applies the same archive-header check before it keeps anything, and writes a backup-<timestamp>.tar.gz in the shape restore.sh reads onto a PersistentVolumeClaim. The PVC carries helm.sh/resource-policy: keep, so helm uninstall leaves the backups behind. One difference from the Compose archives: the chart's backup-metadata holds only created_at (the instant the dump started), database_token (which database it was dumped from) and wal_lsn (that database's WAL position once the dump finished), where Compose's also records the install version, image pin and project name. An archive whose metadata carries no version, whether it came from the chart or from an older Compose run, is treated as carrying no version record, and the restore leaves the install on the release it is already running rather than moving it. A host restoring into a cluster has no Compose install to read the release from, so the Kubernetes recipe in the recovery runbook passes the Helm release's app version to restore.sh with --version.

The restore itself is restore.sh, run from a host that can reach the database, because a restore has questions a template cannot answer: which archive, whether the connection carries the superuser rights the audit event trigger needs, and whether the runtime role still holds its grants afterwards. The steps that follow it (the revocation replay and the external-anchor comparison) run as the chart's post-restore Job, with the release's own configuration. The recovery runbook has the kubectl recipe for both.

backup.image has to carry a pg_dump no older than your server; it defaults to the bundled PostgreSQL image. backup.s3 uploads the same tarball to an S3-compatible bucket instead of a PVC, and needs an image carrying both pg_dump and the aws CLI, which no public image provides, so that path means building one. backup.keep does not apply there: retention is the bucket's lifecycle policy, and a Job that expires objects is a Job holding delete rights on every backup you have.

backup.preUpgrade.enabled=true takes a backup before the migration Job on every helm upgrade, and a failure aborts the upgrade with nothing migrated, which is the point. It needs a destination that already exists, so the upgrade that first enables it alongside backup.persistence.enabled takes no backup: helm finishes every pre-upgrade hook before it creates an ordinary resource, and the PersistentVolumeClaim is one. Every upgrade after that one backs up first. To cover the first one too, point backup.persistence.existingClaim at a claim you already have, or use backup.s3.

The chart's bundled PostgreSQL is a Deployment with a PVC and stock configuration. It has no point-in-time recovery and no backup of its own, so on that path this CronJob is the whole plan and its schedule is your recovery point. On an external database the provider's point-in-time recovery is the better answer and this is a portable second copy.

Your recovery point is the last snapshot

backup.sh takes snapshots. Nothing that ships sets archive_mode or an archive_command: not the Compose files, not the chart, not the scripts, and the bundled postgres:18-alpine runs stock configuration. There is no continuous archive behind the tarballs. Whatever was notarized since the last run is in no copy the product holds. Pick the schedule from the recovery point you need, not from the default.

Point-in-time recovery is a property of your database, and the Server neither provides nor prevents it:

Recovering to a point in time hands you a chain that ends earlier, and a chain truncated from the end still hash-links cleanly, so the walk alone cannot tell you entries are missing. Only evidence held outside the database can say where the chain used to end, and vault_checkpoints is not that: it is an ordinary table in the same database, so it rolls back with the chain and the restored checkpoint agrees with the restored truncated chain. External anchors are the exception. With VAULT_ANCHOR_ENABLED=true each checkpoint is also written to storage the database recovery cannot reach, one key per chain position for the life of the install (vault-anchors/<AGLEDGER_INSTANCE_ID>/<recordId>/<position>.json). The Compose installer and the chart each generate a UUID for the instance id on a first install; vault-anchors/default/... is where a Server anchors only when nothing set one. The id is recorded in the database on first boot, so a restore brings the prefix back with the data, and GET /v1/admin/ops-summary reports the id the database carries under vault.anchoring.instanceId. That storage is object-locked against deletion only when VAULT_ANCHOR_OBJECT_LOCK=true and the bucket was itself created with Object Lock enabled; on a bucket without it, anchor writes still succeed and are still create-only, but nothing stops a later delete.

POST /v1/admin/vault/anchors/verify is not the call that catches a rewind. It checks the anchors of checkpoints the restored database still holds, and never reads an anchor for a position the restore removed. POST /v1/admin/vault/anchors/reconcile is the bucket-to-database comparison: it walks the bucket itself and reports rewound for an anchored position past this database's restored chain head, or missing_locally for a key under this Server's prefix with no matching checkpoint, alongside a posture note on whether the store honours create-only writes and reports Object Lock enabled. A rewound finding refuses every chain write with 409 CHAIN_REWIND_DETECTED until an operator acknowledges it at POST /v1/admin/vault/rewind/acknowledge, so the restored Server does not silently take writes past a position the anchors prove it lost. restore.sh runs that reconciliation for you once anchoring is configured, before it starts the Server, and on Kubernetes the chart's post-restore Job runs it.

Reconcile before the restored Server takes writes: the next entry on a truncated chain takes a position the lost history already signed and delivered. Without anchors, reconcile against evidence outside the chain instead: your SIEM stream of system_audit_log, delivered webhooks, or a counterparty Server's slice of a federated chain. Verify after any restore; see the recovery runbook.

A revocation made after the backup is a row in the database the restore replaces. restore.sh reads every revocation out of that database before it drops it and replays them onto the restored rows once the schema is current, so an API key, ephemeral certificate, federation peer or trusted issuer revoked after the backup comes back revoked rather than live again. On Kubernetes the chart's post-restore Job does the replay, from the export restore.sh hands it.

That covers only a restore that replaces the very database the backup was dumped from. The archive records which one that is (database_token: the PostgreSQL cluster's system identifier and the database's OID), and restore.sh reads revocations out of no other: not a fresh host's own database, not a standby restored from an earlier backup, and not the database a failed attempt at the same restore left behind, since each of those is a new database even when its rows came from the same install. A physical copy of the database (a provider snapshot or point-in-time restore, a clone, a recovery cluster) keeps the same database_token, so the archive also records wal_lsn: a database still behind it is a copy taken before the backup and is not read either, while a copy taken after the backup still matches, and the run says so. A restore onto a host whose previous database is gone has nothing to read, and an archive that records no database_token or wal_lsn (one taken before backups recorded them) cannot show which database it came from. In each of those cases nothing is replayed, and the post-restore step instead lists every credential the restored database still trusts, for you to compare against your own record of what was revoked.

Back up the chain off-box, too

Keep a second, database-independent copy of the chain: the NDJSON dump the offline verifier consumes. This is not a replacement for the database backup - it is the artifact you hand an auditor, and the copy that proves intact without a running Server.

./scripts/vault-dump.sh ./chain-backup

The shipped scripts/vault-dump.sh runs the dump tool inside the Server image, so it needs no source checkout or pnpm, only a reachable database. It prints the dumped row counts and the output directory:

{
  "outDir": "/dump",
  "orgId": null,
  "counts": { "audit_vault": 10, "vault_checkpoints": 0, "vault_signing_keys": 1, "vault_key_statements": 1, "org_admin_reads": 0, "org_admin_reads_checkpoints": 0 }
}

It writes one NDJSON file per entry in counts, including vault_signing_keys.ndjson and vault_key_statements.ndjson: the public-key registry and the signed statements that admit each key travel with the dump, so the chain verifies offline with no further inputs. The audit runbook covers the file contents and how to verify them.

Where it goes is the part that belongs here. Keep the dump alongside the db.dump tarball, not instead of it, and treat it as a separate destination rather than a second file in the same place: the database backup is what you restore a running Server from, and the NDJSON dump is what proves the chain to someone who does not trust your Server, including after that Server is gone. A copy that only exists next to the database it came from proves nothing the database could not have been rewritten to say.

Point @agledger/verify at the directory, and pin both the version and the key. 2.0.0 is the first release that reads a dump from a 2.x Server; an older one fails it once an admin has read a record.

npx -y @agledger/verify@2.0.0 ./chain-backup --trust-anchor <pin>

<pin> is the "Vault signing key pin" install.sh printed, sha256: and the digest of the signing key's public half (node dist/scripts/signing-key-digest.js derives it again from the key). Without --trust-anchor the verdict is VERIFIED, NOT ANCHORED: nothing failed, but every key was taken on the dump's word.

The installer's pin is enough only while every key change on the install was a routine rotation, which links each new key to the one before it. A key retired with force, or a recovery that staged a key under VAULT_TRUST_ANCHORS, cuts that link, and a run with the installer's pin alone then fails. Pass every pin the operator holds instead, each as its own --trust-anchor: the current key's (anchoredFrom on GET /v1/verification-keys) and every digest in VAULT_TRUST_ANCHORS. Pass each VAULT_DISTRUSTED_KEYS entry as --distrusted-key, since the dump does not carry them. The key-compromise runbook lists the pins each recovery path leaves.

Any RFC 9052 library also verifies the per-row cose_sign1 bytes, when you want no AGLedger package in the loop. @agledger/cli's verify subcommand is a different tool: it checks one record's audit export (GET /v1/records/{id}/audit-export) and reports EISDIR if you point it at a dump directory.

What a backup does not contain