Backup
A complete, restorable backup of a Server is three separate things. Two of them are not in the database dump, and the operator decides where each lives.
| What | Where it lives | In the DB dump? |
|---|---|---|
The PostgreSQL database - records, the audit_vault chain, the signing-key registry, provisioning state | Postgres | Yes |
| The customer-held encryption keys (encrypted mode) | Wherever you keep them - the Server never stores them | No |
| The config-as-code provisioning files | Your PROVISIONING_CONFIG_PATH directory / source control | No |
The split is the point. A database dump without the encryption keys is opaque-but-safe - an attacker who steals it cannot read encrypted payloads. The encryption keys without the database are useless. Back up each on its own schedule, to its own place.
This runbook pairs with the recovery runbook - a backup you have never restored is a hope, not a backup.
Back up the database
scripts/backup.sh - from the agledger-ai/install
repository - is the supported path. It writes a timestamped, compressed tarball and prunes to the
last --keep N (default 7):
./scripts/backup.sh # keep last 7
./scripts/backup.sh --keep 30
BACKUP_DIR=/mnt/backups ./scripts/backup.sh
Archives land in backup/ inside the checkout unless BACKUP_DIR names somewhere else, and
--keep takes 1 or more (--help prints the options). Each tarball contains the database dump, the
key-registry export in almost every case, and a metadata file. Under the hood the script runs pg_dump in custom format against the bundled
postgres (or your external DATABASE_URL for Aurora/RDS):
db.dump # pg_dump -Fc, the full database, custom format
vault-public-keys.csv # the signing-key registry, PUBLIC keys only
backup-metadata # created_at, database_token, wal_lsn, compose_project, agledger_version, agledger_image_pin
backup-metadata is a short key=value file recording the release the archive was taken from.
restore.sh reads it to return the install to that release before starting it, and --keep-version
restores the data onto the release the install is on now instead. --keep-version is refused for a
1.x archive on a 2.x install, whose migration would refuse the database. Each archive is named
backup-<compose project>-<timestamp>.tar.gz, and --keep counts and removes only archives
carrying this install's project name, so several installs can share one BACKUP_DIR.
[2026-09-22T02:00:04Z] PostgreSQL backup complete (732K), archive header verified.
The vault-public-keys.csv is public-key metadata only - fingerprints, algorithms, status,
activation and retirement dates. Private signing keys are never written to the database and never
appear in a backup:
key_id,public_key,algorithm,status,activated_at,retired_at
c4dd3e20388b594d,MCowBQYDK2VwAyEAo95XH8DQ...,Ed25519,active,2026-06-09 15:54:50.700979+00,
It is there so that after a restore you can confirm the registry came back intact (see the recovery runbook). A fresh install carries the one active key above; after a key rotation the CSV lists the active and retired keys, because retired keys still verify records signed before the rotation.
For an external database (Aurora, RDS, Cloud SQL), back up with your provider's snapshot mechanism
instead - backup.sh detects an external DATABASE_URL and uses pg_dump directly. The three-way
split above is unchanged.
pg_dump refuses to dump a server newer than itself, and distribution packages lag: Ubuntu 24.04
ships client 16 against the PostgreSQL 18 this product is validated on. backup.sh compares the two
and falls back to a matching client in a container when the host's is older, so the case that needs
your attention is a host with neither an adequate client nor docker. It names the PostgreSQL apt
repository lines when it hits that.
A backup that cannot be read is not written. Before reporting success, backup.sh checks that
db.dump begins with the archive header pg_restore expects. If it does not, the whole backup
directory is removed and the script exits non-zero naming what it found, rather than leaving a
plausible tarball whose only symptom appears during a restore, after the stack is down and the
database is dropped. The key-registry export gets the same treatment against its header row; a bad
one is dropped and the run says so, leaving a one-file tarball with the database dump intact.
On Kubernetes
backup.sh drives docker compose and does not run against a cluster. The chart carries the same
job as a CronJob. Keep its settings in the values file you pass on every upgrade:
# values.yaml
backup:
cronJob:
enabled: true
schedule: "0 2 * * *" # UTC
persistence:
enabled: true
storageClassName: gp3 # a class from `kubectl get storageclass`
keep: 14
helm upgrade --install agledger oci://registry-1.docker.io/agledger/agledger-chart \
--namespace agledger --values values.yaml
--reuse-values is the wrong shortcut here. It re-coalesces the previous release's chart defaults,
which pins backup.image to whatever the chart you installed with shipped.
Name the class unless your cluster has a default one. Without storageClassName, the backup
PersistentVolumeClaim takes the cluster's default StorageClass, and a cluster with none (stock EKS
1.30 and later) leaves the claim Pending. The install still succeeds, because nothing waits on the
claim, and every backup Job then sits Pending behind it, so no backup ever runs.
kubectl get storageclass marks the default with (default); with none marked, set
backup.persistence.storageClassName. helm-install.sh refuses to install when the cluster has
classes but none is the default, and names the setting. For a claim already Pending, delete it (it
holds nothing), name the class, and run the upgrade again: a claim's class cannot be changed in place.
To take a backup now rather than wait for the schedule, start a Job from the CronJob. The chart
names it <release>-agledger-chart-backup, so for the release agledger:
kubectl create job --from=cronjob/agledger-agledger-chart-backup agledger-backup-manual -n agledger
kubectl wait --for=condition=complete job/agledger-backup-manual -n agledger --timeout=30m
kubectl logs job/agledger-backup-manual -n agledger
A Job name can be used once, so give each manual run its own. kubectl get cronjob -n agledger
shows the name if your release is called something else.
The CronJob runs the same pg_dump -Fc, exports the same public-key registry, applies the same
archive-header check before it keeps anything, and writes a backup-<timestamp>.tar.gz in the shape
restore.sh reads onto a PersistentVolumeClaim. The PVC carries helm.sh/resource-policy: keep, so
helm uninstall leaves the backups behind. One difference from the Compose archives: the chart's
backup-metadata holds only created_at (the instant the dump started), database_token (which
database it was dumped from) and wal_lsn (that database's WAL position once the dump finished),
where Compose's also records the install version, image pin and project name. An archive whose
metadata carries no version, whether it came from the chart or from an older Compose run, is treated
as carrying no version record, and the restore leaves the install on the release it is already
running rather than moving it. A host restoring into a cluster has no Compose install to read the
release from, so the Kubernetes recipe in the recovery runbook passes
the Helm release's app version to restore.sh with --version.
The restore itself is restore.sh, run from a host that can reach the database, because a restore
has questions a template cannot answer: which archive, whether the connection carries the superuser
rights the audit event trigger needs, and whether the runtime role still holds its grants
afterwards. The steps that follow it (the revocation replay and the external-anchor comparison) run
as the chart's post-restore Job, with the release's own configuration. The recovery runbook has the
kubectl recipe for both.
backup.image has to carry a pg_dump no older than your server; it defaults to the bundled
PostgreSQL image. backup.s3 uploads the same tarball to an S3-compatible bucket instead of a PVC,
and needs an image carrying both pg_dump and the aws CLI, which no public image provides, so that
path means building one. backup.keep does not apply there: retention is the bucket's lifecycle
policy, and a Job that expires objects is a Job holding delete rights on every backup you have.
backup.preUpgrade.enabled=true takes a backup before the migration Job on every helm upgrade, and
a failure aborts the upgrade with nothing migrated, which is the point. It needs a destination that
already exists, so the upgrade that first enables it alongside backup.persistence.enabled takes no
backup: helm finishes every pre-upgrade hook before it creates an ordinary resource, and the
PersistentVolumeClaim is one. Every upgrade after that one backs up first. To cover the first one
too, point backup.persistence.existingClaim at a claim you already have, or use backup.s3.
The chart's bundled PostgreSQL is a Deployment with a PVC and stock configuration. It has no point-in-time recovery and no backup of its own, so on that path this CronJob is the whole plan and its schedule is your recovery point. On an external database the provider's point-in-time recovery is the better answer and this is a portable second copy.
Your recovery point is the last snapshot
backup.sh takes snapshots. Nothing that ships sets archive_mode or an archive_command: not
the Compose files, not the chart, not the scripts, and the bundled postgres:18-alpine runs stock
configuration. There is no continuous archive behind the tarballs. Whatever was notarized since the
last run is in no copy the product holds. Pick the schedule from the recovery point you need, not
from the default.
Point-in-time recovery is a property of your database, and the Server neither provides nor prevents it:
- Bundled PostgreSQL. Configure continuous WAL archiving on that instance the way you would for
any PostgreSQL you run, and keep it alongside the tarballs rather than instead of them.
restore.shreads a tarball, and that is the path the recovery runbook validates. - External or managed PostgreSQL (Aurora, RDS, Cloud SQL, Azure). Use the provider's
point-in-time recovery. It is the shorter route to a sub-minute recovery point, and it is
independent of
backup.sh, which detects an externalDATABASE_URLand runspg_dumpagainst it either way.
Recovering to a point in time hands you a chain that ends earlier, and a chain truncated from the
end still hash-links cleanly, so the walk alone cannot tell you entries are missing. Only evidence
held outside the database can say where the chain used to end, and vault_checkpoints is not that:
it is an ordinary table in the same database, so it rolls back with the chain and the restored
checkpoint agrees with the restored truncated chain. External anchors are the exception. With
VAULT_ANCHOR_ENABLED=true each checkpoint is also written to storage the database recovery cannot
reach, one key per chain position for the life of the install
(vault-anchors/<AGLEDGER_INSTANCE_ID>/<recordId>/<position>.json). The Compose installer and the
chart each generate a UUID for the instance id on a first install; vault-anchors/default/... is
where a Server anchors only when nothing set one. The id is recorded in the database on first boot,
so a restore brings the prefix back with the data, and GET /v1/admin/ops-summary reports the id the
database carries under vault.anchoring.instanceId. That storage is object-locked against
deletion only when VAULT_ANCHOR_OBJECT_LOCK=true and the bucket was itself created with Object
Lock enabled; on a bucket without it, anchor writes still succeed and are still create-only, but
nothing stops a later delete.
POST /v1/admin/vault/anchors/verify is not the call that catches a rewind. It checks the
anchors of checkpoints the restored database still holds, and never reads an anchor for a position
the restore removed. POST /v1/admin/vault/anchors/reconcile is the bucket-to-database comparison:
it walks the bucket itself and reports rewound for an anchored position past this database's
restored chain head, or missing_locally for a key under this Server's prefix with no matching
checkpoint, alongside a posture note on whether the store honours create-only writes and reports
Object Lock enabled. A rewound finding refuses every chain write with 409 CHAIN_REWIND_DETECTED
until an operator acknowledges it at POST /v1/admin/vault/rewind/acknowledge, so the restored
Server does not silently take writes past a position the anchors prove it lost. restore.sh runs
that reconciliation for you once anchoring is configured, before it starts the Server, and on
Kubernetes the chart's post-restore Job runs it.
Reconcile before the restored Server takes writes: the next entry on a truncated chain takes a
position the lost history already signed and delivered. Without anchors, reconcile against evidence
outside the chain instead: your SIEM stream of system_audit_log, delivered webhooks, or a
counterparty Server's slice of a federated chain. Verify after any restore; see the
recovery runbook.
A revocation made after the backup is a row in the database the restore replaces. restore.sh
reads every revocation out of that database before it drops it and replays them onto the restored
rows once the schema is current, so an API key, ephemeral certificate, federation peer or trusted
issuer revoked after the backup comes back revoked rather than live again. On Kubernetes the
chart's post-restore Job does the replay, from the export restore.sh hands it.
That covers only a restore that replaces the very database the backup was dumped from. The archive
records which one that is (database_token: the PostgreSQL cluster's system identifier and the
database's OID), and restore.sh reads revocations out of no other: not a fresh host's own database,
not a standby restored from an earlier backup, and not the database a failed attempt at the same
restore left behind, since each of those is a new database even when its rows came from the same
install. A physical copy of the database (a provider snapshot or point-in-time restore, a clone, a
recovery cluster) keeps the same database_token, so the archive also records wal_lsn: a database
still behind it is a copy taken before the backup and is not read either, while a copy taken after
the backup still matches, and the run says so. A restore onto a host whose previous database is gone
has nothing to read, and an archive that records no database_token or wal_lsn (one taken before
backups recorded them) cannot show which database it came from. In each of those cases nothing is
replayed, and the post-restore step instead lists every credential the restored database still
trusts, for you to compare against your own record of what was revoked.
Back up the chain off-box, too
Keep a second, database-independent copy of the chain: the NDJSON dump the offline verifier consumes. This is not a replacement for the database backup - it is the artifact you hand an auditor, and the copy that proves intact without a running Server.
./scripts/vault-dump.sh ./chain-backup
The shipped scripts/vault-dump.sh runs the dump tool inside the Server image, so it needs no source
checkout or pnpm, only a reachable database. It prints the dumped row counts and the output
directory:
{
"outDir": "/dump",
"orgId": null,
"counts": { "audit_vault": 10, "vault_checkpoints": 0, "vault_signing_keys": 1, "vault_key_statements": 1, "org_admin_reads": 0, "org_admin_reads_checkpoints": 0 }
}
It writes one NDJSON file per entry in counts, including vault_signing_keys.ndjson and
vault_key_statements.ndjson: the public-key registry and the signed statements that admit each key
travel with the dump, so the chain verifies offline with no further inputs. The audit runbook covers the file
contents and how to verify them.
Where it goes is the part that belongs here. Keep the dump alongside the db.dump tarball, not
instead of it, and treat it as a separate destination rather than a second file in the same place:
the database backup is what you restore a running Server from, and the NDJSON dump is what proves
the chain to someone who does not trust your Server, including after that Server is gone. A copy
that only exists next to the database it came from proves nothing the database could not have been
rewritten to say.
Point @agledger/verify at the directory, and pin both the version and the key.
2.0.0 is the first release that reads a dump from a 2.x Server; an older one fails it once an admin
has read a record.
npx -y @agledger/verify@2.0.0 ./chain-backup --trust-anchor <pin>
<pin> is the "Vault signing key pin" install.sh printed, sha256: and the digest of the signing
key's public half (node dist/scripts/signing-key-digest.js derives it again from the key). Without
--trust-anchor the verdict is VERIFIED, NOT ANCHORED: nothing failed, but every key was taken on
the dump's word.
The installer's pin is enough only while every key change on the install was a routine rotation,
which links each new key to the one before it. A key retired with force, or a recovery that staged
a key under VAULT_TRUST_ANCHORS, cuts that link, and a run with the installer's pin alone then
fails. Pass every pin the operator holds instead, each as its own --trust-anchor: the current key's
(anchoredFrom on GET /v1/verification-keys) and every digest in VAULT_TRUST_ANCHORS. Pass each
VAULT_DISTRUSTED_KEYS entry as --distrusted-key, since the dump does not carry them. The
key-compromise runbook lists the pins each recovery path leaves.
Any RFC 9052 library also verifies the per-row cose_sign1 bytes, when you want no AGLedger package
in the loop. @agledger/cli's verify
subcommand is a different tool: it checks one record's audit export
(GET /v1/records/{id}/audit-export) and reports EISDIR if you point it at a dump directory.
What a backup does not contain
-
Private signing keys. Held in
VAULT_SIGNING_KEY- your secret store, your responsibility. Back it up with your secrets, not your database. Without it a restored Server cannot sign new records, though existing records still verify against the public registry. A key held in AWS KMS (VAULT_SIGNING_KEY_KMS_ARN, see signing with a KMS key) never leaves KMS, so there is nothing to copy: what you protect is the KMS key itself, its key policy and its deletion schedule. A key the restored registry does not carry is registered on the next boot only with a predecessor to vouch for it:VAULT_SIGNING_KEY_PREVIOUSset to a key the registry trusts (which then signs the succession), or, with no such key,VAULT_TRUST_ANCHORSnaming the pins of the keys whose history you vouch for. Without either the process signs nothing (signingKey.gate: "unanchored"). Keep the pin the installer printed (Pin: sha256:<hex>) with the key; the vault public-key CSV each backup writes is the list to compare against when a key's provenance is in question. -
Customer encryption keys. In encrypted mode, the keys that decrypt payloads never reach the Server. They are not in any backup the Server can produce.
-
API_KEY_SECRET. Back it up with the signing key, and for the same reason: the dump holds what it protects and not the secret itself.api_keysrows store an HMAC of each key, keyed on this secret, so restoring the rows under a differentAPI_KEY_SECRETleaves every restored credential failing authentication at once, which is intact data and no way in.If you still hold the old secret, you do not have to re-mint anything. Put it in
API_KEY_SECRET_PREVIOUS(Compose:compose/.env; Helm:secrets.apiKeySecretPrevious) and the Server re-hashes a token under it whenever the primary hash misses, so existing keys authenticate again while new ones are minted under the current secret. That is the same lever a planned secret rotation uses. Re-minting is the path only when the old secret is genuinely gone; the recovery runbook has both. -
The webhook encryption key. Webhook subscription secrets are encrypted with
WEBHOOK_ENCRYPTION_KEY, or withAPI_KEY_SECRETwhen that is unset. The lever to change it is not the_PREVIOUSvariable above unless you never set a dedicated webhook key: rotatingWEBHOOK_ENCRYPTION_KEYneedsWEBHOOK_ENCRYPTION_KEY_PREVIOUS. A stored secret is tried under the current key, then the previous one, then re-encrypted under the current key on its next delivery, so a backup restored under a changed key still delivers once the previous key is in place. Back up whichever of the two you actually encrypt webhook secrets under, alongside the signing key. -
Provisioning YAML. Keep
PROVISIONING_CONFIG_PATHunder source control; it is config, not data.