Day-2 Operations
This page covers the recurring work after install: watching health, scraping metrics, watching how your agents are behaving, rotating the signing key, keeping partitions ahead of growth, reloading config, and upgrading. Two facts shape everything below.
- Two processes. A Server runs an API process and a worker process. The API serves requests; the worker runs scheduled maintenance (partition upkeep, recovery sweeps) and exposes its own metrics. Some gauges live only on the worker, noted where it matters. Both tiers are stateless and scale independently; see the high-availability runbook for the supported multi-replica shape.
- Two key roles. Org-scoped
adminkeys govern one org. Cross-org operator surfaces (system-health, signing-key rotation, provisioning) require aplatformkey. An admin key calling these gets a403that says so:
"detail": "Action 'ROTATE_VAULT_SIGNING_KEY' requires platform role; caller resolved as 'admin'.",
"recoveryHint": "This action is platform-scoped ... mint a platform-role key via POST /v1/admin/api-keys (platform-only)."
Health and readiness probes
Three unauthenticated endpoints, shaped for orchestration probes:
curl -s "$AGLEDGER_API_URL/livez" # liveness: is the process up?
curl -s "$AGLEDGER_API_URL/readyz" # readiness: can it serve (DB reachable)?
curl -s "$AGLEDGER_API_URL/health" # detailed status + version
{"status":"alive","timestamp":"2026-10-01T15:03:20.247Z"}
{"status":"ready","version":"2.0.0","timestamp":"2026-10-01T15:03:20.251Z"}
{"status":"ok","version":"2.0.0","timestamp":"2026-10-01T15:03:20.237Z","signingKey":{"gate":"usable","keyId":"c4dd3e20388b594d"}}
Wire livez to your liveness probe and readyz to your readiness probe. /health/ready is an
alias of readyz. None require auth, so probes need no credentials.
signingKey on /health says whether this process may sign. usable is healthy; unsigned is
the dev shape with no key configured. retired, unregistered, unanchored (a key no signed
statement links to the registry's history; see How rotation works) and
signer_unreachable (a KMS key that is not answering) turn /health to degraded and readyz to
503 with a reason, so a process that cannot sign leaves the load balancer instead of answering
errors. A readyz 503 with "reason":"database unavailable" is the database side of the same
probe.
The aggregate health view
GET /v1/admin/system-health (platform key) is the one-call operator summary: database latency and
pool, every pg-boss queue, the webhook dead-letter count, the worker processes attached, the
AGLedger versions connected to the database, and process memory.
curl -s -H "Authorization: Bearer $AGLEDGER_PLATFORM_KEY" "$AGLEDGER_API_URL/v1/admin/system-health"
{
"status": "healthy",
"degradedReasons": [],
"uptime": 68.95,
"database": { "status": "healthy", "latencyMs": 0.23, "failure": null, "pool": { "total": 4, "idle": 4, "waiting": 0 } },
"queues": {
"phase2-gate": { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
"webhook-delivery": { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
"webhook-delivery-dlq": { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
"maintenance": { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
"federation": { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
"federation-outbound": { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
"federation-outbound-dlq": { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
"cascade-cancel": { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
"vault-scan": { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 }
},
"connectedVersions": [
{ "version": "2.0.0", "connections": 9, "oldestConnectionAt": "2026-10-01T15:02:11.481Z" }
],
"workerConnections": 3,
"webhookDeadLetters": 0,
"process": { "rssMb": 143.53, "heapUsedMb": 59.17, "heapTotalMb": 64.8 },
"timestamp": "2026-10-01T15:03:20.430Z"
}
database.failureis null unlessdatabase.statusisoutage, where it names why the probe did not answer:unreachable,saturated(the database at its connection limit),no_connection,pool_exhausted(this Server's own pool fully taken),shutting_down, orerror.- A queue whose stats could not be read is listed with
nullrather than left out. connectedVersionslists every AGLedger version with a connection open on the database, read frompg_stat_activity. More than one entry means two releases are serving one schema, which is what a rolling upgrade and a part-finished rollback both look like; see version upgrades.nullmeans the view could not be read, which is not the same as nothing being connected.workerConnectionsis the database connections worker processes hold. A running worker holds at least one, so0means no worker is attached. It isnullon a process that enqueues no jobs, or wherepg_stat_activitycould not be read.webhookDeadLettersis every delivery parked in the webhook dead-letter table, an inactive subscription's included;nullmeans the table could not be counted.
A growing queues.*.failed count or a climbing pool.waiting is your earliest signal of trouble.
What moves status
status is the field to alert on, and degradedReasons is why it moved, one line per cause. It
reads degraded when:
- the database cannot serve at all, or answers but its
SELECT 1probe took more than 2 seconds. The first case is visible here only while the calling key is still in the API's auth cache (AUTH_CACHE_TTL_MS); past that, authenticating the call needs the database and the call itself answers 503DATABASE_UNAVAILABLE, so alert on a 503 from this route too; - any dead-letter queue holds work, or the webhook dead-letter table holds any delivery: a dead-lettered job has exhausted its retries, or failed in a way no retry fixes, and stays there until you recover it;
- a queue's stats, or the webhook dead-letter count, could not be read at all, which is the same shape of problem: the check that would catch a backlog did not run;
- no worker process is attached to the database, or one is attached but has stopped taking jobs,
so queued jobs wait. A worker stops taking jobs when its vault signing key is retired, unanchored
or not registered yet, or the KMS holding it is not answering; its own
GET /healthnames which asjobConsumption, and the reason here names the hold askey_retired,key_unanchored,key_unregisteredorsigner_unreachable. A worker with no vault key at all keeps consuming. A worker that is not consuming takes no job, including one queued before it started, and resumes on its own once its key can sign again; - a worker is attached and gives no reason to hold, yet a queue it consumes has held a ready job and started nothing and run nothing inside its lease for over three minutes: the process is attached but not running (a paused container, a wedged event loop), and the reason names the queues and says to unpause or restart it. A worker working through a backlog keeps starting jobs, and an idle install has nothing waiting, so neither reads this way;
- the database role lacks a privilege the engine writes with (what a restore with
--no-privilegesor a hardening script leaves), so every request that needs it answers 500/problems/database-privilege-missing, which is not retryable. The reason names each missing privilege and theGRANTthat restores it, to run as the role that owns the schema, and the next such request succeeds with no restart. Each process also logs it at boot. Readiness does not move on it, so a missing grant fails only the requests that need it rather than taking every process out of rotation; - this process's vault signing key is not usable (
signingKey.gateonGET /health), or a detected rewind refuses chain writes; - the key registry's trust walk reports a finding (
key_statement_invalid,key_closure_invalid,key_window_drift). Nothing is refused, but an auditor walking the published key statements fails exports of what the affected keys signed;POST /v1/admin/vault/scannames each finding and its remedy; - this process's cache-invalidation LISTEN connection is down, so changes other replicas broadcast (key revocations, signing-key retirements, rewind detections, rate-limit exemptions) reach it only as each cache expires;
- a webhook subscription drops its new events at dispatch: its circuit breaker keeps reopening, or
the engine deactivated it with the breaker open, by the auto-disable or a 410 Gone answered while
the breaker was open (
GET /v1/admin/webhooks/healthlists them). Closing the breaker of an engine-deactivated one (PATCH /v1/admin/webhooks/{webhookId}/circuit-breakerwithstate: closed, accepted on an inactive subscription) dismisses it and leaves it inactive. A subscription you delete is not counted, because the delete closes its breaker; - with anchoring on, vault checkpoints have waited for their S3 anchor past two checkpoint intervals
and an hour, so the anchor store is not taking them; or the API and the worker disagree on
whether anchoring is on; or the
org_admin_readscheckpoint sweep is not scheduled; - a provisioning file or entry does not load, or the last reconcile refused an entry whose file
parsed (a schema edit the compatibility check rejects, a name held by a resource provisioning does
not manage).
GET /v1/admin/provisioning/statuslists both inloadErrors, the refused entries until a reload applies them, across replicas and restarts.
Each line in degradedReasons says what clears it. GET /v1/admin/ops-summary carries the same
status and degradedReasons under system, and two posture fields that do not move status:
vault.keyRegistry.findings, the count behind the registry line above, and vault.appendOnly,
which reads enforced: false when the database role in DATABASE_URL owns the chain tables or is a
superuser, so the append-only revokes do not bind it (each process also logs a WARN at boot).
A failed count on a live queue is not one of those conditions and deliberately does not move
status: the job is being retried and clears itself, and a field that flickers stops being read.
That is why the failed column above is worth watching by eye even while status is healthy.
{
"status": "degraded",
"degradedReasons": [
"2716 dead-lettered job(s) in federation-outbound-dlq; they will not be retried until an operator recovers them"
]
}
Dead letters
Both dead-letter surfaces count, and they are different shapes. Queue dead letters (federation,
cascading gates, and the rest) live in pg-boss and are recovered from their own admin route, e.g.
GET /federation/v1/admin/dlq, which lists each entry with the peer and record it belongs to.
Webhook dead letters live in a table, not a queue: a permanent failure (an SSRF refusal, a 410,
a 4xx, an undecryptable secret) never reaches the queue at all, so watching queue names alone would
read healthy through a total webhook outage. Those are at GET /v1/admin/webhook-dlq, and
degradedReasons names that route when they are what moved status.
A receiver that answers 410, or one the engine deactivates after sustained failures, leaves its
subscription inactive, and so does DELETE /v1/webhooks/{webhookId}. That subscription's dead
letters keep counting, and degradedReasons gives them a line of their own, apart from the active
subscriptions' entries. They cannot be retried: a retry answers 422 with
allowedActions: ["discard"], and retry-all leaves them in place and counts them in
skippedInactive. The listing marks each with subscriptionActive: false. To clear one, discard
it; the event itself stays in GET /v1/events and on its record:
curl -s -X DELETE -H "Authorization: Bearer $AGLEDGER_PLATFORM_KEY" \
"$AGLEDGER_API_URL/v1/admin/webhook-dlq/<dlqId>"
The subscription's owner can discard its own entries with
DELETE /v1/webhooks/{webhookId}/dlq/{dlqId}. Deleting a subscription answers 200 with
deadLetters, the count it still holds, and nextSteps leading to the per-subscription listing and
discard.
A provisioning-managed subscription deactivated by a 410 or the auto-disable comes back active on
the next reload while its declaration is in the directory (POST /v1/admin/provisioning/reload),
and its entries can then be retried. A replacement subscription (POST /v1/webhooks) receives new
events only; the discarded ones stay readable from GET /v1/events.
When you recover a queue DLQ with POST /federation/v1/admin/dlq/recover, the recovered count is
the number of rows the call actually removed. A row another worker took in the meantime is not
counted, so treat the number as an instruction to re-check the depth rather than as a total.
The public status page
GET /status needs no key and is rate limited. It reports whether each component can serve, not
whether work is backed up:
curl -s "$AGLEDGER_API_URL/status"
{
"status": "operational",
"components": [
{ "name": "API", "status": "operational" },
{ "name": "Database", "status": "operational", "latencyMs": 0.31 },
{ "name": "Chain writes", "status": "operational" },
{ "name": "Workers", "status": "operational" }
],
"uptime": 68.95,
"timestamp": "2026-10-01T15:03:20.430Z"
}
Each component reads operational, degraded or outage, and the top-level status is the worst
of them. A component that is not operational says why in reason, except a Database at
degraded that answered its probe slowly, which carries latencyMs instead:
- Database at
outage: the same values asdatabase.failureabove. Atdegradedwithprivilege_missing: the role inDATABASE_URLlacks a privilege the engine writes with, andGET /v1/admin/system-healthnames it with itsGRANT. Atdegradedwithconnections_overcommitted: at boot this Server estimated that its pool, the worker loops and pg-boss together need more connections than the database'smax_connections, so under load they take every connection it has; lowerDATABASE_POOL_MAXor the worker concurrency settings, or raisemax_connections, and restart. - Chain writes at
outagewithdatabase_unavailable(the Database component is atoutage),privilege_missing(the role lacks SELECT or INSERT onaudit_vault, which every chain write reads its head from and appends to) orsigning_key_unusable(this process's vault signing key may not sign; seesigningKey.gateonGET /health); atdegradedwithchain_rewind_detected(every chain write answers 409 until a platform operator acknowledges the rewind). - Workers, listed only by an API process that enqueues jobs:
outagewithdatabase_unavailablewhen the database is unreachable;degradedwithno_worker_connected(no worker attached),worker_not_consuming(one attached but not taking jobs; its ownGET /healthnames why asjobConsumption),worker_stalled(one attached that gives no reason and has started no job on a queue holding ready work for over three minutes; system-health names the queues),not_checked(this process's database probe failed for a reason of its own, so workers may still be running),privilege_missing(the role inDATABASE_URLcannot use thepgbossschema, so every job this process enqueues is refused and a webhook push enqueued then is lost, its event still readable fromGET /v1/events;GET /v1/admin/system-healthnames the remedy, which is running the migrate step again), orerror(the check could not be read).
Dead letters, of a queue or of webhooks, do not move /status. A failing receiver is not the
platform being down, and the backlog is not for an unauthenticated page. Alert on
GET /v1/admin/system-health for those.
Metrics
GET /metrics exposes Prometheus metrics on both the API and the worker process. When
METRICS_AUTH_TOKEN is set, which both packaged installs do for you, /metrics on either process
takes that token as a bearer and nothing else: an API key gets 401 (from the API with a
recoveryHint saying an API key is not accepted there, from the worker with no body). With no
token, METRICS_AUTH_REQUIRED (default true in production) decides: the API's /metrics then
takes the same API-key chain as /v1, while the worker, which has no API-key chain, answers 503
rather than serve unauthenticated. Set METRICS_AUTH_REQUIRED=false to serve both openly and
restrict them at the network layer. Under Compose the worker's port is not published to the host
at all, and under the chart it is cluster-internal. All series are prefixed agledger_. The ones
worth alerting on:
| Metric | Watch for |
|---|---|
agledger_vault_integrity_check_results_total{result="broken"} | Any increase: a chain, or the key registry's trust walk, failed periodic verification. Worker only |
agledger_db_pool_waiting_connections | Sustained nonzero: pool saturation |
agledger_pgboss_queue_size{queue=~".*-dlq",state="total"} | Growth: jobs dead-lettering into a DLQ |
agledger_pgboss_queue_size{state="queued"} | Sustained growth on any queue: a worker is down or cannot keep up. Every queue the Server runs reports here, federation included |
agledger_pg_listener_connected | 0 for more than a few minutes on one process: that replica does not hear cross-replica cache invalidations |
agledger_pg_listener_reconnect_failures_total | Increase: a LISTEN reconnect attempt failed |
agledger_vault_checkpoint_skipped_broken_total | Increase: a record went un-anchored. Worker only |
agledger_outbound_ssrf_blocked_total | Increase: outbound calls hitting the SSRF guard |
agledger_federation_zero_row_oldest_candidate_age_seconds | Climbing toward the recovery horizon: crash-orphaned records are aging out unrecovered. Federated deployments only; see Federation delivery and recovery |
agledger_vault_signer_unreachable | 1: the KMS signing key is not answering, and this process answers 503 on every write it would sign. Always 0 with a local key |
The worker serves /metrics on its health port (WORKER_HEALTH_PORT). Under Compose that port is
not published to the host, and the host's port 3001 (AGLEDGER_HOST_PORT) is the API, so read the
worker's series from inside its container:
docker compose exec -e NODE_OPTIONS= agledger-worker /nodejs/bin/node -e \
"fetch('http://localhost:3001/metrics',{headers:{authorization:'Bearer '+process.env.METRICS_AUTH_TOKEN}}).then(r=>r.text()).then(t=>process.stdout.write(t))" \
| grep agledger_vault_integrity_check_results_total
The bundled Prometheus scrapes the worker for you, as its own job. Note that
agledger_vault_integrity_check_results_total and agledger_vault_checkpoint_skipped_broken_total
move on the worker process only. The API's /metrics carries them too, with every series at 0
for as long as it runs, so a zero there says nothing about the chain.
agledger_partition_runway_days (next section) is not on the API's /metrics at all. Scrape both
processes. The federation series are worker-only too.
You do not have to write those alerts yourself. monitoring/alerts/agledger.rules.yml in the
install repository ships a rule set covering silent
drops, chain integrity, partition maintenance, federation delivery, and availability, and the
bundled Prometheus loads them already. Read the file for the current groups and counts rather than
a number quoted here: releases add rules, and every count this page has carried went stale. No
Alertmanager is bundled and no routing is configured, because receivers and escalation
policy are yours to decide; until you point them somewhere the rules evaluate on Prometheus' own
/alerts page, and each carries a severity label of critical or warning as the routing hook.
Every rule also carries a runbook_url annotation pointing at its section of
monitoring/runbooks.md in the same repository: what fired, what to check, what to do, and when
silencing is safe.
On Kubernetes with the Prometheus Operator, the chart renders the same rules and dashboards:
monitoring.prometheusRule.enabled, monitoring.serviceMonitor.enabled and
monitoring.grafanaDashboards.enabled (Grafana sidecar ConfigMaps), all off by default.
monitoring.prometheusRule.runbookUrl repoints every rule's runbook link at your own copy.
Three Grafana dashboards auto-provision with install.sh --with-monitoring: an overview (traffic,
chain throughput, saturation), data-integrity surveillance, and a silent-drop board with a panel per
fire-and-forget path that swallows its error. monitoring/README.md describes all three, including
how to import them into your own Grafana instead of the bundled one.
Watching your agents
An agent that starts behaving differently is usually the first sign of a problem, and the sign is the change itself, in either direction. An acceptance rate that moves from 0.8 to 1.0 deserves the same look as one that moves to 0.6: something changed about the work, the gate, or the agent. The Server reports the change and stops there. It does not score agents, rank them, or decide what a move means. That decision belongs to whoever is watching.
Start with the whole org. An org-scoped admin key (the admin-standard or admin-observer
profile) lists every agent with what it did in the current window, the same counts for the equal
window before it, and the difference:
curl -s "$AGLEDGER_API_URL/v1/agents/drift?window=7" \
-H "Authorization: Bearer $ADMIN_KEY"
{
"window": {
"days": 7,
"currentFrom": "2026-09-01T21:47:40.902Z",
"currentTo": "2026-09-08T21:47:40.902Z",
"baselineFrom": "2026-08-25T21:47:40.902Z",
"baselineTo": "2026-09-01T21:47:40.902Z"
},
"data": [
{
"agentId": "01a082fb-8266-7242-8c57-98f4f2c612dd",
"displayName": "scraper-bot",
"current": { "from": "2026-09-01T21:47:40.902Z", "to": "2026-09-08T21:47:40.902Z", "records": 9, "completions": 0, "verdicts": 0, "accepted": 0, "rejected": 0, "overturned": 0, "acceptanceRate": null, "medianCompletionMs": null },
"baseline": { "from": "2026-08-25T21:47:40.902Z", "to": "2026-09-01T21:47:40.902Z", "records": 0, "completions": 0, "verdicts": 0, "accepted": 0, "rejected": 0, "overturned": 0, "acceptanceRate": null, "medianCompletionMs": null },
"change": { "records": 9, "completions": 0, "verdicts": 0, "overturned": 0, "acceptanceRate": null, "medianCompletionMs": null }
},
{
"agentId": "01a082fb-8258-7801-bd1d-4cbef5765ec9",
"displayName": "invoice-bot",
"current": { "from": "2026-09-01T21:47:40.902Z", "to": "2026-09-08T21:47:40.902Z", "records": 5, "completions": 5, "verdicts": 5, "accepted": 4, "rejected": 1, "overturned": 0, "acceptanceRate": 0.8, "medianCompletionMs": 21 },
"baseline": { "from": "2026-08-25T21:47:40.902Z", "to": "2026-09-01T21:47:40.902Z", "records": 0, "completions": 0, "verdicts": 0, "accepted": 0, "rejected": 0, "overturned": 0, "acceptanceRate": null, "medianCompletionMs": null },
"change": { "records": 5, "completions": 5, "verdicts": 5, "overturned": 0, "acceptanceRate": null, "medianCompletionMs": null }
}
],
"total": 2,
"nextCursor": null,
"hasMore": false
}
Read the two agents side by side. scraper-bot notarizes: its records terminalize on create with
no completion and no gate, so its only signal is volume, and nine records against a baseline of
zero is a new agent or a new workload. invoice-bot is gated: five records, five completions,
five verdicts, one rejected. acceptanceRate is accepted / verdicts, and medianCompletionMs
is the median time from a record's activation to its completion being submitted. Both are null
in a window with nothing to divide, and change is null whenever either side is, so a brand-new
agent reads as counts that went up and rates that do not exist yet, rather than as a rate that
jumped from zero.
Every count is by when the thing happened: records by creation, completions by submission,
verdicts by the moment the verdict landed, disputes by resolution. verdicts is the final verdict
on each completion. A principal-gated record carries the engine's structural pass and then the
principal's verdict on the same completion, and only the principal's counts. A record that is
revised and resubmitted is a new completion and so a new verdict, so a revision cycle shows as two
verdicts with one of each outcome. The verdict count sits beside the rate so one record's cycle
cannot be mistaken for ten records' worth of rejections.
window is the length in days, from 1 to 365, default 7. The current window is the last N days and
the baseline is the N days before it, so the two are always comparable without scaling. The listing
pages by agent with limit and cursor; a cursor is bound to the window it was minted under and
refuses to continue under another.
When an agent moved, ask about that agent. The per-agent read carries the same series overall and then one per type, so a change in the roll-up can be traced to the type that carried it:
curl -s "$AGLEDGER_API_URL/v1/agents/01a082fb-8258-7801-bd1d-4cbef5765ec9/drift" \
-H "Authorization: Bearer $ADMIN_KEY"
{
"agentId": "01a082fb-8258-7801-bd1d-4cbef5765ec9",
"window": { "days": 7, "currentFrom": "...", "currentTo": "...", "baselineFrom": "...", "baselineTo": "..." },
"overall": {
"current": { "from": "...", "to": "...", "records": 5, "completions": 5, "verdicts": 5, "accepted": 4, "rejected": 1, "overturned": 0, "acceptanceRate": 0.8, "medianCompletionMs": 21 },
"baseline": { "...": "..." },
"change": { "...": "..." }
},
"byType": [
{
"type": "principal-gate-generic-v1",
"current": { "from": "...", "to": "...", "records": 5, "completions": 5, "verdicts": 5, "accepted": 4, "rejected": 1, "overturned": 0, "acceptanceRate": 0.8, "medianCompletionMs": 21 },
"baseline": { "...": "..." },
"change": { "...": "..." }
}
]
}
?type= narrows byType to one type; overall still covers everything. The rows behind the
counts are one call further: GET /v1/agents/{agentId}/history lists every record the agent acted
on, newest first, with type, outcome, from and to filters, so the one rejected record above
is ?outcome=reject.
An agent key reads its own drift and its own history with the same calls, and nothing else:
another agent's id answers 403, and the org-wide listing answers 403 naming org-admin as the
role it takes. The scope on all three routes is drift:read, which admin-standard,
admin-observer and agent-full carry. A platform key reads any single agent but has no org to
list, so the org-wide call refuses it and says to use an org-bound admin key.
Nothing here is cached and nothing is computed ahead of time. Each call reads the records,
completions, verdicts and disputes tables for the two windows, so the answer is current to the
request and there is no counter to fall behind or rebuild. There is nothing to configure and no
external dependency, so it works identically air-gapped. Field-level detail for all three routes
is in the API reference under the Agent Drift tag.
Signing-key rotation
Rotation is the load-bearing day-2 task. The guarantee that makes it safe:
Rotating the signing key never breaks verification of already-signed records. Retired keys stay in the published registry, so a record signed under an old key still verifies after any number of rotations. No re-signing, no downtime.
How rotation works
Rotation is two steps, and they are deliberately separate: staging the new key, then retiring the old one once nothing is signing with it.
Stage. Generate a new key with the script in the Server image (add --algorithm es256 for an
ES256 key); it runs from the image you already loaded, so the step works air-gapped:
docker run --rm agledger/agledger:<version> dist/scripts/generate-signing-key.js
Set the new key as VAULT_SIGNING_KEY and the key in use as VAULT_SIGNING_KEY_PREVIOUS, then
restart every api and worker process. A single-replica Helm install counts as more than one
process, since the api and the worker are separate Deployments and both sign. On Helm a
chart-managed Secret change rolls both by itself; with existingSecret set nothing watches the
Secret, so restart both Deployments yourself
(kubectl rollout restart deploy -n <ns> -l app.kubernetes.io/instance=<release>). The first process to boot registers the new key, activates
it beside the key that is already active, and writes a succession statement signed by both
keys. That statement is what makes the new key trusted, to every other process and to every
auditor, and once it exists other processes on the new key need no VAULT_SIGNING_KEY_PREVIOUS.
The boot logs the staging:
{"level":"info","keyId":"6a639248683aab56","activeSigningKeyIds":["affc2b9bfb22144e","6a639248683aab56"],"msg":"Staged vault signing key"}
A process on a new key with no predecessor signer registers nothing and signs nothing: /health
reports signingKey.gate: "unanchored", /health/ready answers 503, a worker stops consuming jobs
(the API's /status then reads Workers degraded with worker_not_consuming), and the rotate
endpoint answers 409 with a recoveryHint. The out-of-band alternative to a predecessor is
VAULT_TRUST_ANCHORS, a comma list of sha256: pins of keys whose history you vouch for: it is the
recovery path when no key in hand links to that history, the process registers under a fresh
genesis, and auditors then need the new key's pin.
Staging writes a KEY_ROTATED entry on the platform chain and a vault.signing_key_rotated row in
system_audit_log, so a key change is on the record whether it arrived by a restart or by the API.
Staging retires nothing. Until you close the old key's window the install has two active keys, which is the normal state of a key change and costs nothing but an extra published key. A rolling restart means some processes are still holding the old key, and every process you have not restarted yet keeps signing inside its own key's published window, so nothing it writes fails verification.
POST /v1/admin/vault/signing-keys/rotate (platform key) performs the same registration on
demand. After a restart the process has already registered its key, so the endpoint reports
already_active. That is the expected answer: use it to confirm the staging and read
activeKeys, which lists every key still able to sign with the last instant each one signed. It
never retires anything.
{
"newKeyId": "6a639248683aab56",
"status": "already_active",
"activeKeys": [
{ "keyId": "6a639248683aab56", "algorithm": "Ed25519", "activatedAt": "2026-09-22T14:02:11.000Z", "lastSignedAt": "2026-09-22T14:05:40.118Z" },
{ "keyId": "affc2b9bfb22144e", "algorithm": "Ed25519", "activatedAt": "2026-05-26T09:12:03.000Z", "lastSignedAt": null }
]
}
A key held in AWS KMS rotates the same way: staging is the new key's ARN in
VAULT_SIGNING_KEY_KMS_ARN (signing.kmsKeyArn on the chart) and a restart, with the outgoing KMS
key in VAULT_SIGNING_KEY_PREVIOUS_KMS_ARN (signing.previousKmsKeyArn), and the process given
kms:GetPublicKey and kms:Sign on both keys for the change; the old key never leaves KMS and signs
only the succession. VAULT_SIGNING_KEY and VAULT_SIGNING_KEY_KMS_ARN
together refuse to boot, and so do the two predecessor variables, so moving from a local key to
KMS is one rotation: put the local key in VAULT_SIGNING_KEY_PREVIOUS, set the ARN, unset
VAULT_SIGNING_KEY, restart, then retire the local key id.
Signing with a KMS key covers the key policy.
Retire. First confirm every process has rolled onto the new key. Read GET /health on each
api and worker process directly, not through a load balancer, and check that signingKey.keyId
is the new key everywhere. That is the check that proves no process still holds the old one.
lastSignedAt on GET /v1/admin/vault/signing-keys is not: it proves only that a key appended to
the chain recently, and reads null on an idle process or one whose only use of the key is webhook
deliveries, certificates or Receipts. Then close the old key's window through a process on the new
key (a process on the key being retired answers 409):
curl -X POST -H "Authorization: Bearer $AGLEDGER_PLATFORM_KEY" \
-H "Content-Type: application/json" -d '{}' \
"$AGLEDGER_API_URL/v1/admin/vault/signing-keys/<old-key-id>/retire"
The response carries retiredKeyId, the retiredAt instant that is now the published end of the
key's window, the keys still active, and the closureDigest of the closure statement the calling
process's key signs over that retiredAt. Anything the retired key admits after the closure counts
for nothing. The call refuses with 422 and names the key while it has appended to the chain
within the last 300 seconds; that refusal is a backstop for what the chain can see, not a substitute
for the /health check. Retiring the only active key is always refused, because it would leave
nothing able to sign: stage the replacement first.
After retirement, an entry written under the old key is a chain break: the Server reports it as
key_expired and an offline verifier as CHAIN_KEY_EXPIRED. Entries written before it stay valid.
A process that keeps holding a retired key does not write such an entry: it stops signing and
answers 503 on /health/ready, so a load balancer takes it out of rotation. The fix is a restart
onto the active key, and a retired key never comes back.
For a leaked key, do not wait. The compromise order retires with {"force": true} immediately
after the restart, which also revokes the key's ephemeral certificates and can leave other keys
untrusted until you pin them. Follow the key-compromise runbook
for that order rather than this one.
Remove VAULT_SIGNING_KEY_PREVIOUS (or VAULT_SIGNING_KEY_PREVIOUS_KMS_ARN) once the old key is
retired, and roll again at your convenience. Finish a key change before a version upgrade starts,
and do not start one during a rollout.
Reading the published keys
After retirement both keys appear at GET /v1/verification-keys: the new one active, the prior one
retired with the exact instants it was active (activatedAt / retiredAt, the values the key
statements sign), still resolvable:
curl -s "$AGLEDGER_API_URL/v1/verification-keys" \
| jq -r '.data[] | "\(.keyId) \(.status) \(.activatedAt) \(.retiredAt)"'
6a639248683aab56 active 2026-09-22T14:02:11.000Z null
affc2b9bfb22144e retired 2026-05-26T09:12:03.000Z 2026-09-22T14:31:52.604Z
The one exception is a key VAULT_DISTRUSTED_KEYS names with an instant earlier than its signed
closure: there retiredAt is that distrust instant, since nothing the key signed after it counts,
and the entry carries distrustedFrom with the instant. Nothing signs that field, so it tells an
auditor what to ask the operator to confirm, not that the window is untampered.
That surface publishes anchored keys only, with anchoredFrom, the sha256: digest of the key the
serving process signs with. It equals the pin the installer printed only until the first rotation;
after one it is the new key's digest. The installer's pin stays a valid trust anchor across routine
rotations, because each new key is admitted by a signed statement from the key before it, so a
verifier given that pin walks the statements forward to the current key. A retirement with
{"force": true} cuts that walk, and its response says which pin to hand auditors instead.
Proving the guarantee
Records signed before the rotation must still verify, and you can check that rather than take it on
trust. Dump the vault and verify it offline by the off-box verification in the
audit runbook, with @agledger/verify 2.0.0 or later (a 2.0 chain needs
that floor, published per key as minVerifierVersion). After a rotation the dump carries both
signing keys and entries signed by each, and the verifier resolves whichever key signed each entry,
so a clean run covers the retired key as well as the active one.
Partition maintenance
Four high-volume tables are range-partitioned by month, and only system_audit_log carries a
DEFAULT catch-all partition. audit_vault, events and webhook_deliveries have none, so a write
past the latest pre-created partition on any of those three has nowhere to land and fails outright:
runway exhaustion there is a write outage, not an overflow. The worker pre-creates upcoming
partitions ahead of the clock and exposes runway as a gauge (agledger_partition_runway_days) for
all four. You can also query the source function directly:
psql "$DATABASE_URL" -c "SELECT table_name, runway_days, default_rows FROM partition_runway();"
table_name | runway_days | default_rows
--------------------+-------------+--------------
audit_vault | 570 | 0
events | 570 | 0
webhook_deliveries | 570 | 0
system_audit_log | 83 | 0
runway_days is days until the latest pre-created partition is reached, and it is worth alerting
on low for all four tables. default_rows is only meaningful on system_audit_log, the one table
with a DEFAULT partition for a row to land in: a nonzero value there means writes are landing in it
and the worker is falling behind, so alert on default_rows > 0 for that table specifically. On
the other three it stays 0 by construction, since there is no DEFAULT partition to reach, so that
alert would never fire and runway_days is what catches trouble instead.
Config-as-code hot reload
If you run with PROVISIONING_CONFIG_PATH set, orgs, agents, webhooks, and contract schemas are
declared in YAML and reconciled on every boot; see the
provisioning runbook for the directory layout. Reload changes without
a restart via SIGHUP or POST /v1/admin/provisioning/reload (platform key). Check current state
first:
curl -s -H "Authorization: Bearer $AGLEDGER_PLATFORM_KEY" "$AGLEDGER_API_URL/v1/admin/provisioning/status"
{"configured":true,"configPath":"/etc/agledger/provisioning","dryRun":false,"prune":false,"lastReloadAt":"2026-06-09T15:59:53.404Z","managed":{"orgs":1,"agents":2,"webhooks":0,"schemas":2,"trustedIssuers":1},"loadErrors":["webhooks/acme.yaml: Environment variable ACME_WEBHOOK_SECRET is not set and has no default"],"loadWarnings":[],"pruneSuppressed":false,"trustedIssuersWarnings":[],"trustedIssuersUnchangedSinceLastLoad":true}
curl -s -X POST -H "Authorization: Bearer $AGLEDGER_PLATFORM_KEY" "$AGLEDGER_API_URL/v1/admin/provisioning/reload"
{
"orgs": { "created": 0, "updated": 1, "pruned": 0 },
"agents": { "created": 0, "updated": 2, "pruned": 0 },
"schemas": { "created": 0, "updated": 0, "pruned": 0 },
"apiKeys": { "created": 0, "skipped": 3, "generated": [] },
"errors": [
{ "resource": "config", "error": "webhooks/acme.yaml: Environment variable ACME_WEBHOOK_SECRET is not set and has no default" }
]
}
Reload is idempotent, and the counts read differently by resource. Every org and agent the files
declare that already exists counts as updated, changed or not; a webhook or schema counts as
updated only when the reload changed it, so an unchanged one counts nowhere. Existing keys count
as skipped. Reload is also fail-open: a single invalid file (here, an unset
${ACME_WEBHOOK_SECRET} substitution in a webhook file of your own) is reported in errors[] while
every valid resource still applies. Newly minted keys appear in apiKeys.generated[] with their raw
value exactly once, in that response body, so capture them then.
Fail-open is the shape to watch. A file that fails to parse is skipped whole: its resources are
silently absent, the rest of the reconcile succeeds, and the Server comes up healthy. status
re-reads the directory on every call and lists every unloadable file in loadErrors[], alongside
the errors from the last trusted-issuers pass this process ran, so that field, not the boot log, is
the durable signal. It also moves status on system health. The
agledger_provisioning_errors gauge (labels stage="load" / stage="reconcile") carries the same
state for alerting and stays non-zero until a clean reconcile clears it.
Three more fields on status answer the questions a clean-looking reload leaves open.
pruneSuppressed is true when prune is on and the directory does not load cleanly: a reload then
prunes nothing, so an org or agent deleted from the YAML to offboard it keeps its API keys active
until loadErrors is empty. trustedIssuersWarnings lists issuers whose OIDC discovery failed or
whose jwks_uri the egress guard refused; each kept the jwks_uri it already held while every other
field applied. trustedIssuersUnchangedSinceLastLoad: true after a reload that changed nothing means
the pod read the same trusted-issuers.yaml bytes as before: on Kubernetes, a ConfigMap mounted
through subPath never receives updates, so mount it as a directory.
A bare ${VAR} with no default is the usual cause: it is a hard parse error when the variable is
unset. Use ${VAR:-default} wherever a sensible default exists, and inject real secrets through the
pod environment (extraEnv / extraEnvFrom on the chart).
Version upgrades
Upgrade with scripts/upgrade.sh from the
agledger-ai/install repository; the
install runbook has the full procedure and the air-gapped path. Migrations run
automatically and are advisory-locked and checksum-verified, so a migration runs once and only once
even across replicas, and a tampered or reordered migration is refused rather than applied. For
diagnostics to attach to a support request, scripts/support-bundle.sh collects health, version,
container state and logs, and .env with credential values redacted by the shape of the name and
the value. Logs are collected as written, so review the archive before you send it.
Restart the process after swapping the image. Already-signed records verify across versions, because the chain format is stable and historical keys resolve.
What each release migrates, and whether its upgrade needs a maintenance window, is in that
release's notes in the agledger-ai/install changelog.
Read the notes for every release you are crossing before you upgrade.
2.0 and a 1.x database
2.0 does not upgrade a 1.x database. The migration refuses a database a 1.x release migrated
before it changes anything, and says so: the checksum it recorded for 001_consolidated.sql is not
the one 2.0 ships. Install 2.0 against a new, empty database, and keep the 1.x database for the 1.x
install that wrote it. A backup does not carry 1.x data across either: restore.sh returns the
install to the release the backup was taken on, --keep-version with a 1.x archive is refused
before anything is stopped or dropped, and an archive that records no version reaches the 2.0
migration, which refuses the restored database the same way.
upgrade.sh refuses an install whose recorded or running version is 1.x at its version check,
before it writes .env, takes a backup or stops anything. Put the checkout back on the 1.x
deploy/ tree afterwards: a docker compose command run from the 2.0 tree starts 2.0 against the
1.x database. When the version was not recorded and the migration is what refuses, the recovery
text names that tree rather than a re-run.
On Helm with the bundled PostgreSQL, a 2.0 chart refuses to render over a release a 1.x chart
installed, so helm upgrade fails before it replaces a pod or rewrites the release Secret, and the
1.x release keeps serving. Install 2.0 as a new release, which gets its own bundled database. The
check reads the running release through lookup, which answers nothing under a renderer
(helm template, Argo CD), so there it cannot refuse: do not sync a 2.0 render onto a 1.x bundled
release, because the migration refuses the database only after the pods are replaced, and
helm rollback does not restore the release Secret the upgrade rewrote. With an external database
the migrate Job is a pre-upgrade hook, so a refused upgrade fails before any 1.x pod is replaced.
Two releases against one schema
Within 2.x, every upgrade migrates the database before the new processes exist, so from the moment
the migration commits until the last process of the old release is gone, two releases serve one
schema. Under the chart's default Recreate strategy that window is the pod swap; under
RollingUpdate it is the whole rollout. Migrations are written to serve the previous release, so
the window needs no ordering. What it does bind is a security control a release introduces (such as
subjectAllowlist or jtiSingleUse on a trusted issuer): only processes carrying that release
enforce it, so it binds every request once connectedVersions on GET /v1/admin/system-health
holds one entry. helm rollback takes the processes back and leaves the schema where it is, which
is safe for the same reason; a worker on the older release fails jobs from schedules the newer
release added and counts them on agledger_maintenance_unknown_task_total until you roll forward
again. The install repository's README covers the same window under "Two releases against one
schema".
The runtime role
On an external database, the upgrade checks the runtime role twice: before the pre-upgrade backup,
and again after migrations. That role can stop being able to serve without the install changing. A
credential-rotation policy that drops and recreates it brings it back without its agledger_app
membership, because role membership does not survive a DROP ROLE. Left unchecked, the first thing
to fail is pg_dump during the backup, which reports the table it was refused and prints its whole
LOCK TABLE statement without naming the role.
Each check names the role and the exact GRANT. Stopping at the first one costs nothing: no backup
has been taken, no migration has run, and neither the version nor the image digest in compose/.env
has moved, so the running install is serving exactly as it was. Stopping at the second one leaves
the upgrade part-done, which it says: the migrations are applied, the worker stays stopped until the
upgrade finishes, and the previous version is still serving against the new schema. Either way,
applying the grant and re-running finishes it, and migrations already applied are skipped.
Migration lock budget under sustained writes
A migration that cannot take a table's lock within MIGRATION_LOCK_TIMEOUT (default 30s) rolls
back whole with lock_not_available, so running it again is safe.
- Both install paths retry it before giving up. On Compose the migrate CLI exits with status 75
for this failure alone, and
upgrade.shruns the migration again, three attempts in all, before it stops with the worker stopped and the previous API serving. On Kubernetes the migrate Job retries any failure up to itsbackoffLimit. On an external database the Job is a pre-upgrade hook, so the previous release keeps serving while it retries. On bundled PostgreSQL it is an ordinary resource: the new pods start beside it and crashloop until a migration lands, because the API and worker each refuse to start against a database that has not applied every migration their image ships, or whose history the migration refuses, and their log names which. On this path Helm records the revision asdeployedeither way, so read the migrate Job's status rather thanhelm statusto tell whether the upgrade finished;helm-install.shwaits on the Job and exits non-zero when it fails. - If it still fails, raise the budget or pick the moment. Set
MIGRATION_LOCK_TIMEOUT=120s(compose/.env, ormigrate.extraEnvon the chart) and re-run the upgrade, or run it in a low-traffic window. On the chart, sizehelm upgrade --timeout(default 5m) to cover every Job attempt, about four times the lock budget plus the migration's own run time; otherwise Helm reports the hook failed while a later attempt is still running.
Request timeouts
CONNECTION_TIMEOUT_MS (default 60000) is the socket inactivity timeout: how long a connection
may go without bytes moving before the server destroys it.
The ordering between the three timeout knobs is what matters. Keep it above HANDLER_TIMEOUT_MS
(default 30000) and below KEEP_ALIVE_TIMEOUT_MS (default 72000). If the socket timeout fires
first, a slow request is severed underneath the handler instead of failing through the route, and
the failure mode is unpleasant: on a bulk write the batch commits, the client sees a dead socket
with no response, and it has no way to tell a committed batch from a lost one.
Recovering from a severed socket: retry the request with the same Idempotency-Key. A commit
that already landed is returned rather than repeated.
Two signals tell you the value is too low for your traffic, and you need both because a severed socket never reaches the normal response path:
- the
agledger_requests_timed_out_totalcounter - a
request severed by socket inactivity timeoutwarning in the API log
Raise CONNECTION_TIMEOUT_MS if you see them on legitimately slow bulk work, keeping it under the
keep-alive value.
Federation delivery and recovery
Skip this section if you run a single Server with no peers. Nothing here applies until federation is configured, and the recovery sweeps below short-circuit before they touch the records table when a Server has no active peers.
Is this peer reachable?
GET /federation/v1/admin/peers answers it. Read lastDeliveryAt (the last outbound message
that got a 2xx), consecutiveDeliveryFailures (reset to 0 on a success) and lastDeliveryError
(cleared on the next success). A frozen lastDeliveryAt with a climbing failure count and a
transport error is a peer that has gone away, with retries still in flight until the job
dead-letters.
One field on the same object does not answer it, and it is worth knowing why: status is
registration state (active / revoked), not reachability. A peer you have never successfully
reached still reads active.
For a count rather than a list, GET /v1/admin/ops-summary reports federation.peers with the
registration counts and, partitioning the active ones, delivering / failing / neverDelivered.
That last split is the one a dashboard needs: active alone reads the same for a peer taking every
delivery and a peer that has never taken one.
For the backlog behind those deliveries, watch agledger_pgboss_queue_size{queue="federation-outbound"}
or read GET /v1/admin/system-health, which reports the same queues as JSON, and whose status
goes degraded once those retries give up and dead-letter.
Expect peers to lag, and size for it
POST /v1/records returns as soon as the record is notarized locally; peers receive their signed
copy afterwards. Your local chain is complete and verifiable the entire time a peer is behind, so
peer lag is a delivery concern, never an integrity one.
The outbound worker drains at a per-pod ceiling of FEDERATION_OUTBOUND_CONCURRENCY * FEDERATION_OUTBOUND_BATCH_SIZE / FEDERATION_OUTBOUND_POLLING_INTERVAL_SECONDS. On the defaults
that is 16 * 4 / 1.0 = 64 legs/sec/pod. A leg is one record
to one peer, so a record shared to two peers costs two legs. Notarizing 4,100 shared records
against two peers queues 8,200 legs and puts peers roughly two minutes behind. Ingest is faster
than drain by design, so a bulk share always builds a queue.
Raise batch size first if you need more throughput on a backlog spread across many peers and
records. It widens the fetch each worker loop already issues, so it lifts the ceiling without
adding database round-trips. It does nothing for a backlog concentrated in a handful of groups,
though: the worker pairs groupConcurrency: 1 with a (peer, record) group id to keep
same-record delivery ordered, which caps every batch at one job per group, so a larger batch size
pulls nothing extra and concurrency or polling interval is the lever instead. Concurrency and
polling interval are the expensive pair and interact super-additively on database CPU: each in
isolation costs roughly 30 to 40 percent on POST /v1/records p50/p95, but together at
(32 / 0.5) they cost about +110 percent p50 against the (16 / 1.0) baseline. Re-measure
record-creation latency after touching any of the three, and prefer scaling out worker pods over
cranking a single pod far past the defaults.
Same-(peer, record) ordering is preserved regardless of these settings, so raising throughput
never reorders state transitions for a given record.
The recovery sweep, and the window that abandons records
Outbound work is enqueued after the record's state change commits. A crash in that gap can leave a committed terminal record with no federation work queued at all. A sweep runs every two minutes on the worker process to find those and re-drive them through the normal publish path: same share gate, same peer set, same idempotency, so recovery can never deliver something the live path would not have.
Two knobs size it, and they are a pair:
| Knob | Default | What it controls |
|---|---|---|
AGLEDGER_FEDERATION_ZERO_ROW_HORIZON_MINUTES | 360 (6h) | How far back the sweep looks |
AGLEDGER_FEDERATION_ZERO_ROW_BATCH_SIZE | 50 | Records re-driven per cycle (~25/min) |
The horizon is an abandonment boundary, not just a scan bound. A record whose updated_at
falls outside the window is never recovered. Six hours covers a crash gap with wide margin, since
the normal case recovers within a sweep interval or two. Widen it only if this Server can be down
longer than that.
Each cycle takes the oldest candidates first, which is the only order that cannot starve the record closest to aging out.
The one number to alert on is agledger_federation_zero_row_oldest_candidate_age_seconds, the
age of the head of that queue. Flat and low is healthy. Climbing toward the horizon means the
candidate pool is refilling faster than one batch drains, which is the only remaining way a
genuinely orphaned record ages out unrecovered. Page well before the horizon.
Which knob to turn when that age climbs: raise AGLEDGER_FEDERATION_ZERO_ROW_BATCH_SIZE.
Widening the horizon in that state makes it worse, because it adds candidates to a pool that is
already draining too slowly. The horizon is the knob for a different problem, a Server that is
legitimately offline longer than six hours.
Records received from peers
A record this Server received from a peer is a read-only view of a row the originating Server owns, and only that Server fans it out. Such records are never recovery candidates: recovering them would deliver redundantly back to the origin and be permanently rejected by third peers.
This matters operationally because a busy receiving Server accumulates those records without
bound. agledger_federation_zero_row_projections_excluded reports how many were held out of the
candidate pool on the last cycle, capped at the batch size. Non-zero is the healthy reading on
any Server receiving federation traffic; it is the exclusion doing its job, not a problem.
Its companion agledger_federation_outbound_projection_skipped_total is a last-resort guard
further down the publish path and is expected to stay at 0 permanently. Treat any increase
there as a regression to report, not as confirmation that the exclusion is working.
All of these series are exposed by the worker process, not the API. Scrape both.