Day-2 Operations

This page covers the recurring work after install: watching health, scraping metrics, watching how your agents are behaving, rotating the signing key, keeping partitions ahead of growth, reloading config, and upgrading. Two facts shape everything below.

"detail": "Action 'ROTATE_VAULT_SIGNING_KEY' requires platform role; caller resolved as 'admin'.",
"recoveryHint": "This action is platform-scoped ... mint a platform-role key via POST /v1/admin/api-keys (platform-only)."

Health and readiness probes

Three unauthenticated endpoints, shaped for orchestration probes:

curl -s "$AGLEDGER_API_URL/livez"          # liveness: is the process up?
curl -s "$AGLEDGER_API_URL/readyz"         # readiness: can it serve (DB reachable)?
curl -s "$AGLEDGER_API_URL/health"         # detailed status + version
{"status":"alive","timestamp":"2026-10-01T15:03:20.247Z"}
{"status":"ready","version":"2.0.0","timestamp":"2026-10-01T15:03:20.251Z"}
{"status":"ok","version":"2.0.0","timestamp":"2026-10-01T15:03:20.237Z","signingKey":{"gate":"usable","keyId":"c4dd3e20388b594d"}}

Wire livez to your liveness probe and readyz to your readiness probe. /health/ready is an alias of readyz. None require auth, so probes need no credentials.

signingKey on /health says whether this process may sign. usable is healthy; unsigned is the dev shape with no key configured. retired, unregistered, unanchored (a key no signed statement links to the registry's history; see How rotation works) and signer_unreachable (a KMS key that is not answering) turn /health to degraded and readyz to 503 with a reason, so a process that cannot sign leaves the load balancer instead of answering errors. A readyz 503 with "reason":"database unavailable" is the database side of the same probe.

The aggregate health view

GET /v1/admin/system-health (platform key) is the one-call operator summary: database latency and pool, every pg-boss queue, the webhook dead-letter count, the worker processes attached, the AGLedger versions connected to the database, and process memory.

curl -s -H "Authorization: Bearer $AGLEDGER_PLATFORM_KEY" "$AGLEDGER_API_URL/v1/admin/system-health"
{
  "status": "healthy",
  "degradedReasons": [],
  "uptime": 68.95,
  "database": { "status": "healthy", "latencyMs": 0.23, "failure": null, "pool": { "total": 4, "idle": 4, "waiting": 0 } },
  "queues": {
    "phase2-gate":             { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
    "webhook-delivery":        { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
    "webhook-delivery-dlq":    { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
    "maintenance":             { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
    "federation":              { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
    "federation-outbound":     { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
    "federation-outbound-dlq": { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
    "cascade-cancel":          { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 },
    "vault-scan":              { "waiting": 0, "active": 0, "delayed": 0, "failed": 0 }
  },
  "connectedVersions": [
    { "version": "2.0.0", "connections": 9, "oldestConnectionAt": "2026-10-01T15:02:11.481Z" }
  ],
  "workerConnections": 3,
  "webhookDeadLetters": 0,
  "process": { "rssMb": 143.53, "heapUsedMb": 59.17, "heapTotalMb": 64.8 },
  "timestamp": "2026-10-01T15:03:20.430Z"
}

A growing queues.*.failed count or a climbing pool.waiting is your earliest signal of trouble.

What moves status

status is the field to alert on, and degradedReasons is why it moved, one line per cause. It reads degraded when:

Each line in degradedReasons says what clears it. GET /v1/admin/ops-summary carries the same status and degradedReasons under system, and two posture fields that do not move status: vault.keyRegistry.findings, the count behind the registry line above, and vault.appendOnly, which reads enforced: false when the database role in DATABASE_URL owns the chain tables or is a superuser, so the append-only revokes do not bind it (each process also logs a WARN at boot).

A failed count on a live queue is not one of those conditions and deliberately does not move status: the job is being retried and clears itself, and a field that flickers stops being read. That is why the failed column above is worth watching by eye even while status is healthy.

{
  "status": "degraded",
  "degradedReasons": [
    "2716 dead-lettered job(s) in federation-outbound-dlq; they will not be retried until an operator recovers them"
  ]
}

Dead letters

Both dead-letter surfaces count, and they are different shapes. Queue dead letters (federation, cascading gates, and the rest) live in pg-boss and are recovered from their own admin route, e.g. GET /federation/v1/admin/dlq, which lists each entry with the peer and record it belongs to. Webhook dead letters live in a table, not a queue: a permanent failure (an SSRF refusal, a 410, a 4xx, an undecryptable secret) never reaches the queue at all, so watching queue names alone would read healthy through a total webhook outage. Those are at GET /v1/admin/webhook-dlq, and degradedReasons names that route when they are what moved status.

A receiver that answers 410, or one the engine deactivates after sustained failures, leaves its subscription inactive, and so does DELETE /v1/webhooks/{webhookId}. That subscription's dead letters keep counting, and degradedReasons gives them a line of their own, apart from the active subscriptions' entries. They cannot be retried: a retry answers 422 with allowedActions: ["discard"], and retry-all leaves them in place and counts them in skippedInactive. The listing marks each with subscriptionActive: false. To clear one, discard it; the event itself stays in GET /v1/events and on its record:

curl -s -X DELETE -H "Authorization: Bearer $AGLEDGER_PLATFORM_KEY" \
  "$AGLEDGER_API_URL/v1/admin/webhook-dlq/<dlqId>"

The subscription's owner can discard its own entries with DELETE /v1/webhooks/{webhookId}/dlq/{dlqId}. Deleting a subscription answers 200 with deadLetters, the count it still holds, and nextSteps leading to the per-subscription listing and discard.

A provisioning-managed subscription deactivated by a 410 or the auto-disable comes back active on the next reload while its declaration is in the directory (POST /v1/admin/provisioning/reload), and its entries can then be retried. A replacement subscription (POST /v1/webhooks) receives new events only; the discarded ones stay readable from GET /v1/events.

When you recover a queue DLQ with POST /federation/v1/admin/dlq/recover, the recovered count is the number of rows the call actually removed. A row another worker took in the meantime is not counted, so treat the number as an instruction to re-check the depth rather than as a total.

The public status page

GET /status needs no key and is rate limited. It reports whether each component can serve, not whether work is backed up:

curl -s "$AGLEDGER_API_URL/status"
{
  "status": "operational",
  "components": [
    { "name": "API", "status": "operational" },
    { "name": "Database", "status": "operational", "latencyMs": 0.31 },
    { "name": "Chain writes", "status": "operational" },
    { "name": "Workers", "status": "operational" }
  ],
  "uptime": 68.95,
  "timestamp": "2026-10-01T15:03:20.430Z"
}

Each component reads operational, degraded or outage, and the top-level status is the worst of them. A component that is not operational says why in reason, except a Database at degraded that answered its probe slowly, which carries latencyMs instead:

Dead letters, of a queue or of webhooks, do not move /status. A failing receiver is not the platform being down, and the backlog is not for an unauthenticated page. Alert on GET /v1/admin/system-health for those.

Metrics

GET /metrics exposes Prometheus metrics on both the API and the worker process. When METRICS_AUTH_TOKEN is set, which both packaged installs do for you, /metrics on either process takes that token as a bearer and nothing else: an API key gets 401 (from the API with a recoveryHint saying an API key is not accepted there, from the worker with no body). With no token, METRICS_AUTH_REQUIRED (default true in production) decides: the API's /metrics then takes the same API-key chain as /v1, while the worker, which has no API-key chain, answers 503 rather than serve unauthenticated. Set METRICS_AUTH_REQUIRED=false to serve both openly and restrict them at the network layer. Under Compose the worker's port is not published to the host at all, and under the chart it is cluster-internal. All series are prefixed agledger_. The ones worth alerting on:

MetricWatch for
agledger_vault_integrity_check_results_total{result="broken"}Any increase: a chain, or the key registry's trust walk, failed periodic verification. Worker only
agledger_db_pool_waiting_connectionsSustained nonzero: pool saturation
agledger_pgboss_queue_size{queue=~".*-dlq",state="total"}Growth: jobs dead-lettering into a DLQ
agledger_pgboss_queue_size{state="queued"}Sustained growth on any queue: a worker is down or cannot keep up. Every queue the Server runs reports here, federation included
agledger_pg_listener_connected0 for more than a few minutes on one process: that replica does not hear cross-replica cache invalidations
agledger_pg_listener_reconnect_failures_totalIncrease: a LISTEN reconnect attempt failed
agledger_vault_checkpoint_skipped_broken_totalIncrease: a record went un-anchored. Worker only
agledger_outbound_ssrf_blocked_totalIncrease: outbound calls hitting the SSRF guard
agledger_federation_zero_row_oldest_candidate_age_secondsClimbing toward the recovery horizon: crash-orphaned records are aging out unrecovered. Federated deployments only; see Federation delivery and recovery
agledger_vault_signer_unreachable1: the KMS signing key is not answering, and this process answers 503 on every write it would sign. Always 0 with a local key

The worker serves /metrics on its health port (WORKER_HEALTH_PORT). Under Compose that port is not published to the host, and the host's port 3001 (AGLEDGER_HOST_PORT) is the API, so read the worker's series from inside its container:

docker compose exec -e NODE_OPTIONS= agledger-worker /nodejs/bin/node -e \
  "fetch('http://localhost:3001/metrics',{headers:{authorization:'Bearer '+process.env.METRICS_AUTH_TOKEN}}).then(r=>r.text()).then(t=>process.stdout.write(t))" \
  | grep agledger_vault_integrity_check_results_total

The bundled Prometheus scrapes the worker for you, as its own job. Note that agledger_vault_integrity_check_results_total and agledger_vault_checkpoint_skipped_broken_total move on the worker process only. The API's /metrics carries them too, with every series at 0 for as long as it runs, so a zero there says nothing about the chain. agledger_partition_runway_days (next section) is not on the API's /metrics at all. Scrape both processes. The federation series are worker-only too.

You do not have to write those alerts yourself. monitoring/alerts/agledger.rules.yml in the install repository ships a rule set covering silent drops, chain integrity, partition maintenance, federation delivery, and availability, and the bundled Prometheus loads them already. Read the file for the current groups and counts rather than a number quoted here: releases add rules, and every count this page has carried went stale. No Alertmanager is bundled and no routing is configured, because receivers and escalation policy are yours to decide; until you point them somewhere the rules evaluate on Prometheus' own /alerts page, and each carries a severity label of critical or warning as the routing hook. Every rule also carries a runbook_url annotation pointing at its section of monitoring/runbooks.md in the same repository: what fired, what to check, what to do, and when silencing is safe.

On Kubernetes with the Prometheus Operator, the chart renders the same rules and dashboards: monitoring.prometheusRule.enabled, monitoring.serviceMonitor.enabled and monitoring.grafanaDashboards.enabled (Grafana sidecar ConfigMaps), all off by default. monitoring.prometheusRule.runbookUrl repoints every rule's runbook link at your own copy.

Three Grafana dashboards auto-provision with install.sh --with-monitoring: an overview (traffic, chain throughput, saturation), data-integrity surveillance, and a silent-drop board with a panel per fire-and-forget path that swallows its error. monitoring/README.md describes all three, including how to import them into your own Grafana instead of the bundled one.

Watching your agents

An agent that starts behaving differently is usually the first sign of a problem, and the sign is the change itself, in either direction. An acceptance rate that moves from 0.8 to 1.0 deserves the same look as one that moves to 0.6: something changed about the work, the gate, or the agent. The Server reports the change and stops there. It does not score agents, rank them, or decide what a move means. That decision belongs to whoever is watching.

Start with the whole org. An org-scoped admin key (the admin-standard or admin-observer profile) lists every agent with what it did in the current window, the same counts for the equal window before it, and the difference:

curl -s "$AGLEDGER_API_URL/v1/agents/drift?window=7" \
  -H "Authorization: Bearer $ADMIN_KEY"
{
  "window": {
    "days": 7,
    "currentFrom": "2026-09-01T21:47:40.902Z",
    "currentTo": "2026-09-08T21:47:40.902Z",
    "baselineFrom": "2026-08-25T21:47:40.902Z",
    "baselineTo": "2026-09-01T21:47:40.902Z"
  },
  "data": [
    {
      "agentId": "01a082fb-8266-7242-8c57-98f4f2c612dd",
      "displayName": "scraper-bot",
      "current":  { "from": "2026-09-01T21:47:40.902Z", "to": "2026-09-08T21:47:40.902Z", "records": 9, "completions": 0, "verdicts": 0, "accepted": 0, "rejected": 0, "overturned": 0, "acceptanceRate": null, "medianCompletionMs": null },
      "baseline": { "from": "2026-08-25T21:47:40.902Z", "to": "2026-09-01T21:47:40.902Z", "records": 0, "completions": 0, "verdicts": 0, "accepted": 0, "rejected": 0, "overturned": 0, "acceptanceRate": null, "medianCompletionMs": null },
      "change":   { "records": 9, "completions": 0, "verdicts": 0, "overturned": 0, "acceptanceRate": null, "medianCompletionMs": null }
    },
    {
      "agentId": "01a082fb-8258-7801-bd1d-4cbef5765ec9",
      "displayName": "invoice-bot",
      "current":  { "from": "2026-09-01T21:47:40.902Z", "to": "2026-09-08T21:47:40.902Z", "records": 5, "completions": 5, "verdicts": 5, "accepted": 4, "rejected": 1, "overturned": 0, "acceptanceRate": 0.8, "medianCompletionMs": 21 },
      "baseline": { "from": "2026-08-25T21:47:40.902Z", "to": "2026-09-01T21:47:40.902Z", "records": 0, "completions": 0, "verdicts": 0, "accepted": 0, "rejected": 0, "overturned": 0, "acceptanceRate": null, "medianCompletionMs": null },
      "change":   { "records": 5, "completions": 5, "verdicts": 5, "overturned": 0, "acceptanceRate": null, "medianCompletionMs": null }
    }
  ],
  "total": 2,
  "nextCursor": null,
  "hasMore": false
}

Read the two agents side by side. scraper-bot notarizes: its records terminalize on create with no completion and no gate, so its only signal is volume, and nine records against a baseline of zero is a new agent or a new workload. invoice-bot is gated: five records, five completions, five verdicts, one rejected. acceptanceRate is accepted / verdicts, and medianCompletionMs is the median time from a record's activation to its completion being submitted. Both are null in a window with nothing to divide, and change is null whenever either side is, so a brand-new agent reads as counts that went up and rates that do not exist yet, rather than as a rate that jumped from zero.

Every count is by when the thing happened: records by creation, completions by submission, verdicts by the moment the verdict landed, disputes by resolution. verdicts is the final verdict on each completion. A principal-gated record carries the engine's structural pass and then the principal's verdict on the same completion, and only the principal's counts. A record that is revised and resubmitted is a new completion and so a new verdict, so a revision cycle shows as two verdicts with one of each outcome. The verdict count sits beside the rate so one record's cycle cannot be mistaken for ten records' worth of rejections.

window is the length in days, from 1 to 365, default 7. The current window is the last N days and the baseline is the N days before it, so the two are always comparable without scaling. The listing pages by agent with limit and cursor; a cursor is bound to the window it was minted under and refuses to continue under another.

When an agent moved, ask about that agent. The per-agent read carries the same series overall and then one per type, so a change in the roll-up can be traced to the type that carried it:

curl -s "$AGLEDGER_API_URL/v1/agents/01a082fb-8258-7801-bd1d-4cbef5765ec9/drift" \
  -H "Authorization: Bearer $ADMIN_KEY"
{
  "agentId": "01a082fb-8258-7801-bd1d-4cbef5765ec9",
  "window": { "days": 7, "currentFrom": "...", "currentTo": "...", "baselineFrom": "...", "baselineTo": "..." },
  "overall": {
    "current":  { "from": "...", "to": "...", "records": 5, "completions": 5, "verdicts": 5, "accepted": 4, "rejected": 1, "overturned": 0, "acceptanceRate": 0.8, "medianCompletionMs": 21 },
    "baseline": { "...": "..." },
    "change":   { "...": "..." }
  },
  "byType": [
    {
      "type": "principal-gate-generic-v1",
      "current":  { "from": "...", "to": "...", "records": 5, "completions": 5, "verdicts": 5, "accepted": 4, "rejected": 1, "overturned": 0, "acceptanceRate": 0.8, "medianCompletionMs": 21 },
      "baseline": { "...": "..." },
      "change":   { "...": "..." }
    }
  ]
}

?type= narrows byType to one type; overall still covers everything. The rows behind the counts are one call further: GET /v1/agents/{agentId}/history lists every record the agent acted on, newest first, with type, outcome, from and to filters, so the one rejected record above is ?outcome=reject.

An agent key reads its own drift and its own history with the same calls, and nothing else: another agent's id answers 403, and the org-wide listing answers 403 naming org-admin as the role it takes. The scope on all three routes is drift:read, which admin-standard, admin-observer and agent-full carry. A platform key reads any single agent but has no org to list, so the org-wide call refuses it and says to use an org-bound admin key.

Nothing here is cached and nothing is computed ahead of time. Each call reads the records, completions, verdicts and disputes tables for the two windows, so the answer is current to the request and there is no counter to fall behind or rebuild. There is nothing to configure and no external dependency, so it works identically air-gapped. Field-level detail for all three routes is in the API reference under the Agent Drift tag.

Signing-key rotation

Rotation is the load-bearing day-2 task. The guarantee that makes it safe:

Rotating the signing key never breaks verification of already-signed records. Retired keys stay in the published registry, so a record signed under an old key still verifies after any number of rotations. No re-signing, no downtime.

How rotation works

Rotation is two steps, and they are deliberately separate: staging the new key, then retiring the old one once nothing is signing with it.

Stage. Generate a new key with the script in the Server image (add --algorithm es256 for an ES256 key); it runs from the image you already loaded, so the step works air-gapped:

docker run --rm agledger/agledger:<version> dist/scripts/generate-signing-key.js

Set the new key as VAULT_SIGNING_KEY and the key in use as VAULT_SIGNING_KEY_PREVIOUS, then restart every api and worker process. A single-replica Helm install counts as more than one process, since the api and the worker are separate Deployments and both sign. On Helm a chart-managed Secret change rolls both by itself; with existingSecret set nothing watches the Secret, so restart both Deployments yourself (kubectl rollout restart deploy -n <ns> -l app.kubernetes.io/instance=<release>). The first process to boot registers the new key, activates it beside the key that is already active, and writes a succession statement signed by both keys. That statement is what makes the new key trusted, to every other process and to every auditor, and once it exists other processes on the new key need no VAULT_SIGNING_KEY_PREVIOUS. The boot logs the staging:

{"level":"info","keyId":"6a639248683aab56","activeSigningKeyIds":["affc2b9bfb22144e","6a639248683aab56"],"msg":"Staged vault signing key"}

A process on a new key with no predecessor signer registers nothing and signs nothing: /health reports signingKey.gate: "unanchored", /health/ready answers 503, a worker stops consuming jobs (the API's /status then reads Workers degraded with worker_not_consuming), and the rotate endpoint answers 409 with a recoveryHint. The out-of-band alternative to a predecessor is VAULT_TRUST_ANCHORS, a comma list of sha256: pins of keys whose history you vouch for: it is the recovery path when no key in hand links to that history, the process registers under a fresh genesis, and auditors then need the new key's pin.

Staging writes a KEY_ROTATED entry on the platform chain and a vault.signing_key_rotated row in system_audit_log, so a key change is on the record whether it arrived by a restart or by the API.

Staging retires nothing. Until you close the old key's window the install has two active keys, which is the normal state of a key change and costs nothing but an extra published key. A rolling restart means some processes are still holding the old key, and every process you have not restarted yet keeps signing inside its own key's published window, so nothing it writes fails verification.

POST /v1/admin/vault/signing-keys/rotate (platform key) performs the same registration on demand. After a restart the process has already registered its key, so the endpoint reports already_active. That is the expected answer: use it to confirm the staging and read activeKeys, which lists every key still able to sign with the last instant each one signed. It never retires anything.

{
  "newKeyId": "6a639248683aab56",
  "status": "already_active",
  "activeKeys": [
    { "keyId": "6a639248683aab56", "algorithm": "Ed25519", "activatedAt": "2026-09-22T14:02:11.000Z", "lastSignedAt": "2026-09-22T14:05:40.118Z" },
    { "keyId": "affc2b9bfb22144e", "algorithm": "Ed25519", "activatedAt": "2026-05-26T09:12:03.000Z", "lastSignedAt": null }
  ]
}

A key held in AWS KMS rotates the same way: staging is the new key's ARN in VAULT_SIGNING_KEY_KMS_ARN (signing.kmsKeyArn on the chart) and a restart, with the outgoing KMS key in VAULT_SIGNING_KEY_PREVIOUS_KMS_ARN (signing.previousKmsKeyArn), and the process given kms:GetPublicKey and kms:Sign on both keys for the change; the old key never leaves KMS and signs only the succession. VAULT_SIGNING_KEY and VAULT_SIGNING_KEY_KMS_ARN together refuse to boot, and so do the two predecessor variables, so moving from a local key to KMS is one rotation: put the local key in VAULT_SIGNING_KEY_PREVIOUS, set the ARN, unset VAULT_SIGNING_KEY, restart, then retire the local key id. Signing with a KMS key covers the key policy.

Retire. First confirm every process has rolled onto the new key. Read GET /health on each api and worker process directly, not through a load balancer, and check that signingKey.keyId is the new key everywhere. That is the check that proves no process still holds the old one. lastSignedAt on GET /v1/admin/vault/signing-keys is not: it proves only that a key appended to the chain recently, and reads null on an idle process or one whose only use of the key is webhook deliveries, certificates or Receipts. Then close the old key's window through a process on the new key (a process on the key being retired answers 409):

curl -X POST -H "Authorization: Bearer $AGLEDGER_PLATFORM_KEY" \
  -H "Content-Type: application/json" -d '{}' \
  "$AGLEDGER_API_URL/v1/admin/vault/signing-keys/<old-key-id>/retire"

The response carries retiredKeyId, the retiredAt instant that is now the published end of the key's window, the keys still active, and the closureDigest of the closure statement the calling process's key signs over that retiredAt. Anything the retired key admits after the closure counts for nothing. The call refuses with 422 and names the key while it has appended to the chain within the last 300 seconds; that refusal is a backstop for what the chain can see, not a substitute for the /health check. Retiring the only active key is always refused, because it would leave nothing able to sign: stage the replacement first.

After retirement, an entry written under the old key is a chain break: the Server reports it as key_expired and an offline verifier as CHAIN_KEY_EXPIRED. Entries written before it stay valid. A process that keeps holding a retired key does not write such an entry: it stops signing and answers 503 on /health/ready, so a load balancer takes it out of rotation. The fix is a restart onto the active key, and a retired key never comes back.

For a leaked key, do not wait. The compromise order retires with {"force": true} immediately after the restart, which also revokes the key's ephemeral certificates and can leave other keys untrusted until you pin them. Follow the key-compromise runbook for that order rather than this one.

Remove VAULT_SIGNING_KEY_PREVIOUS (or VAULT_SIGNING_KEY_PREVIOUS_KMS_ARN) once the old key is retired, and roll again at your convenience. Finish a key change before a version upgrade starts, and do not start one during a rollout.

Reading the published keys

After retirement both keys appear at GET /v1/verification-keys: the new one active, the prior one retired with the exact instants it was active (activatedAt / retiredAt, the values the key statements sign), still resolvable:

curl -s "$AGLEDGER_API_URL/v1/verification-keys" \
  | jq -r '.data[] | "\(.keyId) \(.status) \(.activatedAt) \(.retiredAt)"'
6a639248683aab56 active 2026-09-22T14:02:11.000Z null
affc2b9bfb22144e retired 2026-05-26T09:12:03.000Z 2026-09-22T14:31:52.604Z

The one exception is a key VAULT_DISTRUSTED_KEYS names with an instant earlier than its signed closure: there retiredAt is that distrust instant, since nothing the key signed after it counts, and the entry carries distrustedFrom with the instant. Nothing signs that field, so it tells an auditor what to ask the operator to confirm, not that the window is untampered.

That surface publishes anchored keys only, with anchoredFrom, the sha256: digest of the key the serving process signs with. It equals the pin the installer printed only until the first rotation; after one it is the new key's digest. The installer's pin stays a valid trust anchor across routine rotations, because each new key is admitted by a signed statement from the key before it, so a verifier given that pin walks the statements forward to the current key. A retirement with {"force": true} cuts that walk, and its response says which pin to hand auditors instead.

Proving the guarantee

Records signed before the rotation must still verify, and you can check that rather than take it on trust. Dump the vault and verify it offline by the off-box verification in the audit runbook, with @agledger/verify 2.0.0 or later (a 2.0 chain needs that floor, published per key as minVerifierVersion). After a rotation the dump carries both signing keys and entries signed by each, and the verifier resolves whichever key signed each entry, so a clean run covers the retired key as well as the active one.

Partition maintenance

Four high-volume tables are range-partitioned by month, and only system_audit_log carries a DEFAULT catch-all partition. audit_vault, events and webhook_deliveries have none, so a write past the latest pre-created partition on any of those three has nowhere to land and fails outright: runway exhaustion there is a write outage, not an overflow. The worker pre-creates upcoming partitions ahead of the clock and exposes runway as a gauge (agledger_partition_runway_days) for all four. You can also query the source function directly:

psql "$DATABASE_URL" -c "SELECT table_name, runway_days, default_rows FROM partition_runway();"
     table_name     | runway_days | default_rows
--------------------+-------------+--------------
 audit_vault        |         570 |            0
 events             |         570 |            0
 webhook_deliveries |         570 |            0
 system_audit_log   |          83 |            0

runway_days is days until the latest pre-created partition is reached, and it is worth alerting on low for all four tables. default_rows is only meaningful on system_audit_log, the one table with a DEFAULT partition for a row to land in: a nonzero value there means writes are landing in it and the worker is falling behind, so alert on default_rows > 0 for that table specifically. On the other three it stays 0 by construction, since there is no DEFAULT partition to reach, so that alert would never fire and runway_days is what catches trouble instead.

Config-as-code hot reload

If you run with PROVISIONING_CONFIG_PATH set, orgs, agents, webhooks, and contract schemas are declared in YAML and reconciled on every boot; see the provisioning runbook for the directory layout. Reload changes without a restart via SIGHUP or POST /v1/admin/provisioning/reload (platform key). Check current state first:

curl -s -H "Authorization: Bearer $AGLEDGER_PLATFORM_KEY" "$AGLEDGER_API_URL/v1/admin/provisioning/status"
{"configured":true,"configPath":"/etc/agledger/provisioning","dryRun":false,"prune":false,"lastReloadAt":"2026-06-09T15:59:53.404Z","managed":{"orgs":1,"agents":2,"webhooks":0,"schemas":2,"trustedIssuers":1},"loadErrors":["webhooks/acme.yaml: Environment variable ACME_WEBHOOK_SECRET is not set and has no default"],"loadWarnings":[],"pruneSuppressed":false,"trustedIssuersWarnings":[],"trustedIssuersUnchangedSinceLastLoad":true}
curl -s -X POST -H "Authorization: Bearer $AGLEDGER_PLATFORM_KEY" "$AGLEDGER_API_URL/v1/admin/provisioning/reload"
{
  "orgs":    { "created": 0, "updated": 1, "pruned": 0 },
  "agents":  { "created": 0, "updated": 2, "pruned": 0 },
  "schemas": { "created": 0, "updated": 0, "pruned": 0 },
  "apiKeys": { "created": 0, "skipped": 3, "generated": [] },
  "errors": [
    { "resource": "config", "error": "webhooks/acme.yaml: Environment variable ACME_WEBHOOK_SECRET is not set and has no default" }
  ]
}

Reload is idempotent, and the counts read differently by resource. Every org and agent the files declare that already exists counts as updated, changed or not; a webhook or schema counts as updated only when the reload changed it, so an unchanged one counts nowhere. Existing keys count as skipped. Reload is also fail-open: a single invalid file (here, an unset ${ACME_WEBHOOK_SECRET} substitution in a webhook file of your own) is reported in errors[] while every valid resource still applies. Newly minted keys appear in apiKeys.generated[] with their raw value exactly once, in that response body, so capture them then.

Fail-open is the shape to watch. A file that fails to parse is skipped whole: its resources are silently absent, the rest of the reconcile succeeds, and the Server comes up healthy. status re-reads the directory on every call and lists every unloadable file in loadErrors[], alongside the errors from the last trusted-issuers pass this process ran, so that field, not the boot log, is the durable signal. It also moves status on system health. The agledger_provisioning_errors gauge (labels stage="load" / stage="reconcile") carries the same state for alerting and stays non-zero until a clean reconcile clears it.

Three more fields on status answer the questions a clean-looking reload leaves open. pruneSuppressed is true when prune is on and the directory does not load cleanly: a reload then prunes nothing, so an org or agent deleted from the YAML to offboard it keeps its API keys active until loadErrors is empty. trustedIssuersWarnings lists issuers whose OIDC discovery failed or whose jwks_uri the egress guard refused; each kept the jwks_uri it already held while every other field applied. trustedIssuersUnchangedSinceLastLoad: true after a reload that changed nothing means the pod read the same trusted-issuers.yaml bytes as before: on Kubernetes, a ConfigMap mounted through subPath never receives updates, so mount it as a directory.

A bare ${VAR} with no default is the usual cause: it is a hard parse error when the variable is unset. Use ${VAR:-default} wherever a sensible default exists, and inject real secrets through the pod environment (extraEnv / extraEnvFrom on the chart).

Version upgrades

Upgrade with scripts/upgrade.sh from the agledger-ai/install repository; the install runbook has the full procedure and the air-gapped path. Migrations run automatically and are advisory-locked and checksum-verified, so a migration runs once and only once even across replicas, and a tampered or reordered migration is refused rather than applied. For diagnostics to attach to a support request, scripts/support-bundle.sh collects health, version, container state and logs, and .env with credential values redacted by the shape of the name and the value. Logs are collected as written, so review the archive before you send it.

Restart the process after swapping the image. Already-signed records verify across versions, because the chain format is stable and historical keys resolve.

What each release migrates, and whether its upgrade needs a maintenance window, is in that release's notes in the agledger-ai/install changelog. Read the notes for every release you are crossing before you upgrade.

2.0 and a 1.x database

2.0 does not upgrade a 1.x database. The migration refuses a database a 1.x release migrated before it changes anything, and says so: the checksum it recorded for 001_consolidated.sql is not the one 2.0 ships. Install 2.0 against a new, empty database, and keep the 1.x database for the 1.x install that wrote it. A backup does not carry 1.x data across either: restore.sh returns the install to the release the backup was taken on, --keep-version with a 1.x archive is refused before anything is stopped or dropped, and an archive that records no version reaches the 2.0 migration, which refuses the restored database the same way.

upgrade.sh refuses an install whose recorded or running version is 1.x at its version check, before it writes .env, takes a backup or stops anything. Put the checkout back on the 1.x deploy/ tree afterwards: a docker compose command run from the 2.0 tree starts 2.0 against the 1.x database. When the version was not recorded and the migration is what refuses, the recovery text names that tree rather than a re-run.

On Helm with the bundled PostgreSQL, a 2.0 chart refuses to render over a release a 1.x chart installed, so helm upgrade fails before it replaces a pod or rewrites the release Secret, and the 1.x release keeps serving. Install 2.0 as a new release, which gets its own bundled database. The check reads the running release through lookup, which answers nothing under a renderer (helm template, Argo CD), so there it cannot refuse: do not sync a 2.0 render onto a 1.x bundled release, because the migration refuses the database only after the pods are replaced, and helm rollback does not restore the release Secret the upgrade rewrote. With an external database the migrate Job is a pre-upgrade hook, so a refused upgrade fails before any 1.x pod is replaced.

Two releases against one schema

Within 2.x, every upgrade migrates the database before the new processes exist, so from the moment the migration commits until the last process of the old release is gone, two releases serve one schema. Under the chart's default Recreate strategy that window is the pod swap; under RollingUpdate it is the whole rollout. Migrations are written to serve the previous release, so the window needs no ordering. What it does bind is a security control a release introduces (such as subjectAllowlist or jtiSingleUse on a trusted issuer): only processes carrying that release enforce it, so it binds every request once connectedVersions on GET /v1/admin/system-health holds one entry. helm rollback takes the processes back and leaves the schema where it is, which is safe for the same reason; a worker on the older release fails jobs from schedules the newer release added and counts them on agledger_maintenance_unknown_task_total until you roll forward again. The install repository's README covers the same window under "Two releases against one schema".

The runtime role

On an external database, the upgrade checks the runtime role twice: before the pre-upgrade backup, and again after migrations. That role can stop being able to serve without the install changing. A credential-rotation policy that drops and recreates it brings it back without its agledger_app membership, because role membership does not survive a DROP ROLE. Left unchecked, the first thing to fail is pg_dump during the backup, which reports the table it was refused and prints its whole LOCK TABLE statement without naming the role.

Each check names the role and the exact GRANT. Stopping at the first one costs nothing: no backup has been taken, no migration has run, and neither the version nor the image digest in compose/.env has moved, so the running install is serving exactly as it was. Stopping at the second one leaves the upgrade part-done, which it says: the migrations are applied, the worker stays stopped until the upgrade finishes, and the previous version is still serving against the new schema. Either way, applying the grant and re-running finishes it, and migrations already applied are skipped.

Migration lock budget under sustained writes

A migration that cannot take a table's lock within MIGRATION_LOCK_TIMEOUT (default 30s) rolls back whole with lock_not_available, so running it again is safe.

Request timeouts

CONNECTION_TIMEOUT_MS (default 60000) is the socket inactivity timeout: how long a connection may go without bytes moving before the server destroys it.

The ordering between the three timeout knobs is what matters. Keep it above HANDLER_TIMEOUT_MS (default 30000) and below KEEP_ALIVE_TIMEOUT_MS (default 72000). If the socket timeout fires first, a slow request is severed underneath the handler instead of failing through the route, and the failure mode is unpleasant: on a bulk write the batch commits, the client sees a dead socket with no response, and it has no way to tell a committed batch from a lost one.

Recovering from a severed socket: retry the request with the same Idempotency-Key. A commit that already landed is returned rather than repeated.

Two signals tell you the value is too low for your traffic, and you need both because a severed socket never reaches the normal response path:

Raise CONNECTION_TIMEOUT_MS if you see them on legitimately slow bulk work, keeping it under the keep-alive value.

Federation delivery and recovery

Skip this section if you run a single Server with no peers. Nothing here applies until federation is configured, and the recovery sweeps below short-circuit before they touch the records table when a Server has no active peers.

Is this peer reachable?

GET /federation/v1/admin/peers answers it. Read lastDeliveryAt (the last outbound message that got a 2xx), consecutiveDeliveryFailures (reset to 0 on a success) and lastDeliveryError (cleared on the next success). A frozen lastDeliveryAt with a climbing failure count and a transport error is a peer that has gone away, with retries still in flight until the job dead-letters.

One field on the same object does not answer it, and it is worth knowing why: status is registration state (active / revoked), not reachability. A peer you have never successfully reached still reads active.

For a count rather than a list, GET /v1/admin/ops-summary reports federation.peers with the registration counts and, partitioning the active ones, delivering / failing / neverDelivered. That last split is the one a dashboard needs: active alone reads the same for a peer taking every delivery and a peer that has never taken one.

For the backlog behind those deliveries, watch agledger_pgboss_queue_size{queue="federation-outbound"} or read GET /v1/admin/system-health, which reports the same queues as JSON, and whose status goes degraded once those retries give up and dead-letter.

Expect peers to lag, and size for it

POST /v1/records returns as soon as the record is notarized locally; peers receive their signed copy afterwards. Your local chain is complete and verifiable the entire time a peer is behind, so peer lag is a delivery concern, never an integrity one.

The outbound worker drains at a per-pod ceiling of FEDERATION_OUTBOUND_CONCURRENCY * FEDERATION_OUTBOUND_BATCH_SIZE / FEDERATION_OUTBOUND_POLLING_INTERVAL_SECONDS. On the defaults that is 16 * 4 / 1.0 = 64 legs/sec/pod. A leg is one record to one peer, so a record shared to two peers costs two legs. Notarizing 4,100 shared records against two peers queues 8,200 legs and puts peers roughly two minutes behind. Ingest is faster than drain by design, so a bulk share always builds a queue.

Raise batch size first if you need more throughput on a backlog spread across many peers and records. It widens the fetch each worker loop already issues, so it lifts the ceiling without adding database round-trips. It does nothing for a backlog concentrated in a handful of groups, though: the worker pairs groupConcurrency: 1 with a (peer, record) group id to keep same-record delivery ordered, which caps every batch at one job per group, so a larger batch size pulls nothing extra and concurrency or polling interval is the lever instead. Concurrency and polling interval are the expensive pair and interact super-additively on database CPU: each in isolation costs roughly 30 to 40 percent on POST /v1/records p50/p95, but together at (32 / 0.5) they cost about +110 percent p50 against the (16 / 1.0) baseline. Re-measure record-creation latency after touching any of the three, and prefer scaling out worker pods over cranking a single pod far past the defaults.

Same-(peer, record) ordering is preserved regardless of these settings, so raising throughput never reorders state transitions for a given record.

The recovery sweep, and the window that abandons records

Outbound work is enqueued after the record's state change commits. A crash in that gap can leave a committed terminal record with no federation work queued at all. A sweep runs every two minutes on the worker process to find those and re-drive them through the normal publish path: same share gate, same peer set, same idempotency, so recovery can never deliver something the live path would not have.

Two knobs size it, and they are a pair:

KnobDefaultWhat it controls
AGLEDGER_FEDERATION_ZERO_ROW_HORIZON_MINUTES360 (6h)How far back the sweep looks
AGLEDGER_FEDERATION_ZERO_ROW_BATCH_SIZE50Records re-driven per cycle (~25/min)

The horizon is an abandonment boundary, not just a scan bound. A record whose updated_at falls outside the window is never recovered. Six hours covers a crash gap with wide margin, since the normal case recovers within a sweep interval or two. Widen it only if this Server can be down longer than that.

Each cycle takes the oldest candidates first, which is the only order that cannot starve the record closest to aging out.

The one number to alert on is agledger_federation_zero_row_oldest_candidate_age_seconds, the age of the head of that queue. Flat and low is healthy. Climbing toward the horizon means the candidate pool is refilling faster than one batch drains, which is the only remaining way a genuinely orphaned record ages out unrecovered. Page well before the horizon.

Which knob to turn when that age climbs: raise AGLEDGER_FEDERATION_ZERO_ROW_BATCH_SIZE. Widening the horizon in that state makes it worse, because it adds candidates to a pool that is already draining too slowly. The horizon is the knob for a different problem, a Server that is legitimately offline longer than six hours.

Records received from peers

A record this Server received from a peer is a read-only view of a row the originating Server owns, and only that Server fans it out. Such records are never recovery candidates: recovering them would deliver redundantly back to the origin and be permanently rejected by third peers.

This matters operationally because a busy receiving Server accumulates those records without bound. agledger_federation_zero_row_projections_excluded reports how many were held out of the candidate pool on the last cycle, capped at the batch size. Non-zero is the healthy reading on any Server receiving federation traffic; it is the exclusion doing its job, not a problem.

Its companion agledger_federation_outbound_projection_skipped_total is a last-resort guard further down the publish path and is expected to stay at 0 permanently. Treat any increase there as a regression to report, not as confirmation that the exclusion is working.

All of these series are exposed by the worker process, not the API. Scrape both.