Audit & Verification
A Server exposes two separate audit surfaces. They answer different questions, and an auditor treats them differently. Read this page as two runbooks under one cover.
| Surface | What it holds | Who reads it | Question it answers |
|---|---|---|---|
system_audit_log | Operational events - orgs created, records written, keys rotated, admin reads | Your SIEM / SOC, continuously | "What is happening on this Server?" |
audit_vault | The signed, hash-chained records themselves | An auditor, offline | "Is this record authentic and unaltered?" |
The first surface is operational telemetry: it tells you the Server is behaving. The second is the proof: it stands on its own cryptography, so an auditor can verify the chain without trusting - or even reaching - the Server that produced it. Stream the first into your SIEM. Hand the second to an auditor. Do not substitute one for the other.
Surface 1 - Stream operational events to your SIEM
Two channels carry this feed, and they carry it identically. Pull it yourself from
GET /v1/siem/stream, or configure the Server to push it to a file or an HTTP collector. Both read
the same durable rows through the same projection, so one event is the same bytes on either channel
and carries the same id. A collector reading both collapses the pair on that id instead of
ingesting the event twice.
Pull
Poll GET /v1/siem/stream on an interval and forward the result to your SIEM. The endpoint merges
the event stream and the operational system_audit_log into one time-ordered feed, served as
application/x-ndjson. It requires a key with the audit:read scope (see the
API reference for scope details).
Parameters: since (ISO-8601, starts a walk), cursor (continues one), limit (default 100,
maximum 1000), format (ocsf default, or raw). Send since or cursor, never both.
curl -s -H "Authorization: Bearer $AGLEDGER_API_KEY" \
"$AGLEDGER_API_URL/v1/siem/stream?since=2026-05-01T00:00:00Z&limit=5&format=raw"
{"type":"admin.org_bootstrapped","payload":{"name":"Default","reason":"single-org-install","actorId":"00000000-0000-0000-0000-000000000000","actorRole":"platform","targetType":"org","targetId":"019ead17-fbd4-7381-ac3f-5ee1474830f1"},"timestamp":"2026-06-09T15:54:50.707Z","id":"019ead17-fbd4-7d84-8941-ff140a540fc7"}
{"type":"schema.registered","payload":{"type":"notarize-generic-v1","orgId":"019ead17-fbd4-7381-ac3f-5ee1474830f1","version":1,"category":"general","publisher":"local","compatibilityMode":"backward","fieldMappingCount":0,"actorId":"00000000-0000-0000-0000-000000000000","actorRole":"platform","targetType":"schema_subject","targetId":"019ead17-fbf0-72b2-8817-44467a2ede3d"},"timestamp":"2026-06-09T15:54:50.734Z","id":"019ead17-fc09-7dcd-b152-b238559c6853"}
Set format=ocsf to emit OCSF 1.4.0 events. Splunk takes them directly. Microsoft Sentinel, IBM
QRadar, CrowdStrike and Elastic read the feed through a collector you run in between (Logstash,
Vector, Fluent Bit, or for Sentinel an Azure data collection rule): the Azure Monitor Logs Ingestion
API wants a JSON array body and a Microsoft Entra bearer, and Elasticsearch _bulk wants an action
line before every document, so neither takes the push sink as it is. metadata.uid carries the same
durable row id that the raw shape puts in id, so a correlation rule written against one format
addresses the same event under the other. time is the instant as epoch milliseconds and time_dt
the same instant as RFC 3339 text, which metadata.profiles declares with datetime. Account
Change (class 3001) carries no entity: the payload rides under unmapped, and a credential the
event names goes to user.credential_uid.
curl -s -H "Authorization: Bearer $AGLEDGER_API_KEY" \
"$AGLEDGER_API_URL/v1/siem/stream?since=2026-05-01T00:00:00Z&limit=1&format=ocsf"
{"metadata":{"product":{"name":"AGLedger","vendor_name":"AGLedger","version":"2.0.0"},"version":"1.4.0","log_name":"audit","uid":"019ead17-fbd4-7d84-8941-ff140a540fc7","profiles":["datetime"]},"time":1781020490707,"time_dt":"2026-06-09T15:54:50.707Z","severity_id":3,"class_uid":3001,"category_uid":3,"type_uid":300101,"activity_id":1,"status_id":1,"message":"Org bootstrapped: Default","actor":{"user":{"uid":"00000000-0000-0000-0000-000000000000","type":"platform"}},"user":{"uid":"019ead17-fbd4-7381-ac3f-5ee1474830f1","type":"org"},"unmapped":{"name":"Default","reason":"single-org-install","actorId":"00000000-0000-0000-0000-000000000000","actorRole":"platform","targetType":"org","targetId":"019ead17-fbd4-7381-ac3f-5ee1474830f1"}}
To run a continuous poller, send since once and then follow the cursor: every response carries an
X-AGLedger-Stream-Cursor header, and sending that value back verbatim as ?cursor= resumes
exactly after the last row you were given. Do not advance since to the time of the last event
you forwarded. Every row a transaction writes carries the same created_at, so a time-only
boundary cannot address a position inside that group, and each pass would drop the rest of it. The
cursor carries the row id alongside the instant and steps through. Keep limit modest.
The stream defers rather than skips. A row written by a transaction that has not committed yet is held back from the current page and arrives on a later one, so an empty page is not proof that the feed has caught up. Every page says by how many seconds it stops short of now for that reason:
x-agledger-stream-cursor: 2026-09-16T00:46:38.063676Z_01a0a7ae-11fe-73e6-9ffe-68863953e32a
x-agledger-stream-holdback-seconds: 0
A 503 means the Server could not establish that boundary and refused to serve a page rather than
walk past rows it cannot prove are visible. It carries no holdback header, because it served no
page, and it carries Retry-After. Retry with the same cursor, which is unaffected.
Push
The same feed, driven by the Server instead of by your poller. Turn it on with SIEM_ENABLED=true
and choose a sink: a local NDJSON file (SIEM_FILE_ENABLED, SIEM_FILE_PATH) or an HTTP collector
(SIEM_HTTP_ENABLED, SIEM_HTTP_URL, SIEM_HTTP_AUTH_HEADER). SIEM_FORMAT selects ocsf or
raw for both sinks. SIEM_FLUSH_INTERVAL_MS is the poll interval (default 5000, floored at 1000),
and SIEM_BATCH_SIZE the rows per batch (default 50). The file sink is on by default once
SIEM_ENABLED=true, so set SIEM_FILE_ENABLED=false when you want the HTTP sink alone.
SIEM_ENABLED=true
SIEM_FORMAT=raw
SIEM_FILE_ENABLED=false
SIEM_HTTP_ENABLED=true
SIEM_HTTP_URL=https://collector.internal/agledger
SIEM_HTTP_AUTH_HEADER=Bearer <collector token>
The push channel defers behind an in-flight transaction the way the pull channel does, but only for
SIEM_MAX_HOLDBACK_SECONDS (default 60). Past that it also delivers the committed rows newer than
the held-back ones, so a long transaction such as a backup's pg_dump does not stall the live tail,
and its cursor stays behind the held-back rows until they commit. 0 restores unbounded deferral.
The pull channel has no such cap, because the caller holds the cursor.
The worker is the process that drains both push sinks; the API serves the pull route and polls nothing. The two sinks advance one shared cursor row and each process writes its own sink, so a file sink is the complete feed only under a single worker replica. Past that, each replica's file holds a slice and nothing in either file says so, so use the HTTP sink.
SIEM_HTTP_URL is operator configuration, so the egress guard on it refuses only the cloud
instance-metadata service, whether the URL names that address or a hostname resolving to it.
collector.internal above, a sidecar on localhost, or any collector on a private range is reached
with no SSRF_ALLOW_CIDRS entry. That allowlist governs the URLs API callers register (webhooks,
federation peers, trusted-issuer JWKS), so widening it never changes what the SIEM sink may reach,
nor the reverse.
The HTTP sink POSTs application/x-ndjson batches under the default SIEM_HTTP_MODE=ndjson, which
is what Logstash, Vector, Fluent Bit and Splunk's /services/collector/raw?sourcetype=_json take.
SIEM_HTTP_MODE=hec wraps each event in a Splunk HTTP Event Collector envelope and posts to
/services/collector/event, appending that path when the URL names none; that endpoint reads the
envelope's own time and applies no line truncation, and it answers 400 No data to an unwrapped
NDJSON body. SIEM_HTTP_HEC_SOURCETYPE (default _json), SIEM_HTTP_HEC_INDEX and
SIEM_HTTP_HEC_SOURCE set the envelope fields; the last two are left to the token's defaults when
empty. Under ndjson mode against the raw endpoint, Splunk reads the timestamp out of the line only
with a props.conf stanza for the sourcetype; without one every event carries index time:
[agledger:ocsf]
INDEXED_EXTRACTIONS = json
TIMESTAMP_FIELDS = time_dt
TIME_FORMAT = %Y-%m-%dT%H:%M:%S.%3NZ
TZ = UTC
TRUNCATE = 0
Splunk's shipped TRUNCATE is 10000 bytes and cuts a longer line at the byte, which reaches the
index as unparseable JSON. SIEM_MAX_EVENT_BYTES (default 8192, floor 1024) keeps a line under that
by trimming the copied payload largest-key-first; a trimmed line names the keys it dropped under
agledgerPayloadTruncated, so a short payload and a trimmed one stay distinguishable.
For transport, SIEM_HTTP_HEADERS adds request headers as a JSON object merged over
Authorization ({"X-Splunk-Request-Channel":"<uuid>"} for HEC indexer acknowledgement);
SIEM_HTTP_GZIP=true gzips the body and sends Content-Encoding: gzip, which Splunk HEC accepts on
both collector endpoints; and SIEM_HTTP_CA_FILE names a PEM bundle for a collector behind a
private CA, replacing the system roots for this sender, so include every CA the chain needs.
NODE_EXTRA_CA_CERTS also works, process-wide, and must be set before the process starts.
The file sink needs one more step on a stock Compose install. The image runs under the Node
permission model with no filesystem write grant, so set
NODE_OPTIONS=--allow-fs-write=/var/log/agledger on the agledger-worker service, the one process
that writes this file, and mount a writable volume there owned by uid 65532. Without both, the
worker refuses to start and names the variable and the path in one error. The API serves the pull
channel and touches no file, so it needs neither.
Set that grant on the worker service rather than in .env. A value in .env is read by the API
service too, and by every one-off docker compose run and docker compose exec the shipped scripts
make against it, and a write grant has nothing to do in a process that writes no file. NODE_OPTIONS
reaches every node process in the container, and Node refuses --allow-fs-write from a process that
was not started with --permission; the shipped command and healthcheck both carry it, but a
healthcheck of your own that runs node has to carry --permission as well, or the probe cannot
start and the container never reports healthy while the sink writes fine. On Helm the equivalent
lever is permission.extraAllowArgs, which widens the sandbox argv itself rather than going through
NODE_OPTIONS, and it applies to every container in the release.
What differs between the channels is which rows each may see, never how a row is rendered. A
pull caller sees what its role scopes it to, so an org key sees its own records and an agent key
sees records where it is performer or principal, and within that scope it is served every
system_audit_log row. The push poller forwards on behalf of the install with no role scope, but it
restricts the system_audit_log branch to a fixed list of event types. The five it withholds today
are the four provisioning reconcile events, which fire on every process boot and would make a
rolling restart read as a config change, and agent.reference_added, which the event stream already
carries. A type not on that list never reaches a rule in your SIEM through the push sink; the pull
route serves it.
The feed is operational telemetry, not the tamper-evident proof. A record that appears here has not been independently verified by appearing here; that is the job of Surface 2.
Surface 2 - Verify the chain offline (the real audit)
The audit of record is performed off the Server, against only the published public keys. The verifier has no database, no network, and no AGLedger engine in its dependency tree - so it remains trustworthy even if the Server that produced the chain is later compromised. This is the default posture for a serious audit, and it is fully air-gapped.
Three steps: publish the keys, produce a dump, verify it.
Step 1 - Publish the verification keys
The public signing keys are served unauthenticated and always on. An auditor needs only these.
curl -s "$AGLEDGER_API_URL/v1/verification-keys" | jq -c 'del(.signatureInputTemplate)'
{"data":[{"keyId":"c4dd3e20388b594d","algorithm":"Ed25519","coseAlgorithm":-8,"minVerifierVersion":"2.0.0","publicKey":"MCowBQYDK2VwAyEAo95XH8DQ6ZYqhC761LqlCq0b9wxYgHPyHs67OkQ9Frw=","publicKeyRaw":"o95XH8DQ6ZYqhC761LqlCq0b9wxYgHPyHs67OkQ9Frw=","status":"active","activatedAt":"2026-06-09T17:22:41.108Z","retiredAt":null}],"envelope":"COSE_Sign1","payloadFormat":"application/vnd.in-toto+cbor","canonicalization":"RFC8949-CDE","coseAlgorithm":-8,"signatureAlgorithm":"Ed25519"}
The jq filter drops signatureInputTemplate, the prose recipe for the signing input that Step 3
follows. More than one key can be active at once while a key change rolls through the processes,
so resolve each entry's key by keyId and read algorithm and coseAlgorithm per key; the
document-level pair describes only the most recently activated one. minVerifierVersion is the
oldest @agledger/verify that can verify entries under that key.
The same key set is also published at GET /.well-known/agledger-vault-keys.json. Retired keys
stay in the set with the exact instants they were active (activatedAt / retiredAt), so records
signed before a rotation still verify against the key that actually signed them. An entry's write
time is the database's clock at the insert, set by a trigger rather than by whoever writes the row,
so an entry cannot be dated back into the window of a key that has since been retired. The key
registry is part of the dump in Step 2, so the auditor never has to ask which key signed what. The per-key
algorithm is what lets a chain whose history spans a rotation between algorithms verify entry by
entry.
Step 2 - Produce a dump
scripts/vault-dump.sh - from the agledger-ai/install
repository - is the one component that touches Postgres. It runs the dump tool that already ships
inside the Server image (no source checkout, Node.js, or pnpm on the host), writing a self-contained
set of NDJSON files the verifier consumes. Run it against a live install, then hand the directory to
the auditor.
./scripts/vault-dump.sh ./dump
{
"outDir": "/dump",
"orgId": null,
"counts": {
"audit_vault": 30,
"vault_checkpoints": 0,
"vault_signing_keys": 1,
"vault_key_statements": 1,
"org_admin_reads": 2,
"org_admin_reads_checkpoints": 0
}
}
For a Helm or other non-Compose install, run the shipped tool in a dedicated one-off pod rather than
kubectl exec into the serving API pod: that pod's /tmp is a 64Mi emptyDir, which a populated
vault does not fit, and the dump would compete with live traffic for its limits. The fsGroup makes
the emptyDir writable for the image's non-root uid, and passing command bypasses the image's
--permission argv, which otherwise denies the writes the dump needs.
kubectl run agledger-vault-dump --restart=Never --image=agledger/agledger:<tag> \
--overrides='{"spec":{"securityContext":{"fsGroup":65532},"containers":[{"name":"dump","image":"agledger/agledger:<tag>","command":["sh","-c","/nodejs/bin/node dist/scripts/dump-vault.js /dump && sleep 1800"],"envFrom":[{"configMapRef":{"name":"<release-fullname>"}},{"secretRef":{"name":"<release-fullname>"}}],"volumeMounts":[{"name":"dump","mountPath":"/dump"}]}],"volumes":[{"name":"dump","emptyDir":{}}]}}'
kubectl cp agledger-vault-dump:/dump ./dump
kubectl delete pod agledger-vault-dump
Pass --org <id> to scope the dump to a single org, which leaves out the platform-ops chain. The
output directory holds one file per entry in the counts block:
audit_vault.ndjson the per-record hash chains
vault_checkpoints.ndjson periodic signed checkpoints over the chains
vault_signing_keys.ndjson the public-key registry, with rotation history
vault_key_statements.ndjson the signed statements that admit each key to the registry
org_admin_reads.ndjson the cross-party admin-read log
org_admin_reads_checkpoints.ndjson signed tree heads over the read log
Every row carries a chain_key naming the chain it belongs to: a record id for a record chain,
schema:<orgId> for an org's schema-registration chain, and the all-zero uuid for the platform-ops
chain that holds acts with no record of their own, such as API-key creation. All of them are in the
dump and all of them are verified in Step 3.
The dump is database-independent. Keep a copy alongside your database backup - it is the artifact an auditor verifies, and it does not depend on a live Server to be meaningful. (See the backup runbook for where this fits in a backup schedule.)
Both audit_vault and vault_checkpoints are append-only at the database. Outside an explicit
escape hatch, a DELETE or UPDATE against a row of either raises an exception, and so does a TRUNCATE
or a DROP TABLE on one of audit_vault's partitions - a sql_drop event trigger catches that DDL,
since DDL skips row-level triggers. Postgres lets only a superuser create an event trigger, so on a
managed database whose migrate role has no superuser-equivalent grant the partition-drop guard is
absent and the row-level and TRUNCATE guards are what remain. Checkpoints are the out-of-band high-water mark that makes
terminal truncation detectable at all, since a chain truncated from the end still hash-links
cleanly, so the witness carries the same immutability as the ledger it protects. Each checkpoint
also names which chain it anchors (record, schema, or admin), a keying committed inside the
signed checkpoint payload itself, so it cannot be changed without breaking the checkpoint's
signature.
The hatch is SET agledger.allow_audit_drop = 'on' in the same transaction, and its effect
differs by table. On vault_checkpoints it lets a row DELETE or UPDATE through: the trigger
returns the row unchanged instead of raising, so the statement succeeds. On audit_vault a DELETE
or UPDATE still does not apply under the hatch: the row-level trigger returns NULL, which skips the
row rather than permitting it, so the row survives and the statement reports zero rows affected.
What the hatch opens on audit_vault is the DDL path - TRUNCATE and DROP TABLE on a partition
succeed under it - which is the documented partition-archival route, meant to run in the same
transaction after a checkpoint already covers the positions being archived. Automation reaching for
the hatch to edit or delete a single audit_vault row sees it silently do nothing, not fail loudly.
The engine prunes nothing on its own: there is no retention job over the chain and no retention
setting. audit_vault, events, webhook_deliveries and system_audit_log are partitioned by
month and grow for the life of the install; sizing and archival are your own tooling, through the
hatch above. There is no erasure endpoint either: a record, its chain entries and its events cannot
be deleted through the API. In encrypted mode, destroying the customer-held key changes nothing on
the Server - reads keep returning 200 with the sealed envelope verbatim, chainIntegrity stays
true, and criteria, always Server-plaintext, stays readable. What is lost is the ability to
re-derive evidenceHash from cleartext nobody can produce any more.
Step 3 - Verify with stock libraries
The verification needs no AGLedger software. Each row of audit_vault.ndjson carries its canonical
cose_sign1 envelope (RFC 9052 COSE_Sign1 over an in-toto Statement, signed Ed25519);
vault_signing_keys.ndjson carries the public-key registry. Decode each envelope with any stock
COSE library - go-cose, coset (Rust), or pycose - and verify its Ed25519 signature against the
key resolved by signing_key_id. The Sig_structure is constructed per RFC 9052 §4.4; the
signatureInputTemplate field at /v1/verification-keys documents it exactly.
This example uses Python with cbor2 and cryptography - neither of them ours - to walk the whole
dump. Save it as verify-dump.py:
import json, base64, sys, cbor2
from cryptography.hazmat.primitives.serialization import load_der_public_key
from cryptography.exceptions import InvalidSignature
keys = {k["key_id"]: k["public_key"]
for k in map(json.loads, open(sys.argv[1] + "/vault_signing_keys.ndjson"))}
ok = fail = 0
for row in map(json.loads, open(sys.argv[1] + "/audit_vault.ndjson")):
cose = base64.b64decode(row["cose_sign1"])
protected, _unprotected, payload, signature = cbor2.loads(cose).value
pub = load_der_public_key(base64.b64decode(keys[row["signing_key_id"]]))
sig_structure = cbor2.dumps(["Signature1", protected, b"", payload]) # RFC 9052 §4.4
try:
pub.verify(signature, sig_structure); ok += 1
except InvalidSignature:
fail += 1; print("FAIL pos", row["chain_position"], row["record_id"])
print(f"[{'PASS' if not fail else 'FAIL'}] stock-library offline verification")
print(f" audit_vault entries : {ok + fail}")
print(f" signatures verified : {ok}")
print(f" failures : {fail}")
print(f" signing keys : {len(keys)}")
sys.exit(1 if fail else 0)
python3 verify-dump.py ./dump
[PASS] stock-library offline verification
audit_vault entries : 9
signatures verified : 9
failures : 0
signing keys : 1
The audit_vault row count covers the schema-registration and platform-ops chains alongside your
record chains.
The script exits non-zero on any signature failure, so it drops straight into a CI gate. A few
stricter COSE libraries refuse AGLedger's vendor-private header labels by default - for pycose,
decode with Sign1Message.from_cose_obj(..., allow_unknown_attributes=True); go-cose and coset
accept them as-is. The full library-quirk notes live under "Offline cryptographic verification" in
GET /llms-full.txt.
The packaged verifier
@agledger/verify reproduces the same walk end to end, adding the hash, link and position checks
the stock script does not make, plus the checkpoint and admin-read chains, and walks the signed key
statements from a pin you give it. It has no AGLedger engine in its dependency tree. Pin the
version, and give it the Server's key pin as --trust-anchor: the "Vault signing key pin"
install.sh printed, or what node dist/scripts/signing-key-digest.js derives again from the
signing key. ./scripts/vault-dump.sh prints the same command when it finishes:
$ npx -y @agledger/verify@2.0.0 ./dump --trust-anchor sha256:0e8ee374fb6cc901dd96580dfe0c47f61701aae709591a95420fe0f478156446
[PASS] AGLedger offline verification (dump)
Nothing failed, and every signature was checked under a key the signed key statements
link to a --trust-anchor you gave.
key anchoring
status : walked from sha256:0e8ee374fb6cc901dd96580dfe0c47f61701aae709591a95420fe0f478156446 (write order)
anchored : 0e8ee374fb6cc901
unanchored : (none)
findings : 0
audit_vault chain
records : 7
entries : 26
checkpoints : 0
agent sigs : present=0 verified=0 (none on the chain)
failures : 0
org_admin_reads chain
orgs : 1
leaves : 2
checkpoints : 0
witness cosigned : 0
failures : 0
Without --trust-anchor the same run reads [VERIFIED, NOT ANCHORED]: nothing failed, but every
signing key was taken on the word of the dump itself, which a key written into the Server's
database alone would satisfy, so it is not a trusted verdict. It still exits 0; a gate that needs a
trusted verdict passes --trust-anchor, or reads verdict (trusted, unanchored or failed)
from --report-format json.
Exit 0 is verified, 1 is a verification failure, and 2 means it could not verify at all (a
missing or malformed input, a mistyped pin included), so only 1 is evidence of tampering. Given a
file instead, it verifies one record's audit export (GET /v1/records/{id}/audit-export). An export
carries its own signing keys, so for an independent check supply the keys you fetched separately:
npx -y @agledger/verify@2.0.0 export.json --keys keys.json --require-supplied-keys \
--trust-anchor <pin>
--keys takes the saved GET /v1/verification-keys response as it is. --require-key-id <id>
also rejects an export signed by any key but that one. These key-policy flags apply to an export
file only: a dump carries its own signed key history and refuses them. A supplied key still came
from the Server, so --trust-anchor applies to an export as it does to a dump.
@agledger/cli's verify subcommand reads audit exports only. Pointing it at a dump directory
reports EISDIR:
{"error":true,"code":"FILE_READ_ERROR","message":"Cannot read audit export at ./dump: EISDIR: illegal operation on a directory, read","suggestion":"Pass the path to an audit-export JSON file, or `-` to read it from stdin. Obtain one with `agledger api GET /v1/records/{id}/audit-export`."}
Both packages are optional. For a 2.x Server the verifier must be @agledger/verify 2.0.0 or
later, the first release that reads its dump: an older one fails such a dump once an admin has read
a record, with TENANT_READ_LEAF_HASH_MISMATCH, which reads as tampering and is not. That is why
the commands above pin the version rather than taking whatever npx resolves. The stock-library
path above remains the way to verify a chain with no AGLedger package in the loop at all.
When verification fails
A failure is the verifier doing its job. Tamper with one byte of a signed envelope and the stock-library script above rejects it at that position:
FAIL pos 1 019ead18-c3f9-7b4d-8edf-2dcd8b99fbc7
[FAIL] stock-library offline verification
audit_vault entries : 9
signatures verified : 8
failures : 1
signing keys : 1
A signature mismatch like that is one of a small set of integrity classes a full verifier reports.
The stock-library check above proves the signatures (catching CHAIN_SIGNATURE_INVALID and
CHAIN_SIGNATURE_MISSING_KEY); the in-database scripts/vault-verify.sh and the packaged
@agledger/verify add the hash, link, and position classes. The complete set on the per-record
chain:
| Code | Means |
|---|---|
CHAIN_EMPTY | A chain that should hold entries holds none, so nothing about it was proven |
CHAIN_MALFORMED_ENTRY | A row is missing a field the chain check needs |
CHAIN_GENESIS_INVALID | The first entry of a chain does not link to genesis |
CHAIN_POSITION_GAP | A chain position is missing - an entry was removed |
CHAIN_LINK_BROKEN | An entry's previous_hash does not match the prior entry |
CHAIN_HASH_MISMATCH | The stored hash does not match sha256(cose_sign1) |
CHAIN_SIGNATURE_INVALID | The signature does not verify against its resolved key |
CHAIN_SIGNATURE_MISSING_KEY | The signing key is not in the published registry |
CHAIN_COSE_DECODE_FAILED | The signed envelope is not decodable |
CHAIN_COSE_HEADER_MISMATCH | Chain mechanics in the protected header disagree with the row |
CHAIN_PAYLOAD_BINDING_MISMATCH | The denormalized row diverges from the signed payload |
CHAIN_OIDC_ACTOR_MISMATCH | The recorded actor identity disagrees with the signed claim |
CHAIN_SIGNING_KEY_DRIFT | The row's key column names a different key than the signature-covered kid |
CHAIN_ALG_MISMATCH | The signed header's algorithm is not one the trusted key can produce |
CHAIN_UNSUPPORTED_ALGORITHM | The key's algorithm is beyond this verifier build; upgrade, never a pass |
CHAIN_KEY_EXPIRED | The entry was signed after the key's retirement instant |
CHAIN_KEY_NOT_YET_ACTIVE | The entry was signed before the key's activation instant |
CHAIN_KEY_POLICY_VIOLATION | The entry violates a caller-set key policy (required key id, supplied keys only) |
CHAIN_ACTOR_ATTRIBUTION_MISMATCH | The row's actorId, actorRole or actorOwnerId disagrees with the actor claim signed in the protected header |
CHAIN_AGENT_SIGNATURE_INVALID | An entry's agent signature does not verify against the supplied cert key |
CHAIN_ENTRY_UNSIGNED | An unsigned entry written once the install had begun signing (at or after its first key activation, or after a signed entry on the same chain) |
CHAIN_SIGNING_KEY_UNANCHORED | With --trust-anchor, the entry is signed by a key the signed key statements do not link to a pin |
CHAIN_UNSUPPORTED_ALGORITHM is a capability gap, not a tamper signal. It means the verifier cannot
compute the algorithm the key names, so it fails closed rather than passing something it did not
check. The remedy is a newer verifier, or verifying on a host whose crypto provider carries the
algorithm - a FIPS-mode host reports it for every Ed25519 entry (see
FIPS 140 hosts).
Checkpoint and admin-read chains report their own classes on the same model:
CHECKPOINT_HASH_MISMATCH, CHECKPOINT_ROW_MISSING, CHECKPOINT_SIGNATURE_INVALID,
CHECKPOINT_CLAIM_MISMATCH, CHECKPOINT_UNSIGNED, CHECKPOINT_KEY_UNANCHORED,
TENANT_READ_LEAF_HASH_MISMATCH, TENANT_READ_LEAF_INDEX_GAP, TENANT_READ_SIGNATURE_INVALID,
TENANT_READ_CLAIM_MISMATCH, TENANT_READ_LEAF_UNSIGNED, TENANT_READ_KEY_UNANCHORED,
TENANT_CHECKPOINT_FORK, TENANT_CHECKPOINT_LEAF_COUNT_MISMATCH,
TENANT_CHECKPOINT_ROOT_MISMATCH, TENANT_CHECKPOINT_SIGNATURE_INVALID,
TENANT_CHECKPOINT_CLAIM_MISMATCH, TENANT_CHECKPOINT_UNSIGNED and
TENANT_CHECKPOINT_KEY_UNANCHORED. A _CLAIM_MISMATCH code means the row's columns disagree with
the claim its envelope signs; an _UNSIGNED code means an unsigned row written after the install
began signing; a _KEY_UNANCHORED code means a signing key no pin reaches.
Two of the classes above check attribution rather than integrity, and they are the reason an
export cannot be re-labelled. CHAIN_ACTOR_ATTRIBUTION_MISMATCH cross-checks the row's
actorId, actorRole and actorOwnerId - the attribution an export's own guide tells an auditor
to rely on - against the actor claim signed in the protected header, so an export re-attributed to
another actor fails rather than verifying. CHAIN_AGENT_SIGNATURE_INVALID re-verifies an agent's
own client-side signature offline. It needs the cert's public key. A full dump (not scoped with
--org) carries it: each EPHEMERAL_CERT_ISSUED entry on the platform-ops chain signs its cert's
public key, and @agledger/verify takes those keys from that chain once it has verified clean, and
counts them as vault.certKeysFromChain. For an org-scoped dump or an audit export, supply the keys with --agent-keys <file> on @agledger/verify or agentKeys on
the library call:
npx -y @agledger/verify@2.0.0 ./dump --trust-anchor <pin> --agent-keys agent-keys.json
Where no key is found the check does not run, and the result says so rather than implying it passed:
agentSignatures reports present and verified separately (the agent sigs line in the text
report), so
present > verified on a valid result means some were not checked, never that they failed. The
check applies only where on_behalf_of.validated is true; on a caller-asserted identity the
export cannot tell a passthrough value from a signed one. The Server's own chain verification also
compares each sealed cert against its live cert record (cert_missing, cert_actor_drift,
cert_window_drift, cert_expired), and that record is not exported, so those four have no offline
equivalent.
Note what does not fail verification: editing the convenience JSON in audit_vault.ndjson without
touching the signed envelope. The verifier trusts the signed cose_sign1 artifact as the source of
truth, not the denormalized columns - so a privileged-database edit of the readable payload is
caught as CHAIN_PAYLOAD_BINDING_MISMATCH, not silently accepted.
Findings on the key statements
Given --trust-anchor, the verifier first walks the signed key statements that travel with the
keys (a dump's vault_key_statements.ndjson, an audit export's
exportMetadata.signingKeyStatements) from the pin. A finding there fails the whole run, reported at
position 0, whichever key the entries are signed under:
| Code | When it appears | What the operator does |
|---|---|---|
KEY_STATEMENT_INVALID | A key statement does not verify. On an audit export the usual cause is "the endorser key is unknown": a statement signed by a key no key surface publishes, either a closure by a key the Server reaches but does not trust (such as a successor an attacker staged with a leaked key) or the admission of a trusted key whose endorser a forced retirement cut off. Also a statement that touches no anchored key, one stored after its signer's closure, or a second admission of the same key. | Run the vault scan (POST /v1/admin/vault/scan). keyRegistry.findings names the same statement, and its detail gives the remedy: the signer's pin in VAULT_TRUST_ANCHORS when the signer is honest, which publishes it, or, for a closure, the signer's VAULT_DISTRUSTED_KEYS entry at the closure's write time when it leaked. Set it on every process and restart. |
KEY_CLOSURE_INVALID | A key is listed as retired and no closure the verifier can verify signs that retirement, or a closure counts though its signer stored it after its own retirement or it dates a retirement before its subject's activation. It usually comes with KEY_STATEMENT_INVALID, when the only closure that retires the key is one the verifier cannot verify. | Resolve the scan finding on that closure as above. A key a forced retirement retired and that you have since pinned takes an unforced retire call to sign its retirement (see the key-compromise runbook). |
CHAIN_KEY_WINDOW_DRIFT | A listed activatedAt or retiredAt differs from the window the verifier's own walk signs for the key. There are three causes. On an audit export, the key's listing carries distrustedFrom and its retiredAt is that instant: the Server's VAULT_DISTRUSTED_KEYS cut the window there, earlier than the retirement the key's closure signs, and the verifier was not given that entry, so it walks to the signed retirement. Otherwise, on an audit export, a statement that dates the window is one the verifier cannot verify, so it reads the window differently from the Server. On a dump it can also be a registry column written outside the engine: a dump carries vault_signing_keys rows as stored, and a retired row's retired_at that differs from the signed retirement is reported against that row. | distrustedFrom is the export's own word, and nothing signs it, so it proves nothing by itself. Ask the Server's operator whether VAULT_DISTRUSTED_KEYS names that key at that instant. Only when they confirm it, give the verifier the entry, --distrusted-key sha256:<the key's SPKI digest>@<distrustedFrom> beside the same pins, and the windows agree; off a dump the entry also voids every admission the key signed from that instant. With KEY_STATEMENT_INVALID, resolve that finding; the windows then agree. On a dump alone, the scan reports the same row as key_window_drift and says the column was written out of band. No setting clears it, because a retired row cannot be rewritten: it is evidence of database write access outside the engine, to investigate as such. Entries are held to the signed value, so exports still certify and the verdict on every entry is unchanged. |
The Server grades the key material it publishes the same way before it vouches for an export. While
a walk over that material reports a finding (the scan lists it in keyRegistry.findings, and its
detail says every audit export fails offline verification while it stands), every record's
GET /v1/records/{id}/audit-export answers chainIntegrity: false with chainIntegrityReason: signing_key_material_invalid, and ?integrity=true on the record answers verified: false, though
every entry verified. Not every registry finding is of that kind. A retired row whose retirement no
counting closure signs, a key_window_drift on a column, and a statement no key surface publishes
(one a distrusted key stored after a key you trust retired it, for instance) keep the scan at
healthy: false while every export certifies, because nothing an export carries is wrong. A dump
carries every registry row and statement, signers an export lacks included, so the same registry can
pass as a dump while its exports fail, or the other way round. After the remedy, take a new export or
dump: one taken before it still carries the old statements.
What a distrusted key signed
VAULT_DISTRUSTED_KEYS names a key the operator no longer trusts (a leaked key, or one an attacker
staged with it), from an instant. What such a key signed from that instant until a key you still
trust retired it is accounted for: the operator has said why it cannot be trusted, and the retirement
bounds it. The Server never verifies it, and the scan lists it rather than counting it broken:
- a chain entry it signed, under
distrustedEntries(chain, record or org, position, key), with reasonsigning_key_distrusted; a chain whose only findings are such entries is counted indistrustedSigned(andglobalChains.distrustedSignedfor the record-less chains), in neitherverifiednorbroken; - a key statement it signed, which counts for nothing, under
keyRegistry.distrustedStatements.
healthy does not fold them in: healthy: true beside a non-empty distrustedEntries means nothing
the scan found is unexplained. The checkpoint sweep still counts each chain carrying such an entry on
agledger_vault_key_window_violations_total{reason="signing_key_distrusted"}, so
AGLedgerVaultKeyWindowViolation shows writes under a leaked key. A retirement waits for every
entry and statement under the key still in flight before it stamps its instant, so nothing written
after it can carry a time before it. A record whose chain carries such an entry still answers
chainIntegrity: false with chainIntegrityReason: signing_key_distrusted on its audit export,
since that entry is never presented as verified. What the key signs after that retirement, and what a
distrusted key signs before any key you trust has retired it, is not accounted for: it keeps its own
class (signing_key_unanchored, key_expired, a key-statement finding) and keeps the scan at
healthy: false. While a distrusted key has a registry row and nothing you trust has retired it, the
scan also reports it as key_closure_invalid and names the forced retire call that bounds it.
An offline verifier reads a dump the same way when it is given the same entries as
--distrusted-key and the same pins as --trust-anchor: it lists those entries and statements as
accounted for instead of failing on them. An audit export carries no write times it can hold a key
statement to, so even with the entries an export fails on what the Server accounts for: an entry
signed after the cutoff as CHAIN_KEY_EXPIRED, a statement stored after it as a key-statement
finding. Without the distrust entries a dump and an export also read differently. A dump carries the registry rows as stored and no distrust entry, so the verifier holds the key to the
retirement its closure signs: what it signed before that retirement passes, entries the instant took
away included, and only what it signed after the retirement fails, as CHAIN_KEY_EXPIRED. An audit
export lists the key with the window the Server holds it to, cut at the distrust instant and marked
with distrustedFrom, which a verifier without the entry reports as CHAIN_KEY_WINDOW_DRIFT on that
key and fails the export. Hand auditors the entries with the pins.
On-box reads versus off-box verification
An operator can read the chain on the Server through the admin vault endpoints (see the
API reference for /v1/admin/vault/*). These platform-scoped reads are not notarized.
None of them calls the record-read notarization path that feeds org_admin_reads, and none writes a
system_audit_log row either, except POST /v1/admin/vault/anchors/verify and
POST /v1/admin/vault/anchors/reconcile, which write one (chain.rewind_detected) only when they
detect a fork or a rewind rather than on every call. What org_admin_reads notarizes is cross-party reads of record data, such as
GET /v1/records/{id} and its neighbours, covered in the org-reads transparency log below. Reading
the chain through the admin vault endpoints leaves no cryptographic trail of its own beyond your
access logs; off-box verification is the accountable path.
An auditor does the opposite: they take the dump off the Server and verify it on their own machine with only the public keys. On-box reads are for operations. Off-box verification is the audit. Keep the two roles distinct.
For a single record, GET /v1/records/{id}/audit-export?evidence=true inlines each completion's
evidence body at its COMPLETION_SUBMITTED entry and documents the evidenceHash binding
(SHA-256 over the RFC 8785 (JCS) canonicalization of the evidence JSON) in the export's
verificationGuide.evidenceBinding, so an offline auditor can re-bind a separately-held evidence
body to the signed chain without a live fetch. The response shape and re-binding recipe are in the
API reference.
The org-reads transparency log
Every cross-party read by an org-admin key is notarized into the org_admin_reads chain, and that
chain gets its own periodic signed tree head. This is the symmetric half of the accountability
story: the record chain says what the agents did, and this says who went looking at it.
Five endpoints, all org-scoped. The four reads take an admin or an agent key holding audit:read;
the cosign takes an org-admin key with admin:system:
| Endpoint | What it gives you |
|---|---|
GET /v1/audit/org-reads | The leaves themselves, oldest first in leafIndex order |
GET /v1/audit/org-reads/checkpoints | The signed tree heads, newest first |
GET /v1/audit/org-reads/checkpoints/{id} | One checkpoint envelope by id |
GET /v1/audit/org-reads/checkpoints/{id}/proof?leaf=N | The inclusion path for one leaf |
POST /v1/audit/org-reads/checkpoints/{id}/cosign | Record an external witness signature |
The leaf listing is what makes a checkpoint walkable by someone outside your org. leafIndex is
the order the listing serves and the index the proof endpoint takes, so a third party handed a
signed checkpoint reads the leaves, picks the entry they care about, and asks for its inclusion
path against that checkpoint. Each row carries the leafHash (the RFC 9162 leaf hash of the row's
COSE_Sign1 bytes, defined below), the record read, the key that read it, the filter applied, and the read context.
An empty list on a fresh install is expected
Checkpoints are built by a sweep on the Worker every six hours (0 */6 * * *, fixed, not
configurable). A fresh install returns
{"data":[],"hasMore":false,"nextCursor":null,"total":0}
until the first sweep runs, no matter how many reads the org has already taken. The same applies to
the signedCheckpointRef in a read response: it is null immediately after the read, and the next
sweep populates it. Do not read either as a missing or broken feature.
The checkpoints listing carries a checkpointing block (cron, intervalMinutes, nextRunAt,
lastCheckpointAt, source) so you can tell a sweep that has not run yet from one that is broken.
Read it before reporting missing checkpoints. To tell an empty checkpoint list apart from an org
that has logged no qualifying reads at all, read GET /v1/audit/org-reads: empty there means
nothing has been logged, while leaves with no checkpoint mean the sweep has not reached them.
Which tree this is
This log is an RFC 9162 (Certificate Transparency 2.0, section 2.1) Merkle tree over SHA-256, the
same construction SCITT Receipts carry (/v1/scitt/*, VDS RFC9162_SHA256). It is a separate log
from the SCITT one, so a proof from one does not verify against the other's root. Any RFC 9162
inclusion verifier works unchanged.
Every hash the API returns here is lowercase hex of 32 bytes. Hash the bytes the hex denotes, never the hex text. The leaf data is the row's COSE_Sign1 envelope bytes, so a leaf hash and a node hash are:
leafHash = sha256(0x00 || cose_sign1_bytes)
NODE(L, R) = sha256(0x01 || L || R)
The root over leaves D[0:n] is the leaf hash when n is 1, and otherwise
NODE(root(D[0:k]), root(D[k:n])), where k is the largest power of two less than n. An empty
tree (treeSize 0) has root sha256(""), which is e3b0c442...b855. The same leafHash is the
previous_hash the next leaf's signed claim links to.
Verifying an inclusion proof
path is the RFC 9162 audit path, leaf level first. Its length depends on the leaf's position: a
right-edge leaf of an unbalanced tree skips levels, and a one-leaf tree has an empty path. The walk
is RFC 9162 section 2.1.3.2. leafHash, rootHash and each path entry arrive as hex, so the
listing decodes them before hashing:
import hashlib
def NODE(left, right):
return hashlib.sha256(b"\x01" + left + right).digest()
def verify(proof):
# proof: the JSON body of GET /v1/audit/org-reads/checkpoints/{id}/proof?leaf=N
fn, sn = proof["leafIndex"], proof["treeSize"] - 1
if fn > sn:
return False
r = bytes.fromhex(proof["leafHash"])
for p in map(bytes.fromhex, proof["path"]):
if sn == 0:
return False
if fn & 1 or fn == sn:
r = NODE(p, r)
while not fn & 1 and fn != 0:
fn, sn = fn >> 1, sn >> 1
else:
r = NODE(r, p)
fn, sn = fn >> 1, sn >> 1
return sn == 0 and r == bytes.fromhex(proof["rootHash"])
A proof that passes shows the leaf is under that rootHash. Two comparisons tie it to the rest: the
proof's leafHash should equal the leafHash of the row at that leafIndex in
GET /v1/audit/org-reads, and its rootHash should equal the rootHash of the checkpoint you
fetched by id, whose signature is checked next.
The checkpoint envelope itself is a COSE_Sign1 (RFC 9052) over the tree head, signed by the same
vault key registry the record chain uses, so coseSign1Base64 verifies with the public keys from
GET /v1/verification-keys exactly as Surface 2 above does. An external monitor that wants to pin a
head can post its own signature back with the cosign endpoint; that witness signature is
single-use per checkpoint.
Optional - the SCITT transparency checkpoint
If you operate the SCITT transparency surface (POST /v1/scitt/entries), GET /v1/scitt/checkpoint
returns the current signed tree head for an org's transparency log - a stable anchor an external
monitor can pin and re-check for consistency over time. The checkpoint is org-scoped: use an
org-scoped key (a platform-scope key is refused with 403 ORG_REQUIRED).
curl -s -H "Authorization: Bearer $AGLEDGER_API_KEY" "$AGLEDGER_API_URL/v1/scitt/checkpoint"
{"treeSize":0,"rootHex":"e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855","logId":"019ead17-fbd4-7381-ac3f-5ee1474830f1","iat":1781020588,"kid":"c4dd3e20388b594d","signature":"a130648224123432149bf933086a9a25cab63ee7f8a8e6789e52056a5ae7f4f5f1b1447dc3d3ae8e81bdb0d2ddd31f90a93a3d35c83a851d51011ea82502c90f"}
A treeSize of 0 with the empty-tree root above is a fresh log with no SCITT entries registered
yet. The signature input is ${logId}:${treeSize}:${rootHex}:${iat} over UTF-8 bytes; the logId
prefix is what stops a checkpoint being substituted across orgs, so the three-part form without it
does not verify. Verify against the public key at GET /.well-known/scitt-keys/{kid}.
This surface is independent of the audit_vault chain in Surface 2: the offline verifier is the
proof of the record chain; the SCITT checkpoint is the anchor for the separate SCRAPI transparency
log.