Install AGLedger on Kubernetes
This guide brings up a single AGLedger Server on a generic Kubernetes cluster using the published Helm chart, fronted by TLS, backed by your own PostgreSQL. On OpenShift the same install takes one extra flag; see the OpenShift section below.
Just evaluating? Start here instead. This page is the production Kubernetes install. To get a signed Server running locally in about five minutes with a single command (bundled PostgreSQL, keys generated for you, no cluster), use the quick install on Docker Compose.
Running 1.x? 2.0 installs against a new, empty database. A fresh 2.0 install applies one migration,
001_consolidated.sql, the 2.0 baseline. 2.0 does not upgrade a database a 1.x release migrated: pointed at one, the migration runner applies nothing, says the database was migrated by AGLedger 1.x, and tells you to pointDATABASE_URL(andDATABASE_URL_MIGRATEandDATABASE_URL_DIRECTwhere set) at a new database. Keep the 1.x database for the 1.x install that wrote it. The backup and restore scripts do not carry 1.x data into 2.0 either. Version upgrades in Day-2 operations covers upgrades within 2.x and the migration knobs they use.Release images are multi-arch (
linux/amd64andlinux/arm64) on a Red Hat UBI 10 runtime base (ubi10/nodejs-24-minimal). Pulling by tag selects your platform automatically. If you pin by digest, pin the index digest used below rather than a per-architecture child digest, or you will pin one architecture. On arm64 (Graviton) nodes the runtime is native, and signatures produced there verify against the offline verifier on x86.
The API reference is the OpenAPI document the Server serves at /openapi.json. This guide
does not restate request or response schemas; it links to them.
What you provide
A Server is durable only as far as its database and its signing key. You bring both:
- A PostgreSQL database. AGLedger writes a hash-chained, Ed25519-signed record chain to it.
Use any PostgreSQL 17 or later that serves TLS: in production the Server refuses a
DATABASE_URLwithout a verifiedsslmode(step 3). Connect directly or through a pooler in session mode. Behind a transaction-mode pooler (PgBouncerpool_mode=transaction, RDS Proxy), also setsecrets.databaseUrlDirectto a connection string that bypasses it, or the Server refuses to boot; High availability covers that topology. Three role requirements are worth knowing before you provision it: the role that runs migrations must be a superuser (the schema installs an event trigger, and event triggers require one), and the role the Server runs as must be namedagledger_appunless it also owns the schema objects, because migrations grant table privileges to that name. Step 3 uses the shape the schema is built for: migrate as the owner, serve asagledger_app. A Server that connects as the owner of its tables passes every privilege check on them, so the revokes that keep the audit chain append-only bind nothing, and preflight warns that it can rewrite or delete chain rows. A non-owner runtime role named anything else gets no privileges and the Server fails at boot withpermission denied for table .... If your naming standard forbidsagledger_app, make your role a member of it after migrating:GRANT agledger_app TO "your_role" WITH INHERIT TRUE. Spell out theINHERITclause, because PostgreSQL 16 and later default it to the member role's own setting, so aNOINHERITrole receives nothing from a bare grant. Membership carries the privilege set the schema intends, including the append-only revokes that leaveaudit_vaultand its siblings insert-only (apart from the column-scoped updates a key retirement and a witness cosignature take) and theALTER DEFAULT PRIVILEGESthat keep later migrations reachable. Do not substitute a blanketGRANT ... ON ALL TABLES IN SCHEMA public: it also hands the runtime roleUPDATEandDELETEonaudit_vault, which is the write the chain is protected against. Ifagledger_appdoes not exist at all, the migrating role could not create it and the schema granted nothing to anyone; grant your roleUSAGEon schemapublic,SELECT, INSERT, UPDATE, DELETEon all its tables and the matchingALTER DEFAULT PRIVILEGES, then re-apply the append-only revokes and the column grants that follow them, which the Privileges block at the end of the baseline migration (001_consolidated.sql) lists. Separately from anything the migration grants, the runtime role needsCREATEon the database (GRANT CREATE ON DATABASE agledger TO agledger_app): pg-boss keeps its queue tables in a schema it installs on first start, and the Server starts pg-boss, so without it the Server exits at boot withpermission denied for database .... Creating the schema by hand does not substitute, because pg-boss reads the absence of its own version table as "not installed" and issuesCREATE SCHEMA IF NOT EXISTSregardless, which PostgreSQL refuses without the database-level privilege. When the migration setsagledger_app's password (step 3), it grants this for you. You do not have to get any of this right from memory: on the external-database path the chart runs both checks as apre-install/pre-upgradehook, after migrating and before any workload starts. A role that cannot serve failshelm installbefore the API or Worker is created, and the hook Job's log names the role it connected as and the exactGRANTto run. The same hook also runs the refusals the Server makes when it loads its configuration (NODE_ENV,AGLEDGER_EXTERNAL_URL, the databasesslmode=rule,TRUST_PROXYand the numeric knobs), so a value the API and worker would crash-loop on fails the install there instead, with the Server's own message. Read it withkubectl logs -n <namespace> --tail=-1 -l app.kubernetes.io/instance=<release>,app.kubernetes.io/component=preflight(the instance selector matters if you run more than one release in a namespace). Apply the grant, then re-run. Migrations have already been applied at that point and are skipped on the retry, and no workload was created, so the retry costs nothing../scripts/helm-install.shrunshelm upgrade --install, so a re-run of it reconciles whether or not the first attempt reached a deployed revision. A hand-runhelm installdoes not: it refuses a name that already exists, so either switch tohelm upgrade --installor clear the failed release withhelm uninstall <release>before retrying. If helm reports that another operation is in progress, an earlier run was killed mid-flight:helm rollback <release> -n <namespace>if it has a revision to go back to,helm uninstallif it does not, then re-run. - An Ed25519 vault signing key. The Server signs every record with it. Generate it once, store it like any other root secret, and reuse it across upgrades: the published verification key is derived from it.
Prerequisites
- A Kubernetes cluster (1.27+) and
kubectlpointed at it - PostgreSQL 17 or later, 18 recommended for its native
uuidv7(). The Server image carries its own Node 24 LTS runtime, so nothing is installed on the cluster for it. helm3.x (3.14 or later for./scripts/helm-install.sh) andcosign3.0+- An ingress controller and a way to issue TLS certificates. This guide uses ingress-nginx and cert-manager; any controller and certificate source work.
- A usable
StorageClass, only if you let the chart run PostgreSQL in-cluster (see step 3) or keep backups on a volume. Without a cluster default, name a class:postgres.bundled.storageClassNamefor the bundled PostgreSQL,backup.persistence.storageClassNamefor the backup volume. With neither, the claim staysPendingand the install hangs. Modern EKS (1.30+) ships no default class: name your provisioner's class, for examplegp3, with the EBS CSI driver installed.
1. Verify the release
Releases are keyless-signed: GitHub Actions OIDC -> Sigstore/Fulcio -> the public Rekor transparency log. There is no static public key to fetch. A valid signature binds to the GitHub Actions workflow in agledger-ai/agledger-api that built the artifact, verifiable with no access to the source repository. Requires cosign 3.0+. Verify both the image and the chart before you install.
$ cosign verify \
--certificate-identity-regexp '^https://github\.com/agledger-ai/agledger-api/\.github/workflows/.+@refs/tags/v.+$' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com \
agledger/agledger:2.0.0
$ cosign verify \
--certificate-identity-regexp '^https://github\.com/agledger-ai/agledger-api/\.github/workflows/.+@refs/tags/v.+$' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com \
registry-1.docker.io/agledger/agledger-chart:2.0.0
Expected: each prints Verification for ... -- and a JSON block. The certificate identity in
the output is the GitHub Actions workflow that built the artifact, and the image digest in the
output (sha256:6e297796…) is the index digest you pin in step 3. To additionally check build
provenance, verify the SLSA Build L3 attestation with slsa-verifier verify-image (full recipe
in the install repo's SECURITY.md at https://github.com/agledger-ai/install).
If you install with the scripts rather than by hand, set AGLEDGER_REQUIRE_VERIFY=true. The
scripts verify both the chart and the image, resolving the tag to a digest first so there is no
gap between what was verified and what is deployed. With cosign present and a verification that
fails, they refuse whatever you set. With cosign absent they warn and proceed, which suits an
evaluation; this variable makes that case refuse instead, so set it on production hosts and in CI.
$ AGLEDGER_REQUIRE_VERIFY=true ./scripts/helm-install.sh --version 2.0.0
The lever in the other direction, --skip-verify, exists for local development. Neither it nor
AGLEDGER_SKIP_VERIFY=true belongs on a host that serves anything.
2. Generate the vault signing key
Generate the Ed25519 key with the Server image itself, then keep the private key safe. Create the
namespace first and run the key generator in it: a pod that runs in default says nothing about
the Pod Security admission, quotas and image-pull secrets of the namespace the Server will run in.
$ kubectl create namespace agledger
$ kubectl run agledger-keygen -n agledger --restart=Never --image=agledger/agledger:2.0.0 \
--command -- /nodejs/bin/node dist/scripts/generate-signing-key.js
$ kubectl logs agledger-keygen -n agledger
VAULT_SIGNING_KEY=<base64 Ed25519 private key>
Public key: MCowBQYDK2VwAyEAPInxjno26azIT9i6GVqJag9QuCJFMoDG96iljTd8fHo=
Fingerprint: dcd0573755f045c6
Algorithm: Ed25519
Pin: sha256:dcd0573755f045c6<48 more hex>
$ kubectl delete pod agledger-keygen -n agledger
The fingerprint (dcd0573755f045c6) is the key id the Server later publishes at
/v1/verification-keys and reports as signingKey.keyId on /health. Confirming they match
(step 5) is how you prove the Server is signing with the key you provided. The pin is not secret:
give it to anyone who verifies this install offline. Both are SHA-256 of the same public key, so
the fingerprint is always the first 16 hex characters of the pin. The pin stays valid across later
key rotations, and dist/scripts/signing-key-digest.js derives it again from VAULT_SIGNING_KEY.
On a FIPS-mode cluster, add --algorithm es256; read FIPS 140 hosts first.
3. Connect PostgreSQL and install
Make a PostgreSQL reachable from the namespace that serves TLS. Then install the chart: the
migration connects as the owner, the Server connects as agledger_app, and the signing key goes in
from a file.
Simpler path for evaluation: let the chart run PostgreSQL. Set
postgres.bundled.enabled: trueinvalues.yamland omitdatabase.externalUrlandsecrets.databaseUrlMigrate: the chart provisions an in-clusterpostgres:18-alpinewith a 10Gi PersistentVolumeClaim, on a private connection it exempts from the TLS requirement, and sets up both roles itself (the migration runs as the owner, the API and worker asagledger_appunder a password derived from the owner's). Bundled PostgreSQL is not recommended for production. It requires a usableStorageClass: with no cluster default and no explicitpostgres.bundled.storageClassName, the claim staysPendingand the install hangs. On EKS 1.30+ (no default class) set the class explicitly.postgres: bundled: enabled: true password: <change-me> # the owner's password; agledger_app's is derived from it storageClassName: gp3 # required on EKS 1.30+; "" uses the cluster defaultConfirm the claim binds before waiting on the workloads:
kubectl get pvc -n agledgershould show the-pgdataclaimBound, notPending. On a first bundled install the API and worker start beside the migration, before it has givenagledger_appits password, so they exit onpassword authentication failed for user "agledger_app"and restart, usually two or three times, until the migrate Job reportsComplete 1/1. That is expected and clears on its own. Choose one database path, not both. When both are set,database.externalUrlwins and the chart renders no bundled PostgreSQL. The steps below use an external PostgreSQL.
Generate the password agledger_app will log in with. The migration sets it on every run, so it
goes in two places in values.yaml:
$ openssl rand -hex 24 # <app-password>
values.yaml:
image:
digest: "sha256:6e29779624951807dbcf6adeb5c6c7b08da153339fab7ec6e1210e4ebfc66a05" # 2.0.0
database:
# The Server's connection, as agledger_app.
# In production the Server refuses to boot unless sslmode is require, verify-ca or verify-full.
# Use verify-full, and set config.nodeExtraCaCerts when the database's CA is not one the image
# already trusts.
externalUrl: "postgresql://agledger_app:<app-password>@<db-host>:5432/agledger?sslmode=verify-full"
secrets:
# The migration's connection, as the superuser that owns the schema.
databaseUrlMigrate: "postgresql://<owner-role>:<owner-password>@<db-host>:5432/agledger?sslmode=verify-full"
migrate:
extraEnv:
# Every migration run gives agledger_app its login under this password, CREATE on the
# database and ownership of the pg-boss schema. Same value as in database.externalUrl.
- name: AGLEDGER_APP_ROLE_PASSWORD
value: "<app-password>"
config:
# Required. The Server's signed issuer identity, and permanent: rows already written keep the
# issuer they were signed with, so changing it later splits the chain rather than correcting it.
# A production render with neither this nor an ingress host is refused rather than defaulted.
externalUrl: "https://agledger.k8s.example"
# The addresses the ingress controller connects from (your cluster's pod CIDR; 10.244.0.0/16 is
# the kind and flannel default), so X-Forwarded-For is read only when it wrote it. Left false,
# the lockout, rate limits, API key allowedIps and audit rows all see the controller's address
# instead of the client's. Never `true`.
trustProxy: "10.244.0.0/16"
ingress:
enabled: true
className: nginx
annotations:
cert-manager.io/cluster-issuer: agledger-selfsigned
hosts:
- host: agledger.k8s.example
paths: [{ path: /, pathType: Prefix }]
tls:
- secretName: agledger-tls
hosts: [agledger.k8s.example]
A values file written for 1.x can fail to render. Two refusals catch it, each with text that
names the fix. A networkPolicy.ingressFrom entry needs a namespaceSelector beside its
podSelector; 1.x filled in namespaceSelector: {} itself, and 2.0 refuses the entry instead:
networkPolicy.ingressFrom: every entry needs a namespaceSelector beside its podSelector. A pod label
matched in every namespace admits any pod that carries it, and any workload can give its pod that
label. Name the namespace the pod runs in, for example namespaceSelector: {matchLabels:
{kubernetes.io/metadata.name: ingress-nginx}}, or write namespaceSelector: {} to admit every
namespace deliberately.
Name the namespace your ingress controller runs in. The chart's default ingressFrom already does
this for ingress-nginx in ingress-nginx and Traefik in traefik or kube-system, so a values
file that does not set ingressFrom is unaffected. The second refusal is ingress.enabled with no
ingress.tls: the message starts ingress.enabled is true with no ingress.tls and lists the three
ways out, which are a certificate under ingress.tls as above, an AWS ALB certificate-arn
annotation, or ingress.allowPlainHttp=true when something in front of the Ingress terminates TLS.
Write the key to a file rather than passing it inline. --set puts the value in helm's
argv, where ps shows it to every other user on the machine for the length of the install,
and the vault signing key is the private key every record is signed with.
$ umask 077 && printf %s '<vault-key>' > vault-key
$ helm install agledger oci://registry-1.docker.io/agledger/agledger-chart \
--version 2.0.0 --namespace agledger \
--values values.yaml --set-file secrets.vaultSigningKey=vault-key
NAME: agledger
STATUS: deployed
REVISION: 1
$ rm vault-key
Use helm upgrade --install rather than helm install if you expect to re-run this command. It
reconciles an existing release instead of refusing its name, which is what makes a corrected value
a one-command fix rather than an uninstall. On a bundled-PostgreSQL release an uninstall takes the
database with it, so the difference is not cosmetic.
The chart runs schema migrations in their own Job, then starts the API and worker. A fresh install
applies the baseline migration in that single Job, which reports Complete 1/1. Re-running it
against a database already at that head applies nothing and reports a migration count of 0, so a
repeated helm upgrade is safe.
On the external-database path that Job is a pre-install,pre-upgrade hook, so a migration pod that
cannot schedule fails the whole install or upgrade, and a second hook Job, preflight, runs after
it as the runtime-role gate described in What you provide. On the bundled-PostgreSQL path the
migrate Job is an ordinary chart resource and no preflight Job renders, because helm creates the
bundled PostgreSQL only after every pre-install hook has succeeded. Because it is not a hook there,
helm records the revision as deployed once it has applied the manifests, whether or not the
migration then succeeds, so deployed on that path does not mean the migration ran.
./scripts/helm-install.sh waits for the Job and exits non-zero when it fails, printing the Job's
log and, when the revision replaced the pods that were serving, the helm rollback that restores
service. Fix what the log names, then re-run.
If your API pods are pinned to a node class (architecture, taints), set migrate.nodeSelector /
migrate.tolerations / migrate.affinity to match. Each is empty by default and falls back to the
matching api.* value, so an install that pins only the API is already consistent.
Sizing an upgrade window
Three environment variables size a migration window: MIGRATION_STATEMENT_TIMEOUT,
MIGRATION_MAINTENANCE_WORK_MEM, and MIGRATION_LOCK_TIMEOUT. What each one does and how to
estimate your own cost are in Version upgrades in
Day-2 operations. Read it before upgrading a mature install.
migrate.extraEnv delivers them to the migration Job (it is a separate list from the top-level
extraEnv, which goes to the API and worker):
migrate:
extraEnv:
- name: MIGRATION_STATEMENT_TIMEOUT
value: "60min"
- name: MIGRATION_MAINTENANCE_WORK_MEM
value: "256MB"
Do not set DATABASE_URL here: the Job already receives it from the chart's Secret, and the chart
refuses to render a second channel for it. Point the migration at its own role with
secrets.databaseUrlMigrate instead.
On Compose the same knobs are read from .env; see the migration section of .env.example.
$ kubectl rollout status deploy/agledger-agledger-chart-api -n agledger --timeout=180s
deployment "agledger-agledger-chart-api" successfully rolled out
$ kubectl rollout status deploy/agledger-agledger-chart-worker -n agledger --timeout=120s
deployment "agledger-agledger-chart-worker" successfully rolled out
4. Create the platform API key
The platform key is the first credential; you use it to provision organizations and agents. It is printed once.
$ kubectl exec deploy/agledger-agledger-chart-api -n agledger -- \
env NODE_OPTIONS= /nodejs/bin/node dist/scripts/init.js --non-interactive
✓ API_KEY_SECRET found in environment
✓ VAULT_SIGNING_KEY found in environment (Ed25519)
✓ AGLEDGER_FEDERATION_SIGNING_KEY generated (Ed25519)
⚠ Read-only filesystem: .env could not be written.
3 secret value(s) already set in this environment are shown by name only.
The secrets generated by this run are printed in full below and exist nowhere else. Capture them.
--- .env content ---
...
DATABASE_URL=<already set in this environment; not reprinted>
API_KEY_SECRET=<already set in this environment; not reprinted>
VAULT_SIGNING_KEY=<already set in this environment; not reprinted>
AGLEDGER_FEDERATION_SIGNING_KEY=<printed in full>
---
✓ Platform API key created (ID: ...)
✓ Database query OK
│ Platform Key: agl_plt_<store-this-securely>
The chart runs the container with a read-only root filesystem, so init.js cannot write a .env
and prints one to stdout instead. Everything you supplied in steps 2 and 3 is shown by name only:
your database password, API_KEY_SECRET and VAULT_SIGNING_KEY are already in your Secret and are
not reprinted.
Two things in that output are secret and exist nowhere else: the platform key, and the
federation keypair this run generated. The federation key is optional. Capture it into your Secret
as AGLEDGER_FEDERATION_SIGNING_KEY if you intend to federate; otherwise ignore it, and a later
run generates a fresh pair. Either way, do not run this step in a CI job whose logs are retained.
5. The readiness gate: up and signing
A Server is ready when it is healthy and signing with your key.
$ kubectl port-forward svc/agledger-agledger-chart -n agledger 3001:80 &
$ curl -s http://localhost:3001/health
{"status":"ok","version":"2.0.0","timestamp":"...","signingKey":{"gate":"usable","keyId":"dcd0573755f045c6"}}
$ curl -s http://localhost:3001/health/ready
{"status":"ready","version":"2.0.0","timestamp":"..."}
signingKey.gate is usable and keyId is the keygen fingerprint from step 2: this process holds
your key and may sign with it. The published key list says the same from the verifier's side:
$ curl -s http://localhost:3001/v1/verification-keys
{
"data": [{ "keyId": "dcd0573755f045c6", "algorithm": "Ed25519",
"publicKey": "MCowBQYDK2VwAyEAPInxjno26azIT9i6GVqJag9QuCJFMoDG96iljTd8fHo=",
"status": "active", "coseAlgorithm": -8, "minVerifierVersion": "2.0.0", ... }],
"envelope": "COSE_Sign1", "signatureAlgorithm": "Ed25519", ...
}
The Server is signing with your key, COSE_Sign1 / Ed25519. minVerifierVersion is the oldest
offline verifier that checks what this Server signs: give anyone verifying this install
@agledger/verify 2.0.0 or later, with the pin from step 2.
Every feature is enabled whether or not a license key is installed; the license is contractual. With no key installed the Server reports itself unlicensed, which licenses evaluation, development and testing only:
$ curl -s -H "Authorization: Bearer <platform-key>" http://localhost:3001/v1/admin/license
{"validity":"unlicensed","tier":"unlicensed",
"features":["custom_schemas","expression_engine","delegation_chains","audit_export",
"compliance_reports","encrypted_mode","entity_references","proposals","federation"],
"notice":{"kind":"unlicensed", ...},"source":"none", ...}
An external database is licensed by Enterprise Edition, one license per database instance. Put
the key in secrets.license (the compact agl_<tier>_v1_... string) and upgrade the release,
which rolls the pods onto it. To rotate a key later without a restart, mount it from a Secret of
your own with license.keyFile and call POST /v1/admin/license/reload.
6. TLS
With ingress-nginx and cert-manager, the chart's Ingress requests a certificate and serves it. A self-signed ClusterIssuer is portable and needs no public DNS; for production use an ACME (Let's Encrypt) or your-CA issuer.
$ kubectl get certificate -n agledger
NAME READY SECRET AGE
agledger-tls True agledger-tls ...
$ curl -sk https://agledger.k8s.example/health # through the ingress
{"status":"ok","version":"2.0.0","timestamp":"...","signingKey":{"gate":"usable","keyId":"dcd0573755f045c6"}}
The scripted path
./scripts/helm-install.sh does steps 1 to 3 in one command, and it is the path to use if you
expect to correct a value and run it again. It resolves and verifies the chart and image, generates
a vault signing key only when the release does not already have one, installs with
helm upgrade --install --reset-then-reuse-values (helm 3.14 or later, which it checks for), and
waits on the rollout. It does not run step 4: it prints the kubectl exec ... init.js command for
it, which you run yourself.
$ ./scripts/helm-install.sh --version 2.0.0 --namespace agledger --release agledger \
--external-url https://agledger.k8s.example \
--db 'postgresql://agledger_app:<pw>@db.example:5432/agledger?sslmode=verify-full'
--external-url is the Server's issuer, the same permanent value as config.externalUrl in step 3.
Left out, the script asks for it at the terminal, with https://localhost as the answer Enter
gives. With no terminal to ask on (CI, a detached run) and nothing in your --set flags, values
file or existing release naming one, it refuses and installs nothing, rather than make
https://localhost the issuer of a Server that has a domain. Pass
--external-url https://localhost to choose that on purpose.
On a first install the script generates the vault signing key and prints it with its pin at the
end of the run, along with the fingerprint /health and /v1/verification-keys report. If the
cluster cannot pull the image, it stops at key generation, before anything is installed, and names
the image and the kubelet's reason. If the rollout wait fails after it generated a key, it names
the release Secret holding that key and the command to read it; a re-run reuses the key and prints
it.
--db takes the connection string as a separate argument (there is no --db= form) and hands it
to helm through --set-file from a mode-600 temp file, so the password never reaches helm's argv
where ps would show it.
--db applies a database CA on its own. A managed PostgreSQL presents a certificate the
container has no root for, and sslmode=verify-full then fails the handshake during the migration
hook. So when you name a database and have not named a CA, the script sets
config.nodeExtraCaCerts to /etc/ssl/certs/rds-global-bundle.pem, the AWS RDS and Aurora root
bundle baked into the image. Precedence, highest first: your own
--set config.nodeExtraCaCerts, your own values file, a CA the release already carries, then that
default. --ca-cert <path> names a different bundle and --no-ca-cert clears the value outright;
both win over everything above and both apply to a reconcile exactly as to a first install. If the
migration hook fails on a TLS handshake, the script prints the hook's own log and then the re-run
that fixes it, with --ca-cert already filled in.
A re-run with no --version stays on the version the release is on. It reads that version off
the release rather than from Docker Hub, and says which revision it read it from, so a re-run to
fix a CA or an issuer cannot also move the release between images. With --image, the image tag
stays the one the release runs, which on a mirror can differ from the chart version; pass
--version to move it. Once the release exists, a re-run also stops prompting and stops generating
a signing key.
Secrets across a re-run
A repeated helm upgrade preserves API_KEY_SECRET, VAULT_SIGNING_KEY, POSTGRES_PASSWORD and
METRICS_AUTH_TOKEN: the chart reads the existing release Secret and keeps what it finds. Losing
the vault signing key would break verification of every record already written, so this matters
more than the others.
A re-run that changes no value leaves the API and worker pods running, the first re-run after a
fresh install included. Their checksum/config and checksum/secret annotations hash the
ConfigMap and the Secret as the chart writes them, so a value change in either rolls the pods and
a re-run that writes both unchanged does not.
AGLEDGER_INSTANCE_ID is kept the same way, from the release's ConfigMap. The chart generates a
UUID for it on first install, so every release has its own: it is the prefix the Server's external
anchors are written under (vault-anchors/<id>/) and, with federation on, the identity peers
store. The Server records the id it first boots with in its database and refuses a later boot under
a different one, naming both. Set the instanceId value when the chart cannot read its ConfigMap
back: under a renderer (below), or for a release installed over a database another release wrote,
such as a restore into a new release, where the post-restore Job prints the database's id.
That preservation only works under a real helm install or helm upgrade. A renderer, meaning
Argo CD, Flux, or helm template | kubectl apply, cannot read cluster state, so the chart sees no
existing Secret and generates fresh values on every sync. On any GitOps path, supply the secrets
yourself through secrets.existingSecret, pin instanceId to a UUID (with existingSecret the
render refuses until you do), and set secrets.gitops: true, which turns each silent regeneration,
the instance id's included, into a named render failure instead.
External database
Step 3 is the external-database path, and it holds for any PostgreSQL 17 or later: managed,
HA or self-managed. Use sslmode=verify-full. On this release require and verify-ca validate
the certificate and hostname exactly as verify-full does, but the PostgreSQL driver's next major
release gives them libpq's meaning, under which require validates nothing; verify-full is the
one spelling that means the same in both. When the server certificate chains to a CA the image
does not already trust, set config.nodeExtraCaCerts to the CA bundle's path inside the
container. The image carries the AWS RDS / Aurora bundle; for Amazon Aurora or RDS, see the
AWS install guide. External-database licensing is per database instance.
database.externalUrl wins over the chart's bundled PostgreSQL. Set both and the release runs no
bundled database and is an external-database release in every other respect too: hook migration,
runtime-role preflight gate, and no TLS exemption, so the production sslmode= requirement
stands. Set neither and the render fails by name: database.externalUrl is required when postgres.bundled.enabled is false.
The migration role needs superuser. The schema installs an event trigger,
agledger_block_audit_drop, which is what stops the audit chain being dropped, and
CREATE EVENT TRIGGER is superuser-only in PostgreSQL. A managed database hands you an owner
role with CREATEDB and not this, so migrations fail with permission denied to create event trigger unless you grant it:
| Provider | Grant |
|---|---|
| Amazon RDS / Aurora | GRANT rds_superuser TO <role>; |
| Google Cloud SQL | GRANT cloudsqlsuperuser TO <role>; |
| Azure Database for PostgreSQL | GRANT azure_pg_admin TO <role>; |
| Self-managed | ALTER ROLE <role> SUPERUSER; |
Give that role to migrations only: set secrets.databaseUrlMigrate (Helm) or
DATABASE_URL_MIGRATE (Compose) to its URL, and leave database.externalUrl / DATABASE_URL as
agledger_app.
agledger_app needs a login. The baseline migration creates it without a usable one. Give it
one the way step 3 does: set AGLEDGER_APP_ROLE_PASSWORD where the migration runs
(migrate.extraEnv on the chart, or that key in secrets.existingSecret) and put the same password
in the agledger_app URL. Every migration run then sets agledger_app's login to that password and
grants it CREATE on the database. To manage the password yourself instead, leave
AGLEDGER_APP_ROLE_PASSWORD unset and run ALTER ROLE agledger_app LOGIN PASSWORD '<app-password>';
as the migration role once the first migration has created the role. Until then the preflight hook
cannot log in as agledger_app and fails the install; re-run it after the ALTER ROLE. Whatever
role you serve as, the requirements in What you provide hold on every deployment path.
Supplying your own Secret. With secrets.existingSecret set, the chart writes no Secret and
reads its keys from yours, under these names: DATABASE_URL and API_KEY_SECRET (required),
VAULT_SIGNING_KEY (required unless signing.kmsKeyArn is set), and optionally
VAULT_SIGNING_KEY_PREVIOUS, WEBHOOK_ENCRYPTION_KEY, DATABASE_URL_MIGRATE,
DATABASE_URL_DIRECT, AGLEDGER_APP_ROLE_PASSWORD, AGLEDGER_LICENSE, AGLEDGER_LICENSE_KEY
and METRICS_AUTH_TOKEN (required if you enable monitoring.serviceMonitor). The
secrets.existingSecret comment in the chart's values.yaml says what each one does. Restart both
Deployments after any change to it; nothing in the chart notices.
Backups of an external database need a client at least as new as the server. pg_dump refuses
to dump a server newer than itself, and distributions lag: Ubuntu 24.04 ships client 16 against the
PostgreSQL 18 this product is validated on. backup.sh reconciles the two for you, falling back to
a matching client in a container when the host's is older, so the only case that needs your
attention is a host with neither an adequate client nor docker.
OpenShift
Add --set openshift.enabled=true to the helm install above. Everything else on this page is
unchanged.
helm install agledger oci://registry-1.docker.io/agledger/agledger-chart \
--version 2.0.0 --namespace agledger --create-namespace \
--set openshift.enabled=true \
--values values.yaml
Without it, admission rejects every pod before it starts:
unable to validate against any security context constraint:
runAsUser: Invalid value: 65532: must be in the ranges: [1000700000, ...]
OpenShift's default restricted-v2 SCC assigns each namespace its own uid range and admits pods
with MustRunAsRange. The chart's stock pod security context asks for uid/gid 65532, which is
outside that range on essentially every cluster. The flag drops that block from every pod the chart
creates (the api and worker Deployments, the migrate Job, and on the external-database path the
runtime-role preflight Job) and lets the platform assign a uid instead. It also widens the api
NetworkPolicy to admit the openshift-ingress namespace, where OpenShift's own router runs. The
bundled PostgreSQL follows the same rule: its pod-level block asks for uid/gid 70, the postgres
user of the upstream image, and the flag drops that too.
What goes is runAsNonRoot: true alongside the uid/gid keys and the RuntimeDefault seccomp
profile, so non-root stops being asserted in the manifest. It still holds, by a different guarantee: on OpenShift restricted-v2 enforces it
directly along with a default seccomp profile, and elsewhere it rests on the image's own
USER 65532:65532. Every container-level control is untouched either way, so each container keeps
its read-only root filesystem, no privilege escalation, and all capabilities dropped.
The image supports an arbitrary uid: it declares USER 65532 for plain Kubernetes and Docker, but
nothing in it is owned by or hardcoded to that uid, its files are world-readable, and it needs no
passwd entry.
./scripts/helm-install.sh sets the flag for you when it sees the security.openshift.io API group
on the cluster, unless you already set openshift.enabled yourself, whether by --set or in a
values file it was passed. The chart also ships
values-openshift.yaml, which sets the same flag for anyone who prefers a values file, though
reaching it means helm pull --untar first:
helm pull oci://registry-1.docker.io/agledger/agledger-chart --version <ver> --untar
helm install agledger ./agledger-chart -n agledger \
-f ./agledger-chart/values-openshift.yaml \
-f values.yaml
What was measured. Under the constraints
restricted-v2imposes on a container (--user 1000700000:0 --read-only --cap-drop ALL --security-opt no-new-privileges), the shipped image applies its migrations, boots healthy, serves the unauthenticated discovery surfaces, mints a platform key, and notarizes a record, so signing works under an assigned uid. The rendered manifests were checked directly: the flag removes the pod-level uid request from every workload and leaves container hardening identical to a stock render. Admission itself was not exercised on a live OpenShift cluster; that part follows from the SCC rules.
Run more than one Server
You can run more than one Server and link them so chains reference records across Servers (we
call linking Servers federation). Each Server is a full, independent install of this guide with
its own database, database role, signing key and instance id; linking is configured after both are
healthy. The chart generates a distinct instance id per release, so two Servers can anchor to one
bucket without reading each other's anchors; do not copy one release's instanceId into another.
If you host several Servers on one PostgreSQL cluster, give each its own role: PostgreSQL roles
are cluster-wide, so Servers sharing the agledger_app role share one credential across all of
their databases.
A per-Server role does not exempt you from the role requirement in What you provide. Each
Server's role still has to own its database's schema objects or be a member of agledger_app:
GRANT agledger_app TO "server_a_app" WITH INHERIT TRUE, run after migrating, because it is the
migration that creates agledger_app and grants it anything.
Membership is what makes the two rules meet, and it is also the reason separate role names alone
do not isolate the databases. Roles and their memberships are cluster-wide, and agledger_app is
one role, created once by whichever database migrated first: every later database's migration
grants to that same role, so a member of it inherits privileges in all of them. If isolation is the
point of the separate roles, take away the reach as well as the name. Revoke CONNECT on each
database from PUBLIC and grant it only to that database's own role.
Air-gapped install
Nothing in install or runtime depends on agledger.ai, Docker Hub, or npm. For a restricted
network, mirror the image (agledger/agledger:2.0.0) and chart
(agledger/agledger-chart:2.0.0) into your internal registry and set image.repository (and
image.pullSecrets) to it. The chart runs three more images, each with its own override:
postgres.bundled.image (also the migrate Job's wait-for-database container), backup.image, and
tests.image. The Server makes outbound calls only to endpoints you configure: webhook receivers,
federation peers, OIDC issuers, and optional integrations such as an OTLP collector or AWS KMS.
The signature verifies inside the enclave. Each release attaches
agledger-2.0.0-offline-verification.tar.gz to the
install repository's release: the
signature and attestation bundles, the signed index and platform manifests, and the Sigstore
trusted_root.json they verify against. Check that archive's own signature with
cosign verify-blob while you still have a network, carry it in, and run
./scripts/verify-release.sh --bundle-dir against your mirrored image. No registry, Rekor or
Docker Hub is needed at that point. The procedure, including the no-registry enclave case, is the
install repository's air-gap guide
(air-gap/README.md).
Mirror the index, not one architecture. A plain docker pull and docker push copies only the
image for the machine doing the mirroring, so the digest your registry then reports is a
per-architecture child, not the multi-arch index digest pinned above. Pin that child and you have
pinned one architecture, which will not schedule on nodes of the other. Copy the whole index with
oras cp -r docker.io/agledger/agledger:2.0.0 <your-registry>/agledger:2.0.0, which also carries
the signature and attestations as OCI referrers so your registry serves them, then read the index
digest back from your registry and pin that.
Uninstall
$ helm uninstall agledger -n agledger
release "agledger" uninstalled
This leaves your database and signing key intact, so a reinstall against the same database and
key resumes the same chain. The release's ConfigMap and Secret are Helm hooks, which
helm uninstall leaves in place, so a reinstall under the same release name and namespace keeps
the instance id. A reinstall under another name against the same database needs instanceId set to
the id the database carries, or the Server refuses to boot and names both.
Next
Keep the Server healthy over time with Day-2 Operations, and set up a backup and recovery routine.