Install AGLedger on Amazon EKS

This guide brings up a single AGLedger Server on Amazon EKS, reachable on a real hostname over TLS, using AWS-native services: an AWS Load Balancer Controller Application Load Balancer (ALB), an AWS Certificate Manager (ACM) certificate, and Amazon Aurora PostgreSQL as the database. It is the AWS companion to Install AGLedger on Kubernetes; read that first for the base model (signing key, readiness gate). The examples use the hostname agledger-aws.agledger.ai; substitute your own.

A fresh 2.0 install applies one migration, 001_consolidated.sql, the 2.0 baseline, and the migrate Job reports Complete 1/1. 2.0 does not upgrade a 1.x database in place. The migration refuses an Aurora database a 1.x release migrated before it changes anything: install 2.0 on a new, empty Aurora database and keep the 1.x database for the 1.x install that wrote it. Version upgrades in Day-2 operations covers that refusal and the migration knobs for later 2.x upgrades.

Pin the index digest, not a per-architecture child. Release images are multi-arch (linux/amd64 and linux/arm64), and the digest pinned below is the multi-arch index. It resolves per-node: on an arm64 (Graviton) node the runtime is native arm64, migrations apply, and signatures produced there verify against the offline verifier on x86.

On the external-database path the migrate Job is a pre-install and pre-upgrade hook, so a migration pod that cannot schedule fails the whole install. If your API pods are pinned to a node class, set migrate.nodeSelector / .tolerations / .affinity, each of which defaults to the matching api.* value.

AWS prerequisites

1. Prepare the Amazon Aurora database

Create a dedicated database and an application role. Run this from inside the VPC (Aurora is not publicly reachable), for example a one-off psql pod connecting as the Aurora master user.

CREATE ROLE agledger_aws_app LOGIN PASSWORD '<app-password>';
GRANT agledger_aws_app TO agledger;          -- master must be a member to set ownership
CREATE DATABASE agledger_aws OWNER agledger_aws_app;

Migration privilege (important). The migration role needs superuser, and on Aurora / RDS that is rds_superuser. A plain database-owner role is not enough: the migration fails with permission denied to create event trigger.

GRANT rds_superuser TO agledger_aws_app;

Serve as a role that does not own the tables (recommended). With the commands above, the API and worker connect as the role that owns every table, and an owner passes every privilege check on its own tables. The schema's UPDATE, DELETE and TRUNCATE revokes on the audit chain's append-only tables do not bind it, so it can rewrite or delete chain rows, and tamper-evidence rests on signatures and external anchors alone. The install works; the preflight Job's warning, which helm-install.sh prints at the end of the run, and GET /v1/admin/ops-summary (vault.appendOnly.enforced: false) both say so. To close the gap, keep agledger_aws_app for migrations and serve as agledger_app, the role the baseline migration creates:

agledger_app is one role for the whole Aurora cluster, and every migration sets its password. Where other Servers on the cluster already use it, give this one the password they use, or serve as a role of this Server's own that holds its grants: CREATE ROLE <role> LOGIN PASSWORD '<generated>'; GRANT agledger_app TO <role> WITH INHERIT TRUE;

Why the superuser grant is needed, what the runtime role requires of its own, and the pg_dump client-version rule for backups are in External database on the Kubernetes install guide.

The connection string uses sslmode=verify-full. The production boot check also accepts require and verify-ca, and on this release the PostgreSQL driver treats both as verify-full: the Aurora certificate is checked against the CA bundle and its hostname against the endpoint, so an untrusted CA fails with self-signed certificate and a hostname mismatch fails too. With either of those modes set, the driver logs a SECURITY WARNING at boot saying so, and that its next major release gives them libpq's meaning, where require performs no certificate validation. verify-full means the same in both, so use it. The agledger image bundles the AWS RDS / Aurora root CA at /etc/ssl/certs/rds-global-bundle.pem; set config.nodeExtraCaCerts to that path so the Server can validate the Aurora server certificate.

Adding uselibpqcompat=true to the connection string opts into libpq's meaning now. The Server refuses to boot on uselibpqcompat=true with sslmode=require or sslmode=verify-ca and no sslrootcert=, because that combination either skips validation or fails every connection.

2. Values for the AWS path

aws-values.yaml:

image:
  digest: "sha256:6e29779624951807dbcf6adeb5c6c7b08da153339fab7ec6e1210e4ebfc66a05"  # 2.0.0
database:
  poolMax: 20
config:
  externalUrl: "https://agledger-aws.agledger.ai"          # the Server's signed issuer identity
  nodeExtraCaCerts: "/etc/ssl/certs/rds-global-bundle.pem"  # bundled AWS RDS / Aurora CA
  trustProxy: "10.0.0.0/16"   # the ALB's addresses: the same VPC CIDR as networkPolicy.ingressCIDRs
ingress:
  enabled: true
  className: alb
  annotations:
    alb.ingress.kubernetes.io/scheme: internet-facing
    alb.ingress.kubernetes.io/target-type: ip
    alb.ingress.kubernetes.io/healthcheck-path: /health
    alb.ingress.kubernetes.io/listen-ports: '[{"HTTP":80},{"HTTPS":443}]'
    alb.ingress.kubernetes.io/ssl-redirect: "443"
    alb.ingress.kubernetes.io/certificate-arn: "arn:aws:acm:us-west-2:<acct>:certificate/<id>"
  hosts:
    - host: agledger-aws.agledger.ai
      paths: [{ path: /, pathType: Prefix }]
  tls: []   # TLS terminates at the ALB via the ACM cert above; no in-cluster TLS secret
networkPolicy:
  ingressCIDRs:
    - 10.0.0.0/16   # your VPC or subnet CIDR(s); see note below

NetworkPolicy on the ALB path. The chart's default NetworkPolicy admits ingress only from ingress-controller pods (ingress-nginx / traefik). With target-type: ip, the ALB sends traffic and health checks from VPC elastic network interfaces, not from a pod, so the default policy blocks the ALB and targets never go healthy. Set networkPolicy.ingressCIDRs to your VPC or subnet CIDRs rather than disabling the policy: it covers egress as well as ingress, so networkPolicy.enabled: false also removes every egress restriction on the API pod, not just the ingress block the ALB needs past.

Anchoring to S3. The chart's egress admits port 443 on public addresses, which covers S3 and an S3 gateway endpoint. An S3 interface endpoint answers on private addresses, and the default policy refuses them: add its subnet CIDRs on 443 under networkPolicy.extraEgress. The anchor prefix is the release's AGLEDGER_INSTANCE_ID, which the chart generates per release, so two releases can share one anchor bucket; Anchoring has the variables.

Client addresses behind the ALB. Every request reaches the API from an ALB node, so without config.trustProxy the Server reads the ALB's address as the client's. The refused-credential lockout, rate-limit buckets, API key allowedIps and audit-row addresses then all key on the ALB: one agent retrying a stale key locks out every new credential arriving through that ALB node for the window. Set config.trustProxy to the CIDR the ALB connects from (your VPC CIDR, as above) so the Server reads X-Forwarded-For only when the ALB wrote it. Do not set it to true, which trusts whatever the client sends in that header.

Rolling updates behind the ALB. With target-type: ip, also label the namespace so the AWS Load Balancer Controller injects its pod readiness gate (kubectl label namespace agledger-aws elbv2.k8s.aws/pod-readiness-gate-inject=enabled). Without it, a rolling update counts a new pod Ready while its ALB target is still registering, and the old pod is removed before the new one receives traffic.

3. Install

Generate the vault signing key and create the platform key exactly as in Install (steps 2 and 4). Keep the Aurora URL and the signing key out of the values file, and out of the command line too: --set puts a value in helm's argv, where ps shows it to every other user on the machine for the length of the install. Both of these carry a secret, the database password and the private key every record is signed with, so pass them as files.

$ kubectl create namespace agledger-aws
$ umask 077
$ printf %s 'postgresql://agledger_aws_app:<pw>@<aurora-endpoint>:5432/agledger_aws?sslmode=verify-full' > db-url
$ printf %s '<vault-key>' > vault-key
$ helm install agledger oci://registry-1.docker.io/agledger/agledger-chart \
    --version 2.0.0 --namespace agledger-aws \
    --values aws-values.yaml \
    --set-file database.externalUrl=db-url \
    --set-file secrets.vaultSigningKey=vault-key
NAME: agledger
STATUS: deployed
REVISION: 1
  API URL:  https://agledger-aws.agledger.ai

$ rm db-url vault-key

The scripted alternative, and the Aurora CA it applies for you

./scripts/helm-install.sh --db <url> does the same install and handles the Aurora certificate without you naming it. When you pass a database and have not named a CA, it sets config.nodeExtraCaCerts to /etc/ssl/certs/rds-global-bundle.pem, the AWS RDS and Aurora root bundle baked into the image, which is exactly what sslmode=verify-full needs to validate the Aurora server certificate. Without it, the migration hook fails the TLS handshake before anything starts.

$ ./scripts/helm-install.sh --version 2.0.0 --namespace agledger-aws --release agledger \
    --values aws-values.yaml \
    --db 'postgresql://agledger_aws_app:<pw>@<aurora-endpoint>:5432/agledger_aws?sslmode=verify-full'

--db takes the URL as a separate argument and passes it to helm with --set-file, so the password never reaches helm's argv. A CA you name yourself, in --set config.nodeExtraCaCerts or in a values file, wins over the default; so does a CA the release already carries. --ca-cert <path> names a different bundle and --no-ca-cert clears it, and both apply to a re-run exactly as to a first install.

The script installs with helm upgrade --install, so correcting a value and running it again reconciles the release rather than refusing its name. If the migration hook fails on a TLS handshake, it prints the hook's log and then the exact re-run with --ca-cert filled in. A re-run with no --version keeps the version the release is already on, read off the release rather than from Docker Hub, so fixing a CA cannot also move the image.

$ kubectl rollout status deploy/agledger-agledger-chart-api -n agledger-aws --timeout=180s
deployment "agledger-agledger-chart-api" successfully rolled out
$ kubectl get pods,jobs -n agledger-aws
pod/agledger-agledger-chart-api-...      1/1   Running
pod/agledger-agledger-chart-migrate-...  0/1   Completed
pod/agledger-agledger-chart-worker-...   1/1   Running
job.batch/agledger-agledger-chart-migrate   Complete   1/1   7s

4. Point DNS at the ALB

The Ingress provisions an ALB; read its hostname and create a DNS record for your host (CNAME, or a Route 53 alias) pointing at it.

$ kubectl get ingress -n agledger-aws
NAME                      CLASS   HOSTS                      ADDRESS
agledger-agledger-chart   alb     agledger-aws.agledger.ai   k8s-agledger-agledger-....us-west-2.elb.amazonaws.com

Wait for the ALB target to register as healthy:

$ aws elbv2 describe-target-health --target-group-arn <tg-arn> \
    --query 'TargetHealthDescriptions[].TargetHealth.State'
[ "healthy" ]

5. The readiness gate: named, TLS-terminated, signing

Reach the Server on its real hostname over HTTPS. The ALB serves the ACM certificate.

$ curl -s https://agledger-aws.agledger.ai/health
{"status":"ok","version":"2.0.0","timestamp":"...","signingKey":{"gate":"usable","keyId":"c4ddafd6bf06f1ef"}}

$ echo | openssl s_client -connect agledger-aws.agledger.ai:443 \
    -servername agledger-aws.agledger.ai 2>/dev/null | openssl x509 -noout -subject -issuer
subject=CN = *.agledger.ai
issuer=C = US, O = Amazon, CN = Amazon RSA 2048 M04

$ curl -s -o /dev/null -w "HTTP %{http_code} -> %{redirect_url}\n" http://agledger-aws.agledger.ai/health
HTTP 301 -> https://agledger-aws.agledger.ai:443/health

The served certificate is issued by Amazon (ACM), and plain HTTP is redirected to HTTPS by the ssl-redirect annotation. The keyId in /health is your vault-key fingerprint, and the published key list carries the same one:

$ curl -s https://agledger-aws.agledger.ai/v1/verification-keys
{"data":[{"keyId":"c4ddafd6bf06f1ef","algorithm":"Ed25519","status":"active",…}],…}

(Response elided. The full key entry and the rest of the readiness gate are in step 5 of the Kubernetes install guide.)

The Server is healthy, reachable on its name over an ACM-issued certificate, backed by Aurora, and signing with your key.

Upgrading the Aurora engine (PostgreSQL 17 → 18)

Aurora supports an in-place major-version upgrade (set the cluster's engine version to 18.x with allow_major_version_upgrade). The Server's data, audit chain, and signatures carry through it unchanged, and the API reconnects on its own once the database returns, with no pod restart. Plan for a short write outage during the upgrade and run it between workloads, not during one.

Adopting native uuidv7 after the upgrade. On PostgreSQL 17 the schema migration installs a small uuidv7() polyfill in the public schema (PostgreSQL gained a native uuidv7() in 18). A fresh 18 install never creates it. But an in-place upgrade does not adopt the native function automatically: the polyfill persists, and every table's id column default stays bound to it, so inserts keep using the polyfill.

No shipped migration adopts the native function, so for any 2.0 install whose schema was created on 17, run the block below once, as a superuser, against your application database after the upgrade. It is idempotent and self-discovering: it acts only when native uuidv7 exists and the polyfill is present, and it finds the columns itself.

-- Re-point every uuidv7() column default so it re-resolves to the native pg_catalog.uuidv7,
-- then remove the now-unused polyfill. No-op on a fresh 18 install or an already-remediated DB.
DO $$
DECLARE r record;
BEGIN
  -- only act if BOTH native and the public polyfill exist
  IF EXISTS (SELECT 1 FROM pg_proc p JOIN pg_namespace n ON n.oid = p.pronamespace
             WHERE p.proname = 'uuidv7' AND n.nspname = 'pg_catalog')
     AND EXISTS (SELECT 1 FROM pg_proc p JOIN pg_namespace n ON n.oid = p.pronamespace
                 WHERE p.proname = 'uuidv7' AND n.nspname = 'public') THEN
    FOR r IN
      SELECT n.nspname AS sch, c.relname AS tbl, a.attname AS col
      FROM pg_attrdef ad
      JOIN pg_class c       ON c.oid = ad.adrelid
      JOIN pg_namespace n   ON n.oid = c.relnamespace
      JOIN pg_attribute a   ON a.attrelid = ad.adrelid AND a.attnum = ad.adnum
      WHERE pg_get_expr(ad.adbin, ad.adrelid) ILIKE '%uuidv7%'
    LOOP
      EXECUTE format('ALTER TABLE %I.%I ALTER COLUMN %I SET DEFAULT uuidv7()', r.sch, r.tbl, r.col);
    END LOOP;
    DROP FUNCTION public.uuidv7();
  END IF;
END $$;

After it runs, new inserts use native uuidv7() and the public.uuidv7 function is gone. (Do not DROP FUNCTION public.uuidv7() on its own: the column defaults depend on it, so a bare drop fails and CASCADE would strip the defaults and break inserts. The re-point above is what makes the drop safe.)

AWS licensing

An external database such as Aurora is licensed by Enterprise Edition, one license per database instance; the free Developer Edition key covers the bundled PostgreSQL only. With no key applied the install is Unlicensed Use, which the license permits for evaluation, development and testing only, and the Server logs a license warning on boot and once a day after that. From 45 days after the database was initialized, the same notice also reaches every successful /v1/admin/* response (a Warning header, plus a nextSteps entry on every mutation), /v1/admin/ops-summary, /v1/conformance, and the SIEM stream as a license.notice_escalated audit row about once a day, until a key is applied. A Developer Edition key on Aurora, which is outside that edition's scope, gets the same notice. Nothing is gated. Step 5 of the Kubernetes install guide shows the license read-back.

To license through AWS, subscribe on the AWS Marketplace listing, then set marketplace.productId to prod-gdyk7ehkopbnm and name an IRSA role in marketplace.serviceAccountAnnotations. The role needs license-manager:CheckoutLicense and license-manager:CheckInLicense. With a product ID set, the Server calls License Manager CheckoutLicense every 15 minutes, and at boot as well when you hold no license key of your own, since a key you hold settles the tier locally. That call goes to AWS, not to AGLedger, carries no record, agent or usage data, and fails open, so an unreachable entitlement service blocks nothing. The listing's usage instructions carry the full command, including the Marketplace ECR repository (709825985650.dkr.ecr.us-east-1.amazonaws.com/ag-ledger/agledger) and its docker login. That repository holds the linux/amd64 image only, a byte-for-byte copy of the amd64 image in the signed Docker Hub release, so schedule it on amd64 nodes.

Out-of-band delivery is the alternative and reaches no network: secrets.license for a compact agl_ent_v1_... key, or license.keyFile.* to mount a PEM from a Kubernetes Secret.

Signing with a KMS key

By default the chart holds the vault signing key as key material in its Secret. On AWS the alternative is a key whose private half never leaves KMS: create an asymmetric key with key spec ECC_NIST_P256 and usage SIGN_VERIFY, grant the Server's IAM role kms:GetPublicKey and kms:Sign on it, and point the chart at it instead of secrets.vaultSigningKey:

signing:
  kmsKeyArn: "arn:aws:kms:us-west-2:<acct>:key/<id>"
serviceAccount:
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::<acct>:role/agledger-signing
config:
  allowNonDefaultSigningAlg: true   # KMS signs ES256, not Ed25519

The Server fetches the public half at boot, registers it under the same fingerprint scheme as a local key, and signs every chain entry, checkpoint, Receipt, signed webhook delivery and certificate with one KMS Sign call. /v1/verification-keys then lists the key with algorithm ES256, and GET /v1/admin/vault/signing-keys reports signer.backend aws-kms. KMS signs no Ed25519, so this is an ES256 chain, which is what the opt-in acknowledges. Every consumer verifying it needs @agledger/verify 2.0.0 or newer, as on any 2.x install. Rotation is the same stage-then-retire sequence as any key change, with a new ARN and a restart as the staging step.

If KMS stops answering, the Server refuses to write rather than write unsigned: writes answer 503, /health/ready reports signingKey.gate signer_unreachable so the ALB routes around the pod, and the process resumes on its own once a probe Sign is answered. The chart's alert rules carry AGLedgerVaultSignerUnreachable for it. KMS Sign quotas are per account and region and shared with every other caller, so signing.remoteMaxPerSecond caps each pod and the sum across api and worker replicas has to sit under the quota.

Air-gapped / private registries

To run from Amazon ECR instead of Docker Hub, mirror agledger/agledger:2.0.0 into ECR, set image.repository to the ECR repository and image.pullSecrets (or use the node role / IRSA for ECR pull), and pin image.digest.

Mirror the whole index, not the one architecture your machine pulls, and pin the index digest your registry reports back. The copy commands are in Air-gapped install on the Kubernetes install guide, along with the offline verification flow and the install repository's air-gap guide (air-gap/README.md) that carries it. An ECR mirror that holds only the amd64 child runs on amd64 nodes when you pin the child digest; the index digest does not exist in that repository.