Blog
Research findings, compliance guidance, and engineering insights.
Resetting an AI Agent Mid-Delegation: Catching Up From the Record
Agents start fresh on every run. We reset one right after it delegated: with a record it caught up on the work 21 times in 24, without one 8 times in 23.
The agent you authorized in the morning is not the agent that acted in the afternoon
Anthropic traced six weeks of quality reports to three product-layer changes with no new model weights. Authority was granted to a composition that changed.
Three ways to measure a Postgres queue table, three different answers
An 11.5-hour heap series at 8,000 jobs/s, measured three standard ways: +3.2% a day, settled, and a 0.84% band. The heap is bounded. The indexes reached 2.4 GB and were still growing.
Retention does not protect a Postgres queue from a pinned xmin horizon
Three staggered 30-minute snapshots took a pg-boss job table from 1.1% to 54% dead tuples at an identical completed workload. Both arms ran exactly 59 autovacuums.
A local agent in front of a cloud agent: gpt-oss-120b on llama.cpp
A self-hosted gpt-oss-120b handles two classification jobs on an always-on ops host and decides whether the cloud agent runs at all. The llama.cpp systemd unit, the three-gate ladder, per-job effort settings, and where the line between local and cloud sits.
We Shipped A2A 1.0 Support Twice. The First Time Was Wrong.
Adopt, revert, negotiate: we implemented the A2A 1.0 wire format from the spec, tore it out six weeks later because no deployed client could talk to it, then reintroduced it as a per-request negotiated dialect. What that arc taught us about specs versus ecosystems.
Lessons from an In-Place PostgreSQL 17 to 18 Upgrade (the uuidv7 Default That Did Not Switch Over)
An in-place Aurora 17.9 to 18.3 upgrade kept inserting with the uuidv7 polyfill because column defaults are OID-pinned. The trap, the re-point-then-drop fix, the benchmarks (native is 2.6x faster and packs a 33%-smaller index), and a ~10-minute write-outage that self-healed.
Your people are already using AI. Make being compliant the easy path.
Shadow AI is a path problem, not a discipline problem: 90% of executives feel in control while 52% of workers use unsanctioned AI. Pave the compliant path so agents take it on their own, and the audit-ready record accrues as a byproduct.
Settlement Evidence for Agentic Commerce: What AP2, ACP, and x402 Leave Out of Scope
The 2026 agentic-payment rails sign the authorization moment, then declare dispute evidence and settlement out of scope - in their own spec language. AP2, Verifiable Intent, Stripe SPTs, and x402 mapped against the settlement-evidence layer none of them carries.
Can an AI Agent Be Trusted to Write Its Own Audit Log? We Measured It.
Four production-tier models processed payment batches with forced write failures, then wrote their own audit reports. Three of four claimed success for writes that never happened - up to 47%. The independent signed chain caught every phantom record.
Durable Intent, Measured: We Reset Four Agents and Asked Them to Finish Their Own Work
A cold-recovery experiment: four agents restarted with zero memory had to discover and finish their own work from the signed ledger alone. All four stalled behind a judgment gate; a deterministic gate flipped the weakest model from 0/3 to 3/3.
The 88,000-Token Crash: Making gpt-oss-120b Survive Its Full 128K Context on Strix Halo
Past ~80k tokens of prefill, llama.cpp on Strix Halo trips the amdgpu GPU watchdog: ring reset, vk::DeviceLostError, core dump. Flash attention does not fix it; -ub 512 does. The crash matrix, the throughput bill, and 3/3 needle-in-a-haystack at 110k tokens.
SCITT as the Article 12 Implementation Pattern for Autonomous AI Agents
EU AI Act Article 12 requires automatic event logging for traceability, and names no standard. SCITT is the IETF-track standard designed for this evidence shape. The article-by-article mapping, the gaps SCITT explicitly leaves open, and a concrete implementation pattern.
The advisory-lock self-deadlock Postgres can't see
PostgreSQL's deadlock detector walks the row-lock wait graph. Advisory locks held across `await` boundaries can form cycles through application code that the detector cannot see. Reproduction, real diagnostic output, and the structural fix.
Near Frontier-Quality LLM, No Cloud, No Subscription, Unlimited Tokens: gpt-oss-120b on Strix Halo + Ubuntu 26.04
A $2,300 96 GB Strix Halo box runs gpt-oss-120b at ~48 tok/s with a stable full 128K context on llama.cpp + Vulkan. Rewritten after a full rebuild - including what the May version of the post got wrong.
pg-boss in production: footguns we hit and how to avoid them
Eleven operational footguns we hit running pg-boss at AGLedger. Four are API-time and seven configuration-time. Self-contained reproductions, citations, and the patterns we settled on.
Cutting PostgreSQL Audit-Report Query Time 44% with GROUPING SETS and Materialized CTEs
Six aggregations per request became two. Total DB time dropped 44% under sustained load. The anti-pattern, the rewrite, and the caveat about why wall-clock p50 did not move.
AGLedger Performance at Scale: A Simulated Quarter, Measured and Verified Offline
A quarter of claims operation on the published 1.2.0 image: 353,861 API calls in 86 minutes, zero failures, then 265k records verified offline in under a minute on a desktop CPU.