Resetting an AI Agent Mid-Delegation: Catching Up From the Record
By Michael Cooper · Founder
Updated 2026-09-26 with a second run: the work defined in the customer's own record types, nothing handed back to the agent after the reset, and no output shape from the harness. Its numbers replace the first run's.
An agent runs when something triggers it: a message arrives, a schedule fires, an event lands. It does its part of the task and the process ends, and the next trigger starts it with an empty context. That is the design, not a fault: a fresh start costs less and reasons better than a context carrying every earlier step. Long work spans many runs. The agent delegates part of a job and ends; the next run checks status (not done) and ends; a later run checks again (done) and carries on.
Each fresh start takes everything the agent knew about the work: who asked for it, who authorized it, what it sent to other agents, and where each piece stands. After every one, the agent has to catch up from something outside its context. Between independent agents, which A2A 1.0 describes as “independent, potentially opaque AI agent systems” that do not share internal state, memory, or tools, nobody else holds that for it. Whatever it did not write down is gone.
That is why AGLedger has agents write the work down before, during and after, and read the record to catch up. We tested one step of that loop: three independent agents passed a purchase request down a two-hop chain, and we reset the middle agent's context right after it delegated. 48 runs across three model vendors, half on A2A alone and half with AGLedger records alongside.
With the record, the agent caught up on the work it had already delegated before sending anything in 21 of 24 runs. On A2A alone, 8 of 23; most of the rest sent the work out a second time.
The setup
- Orchestrator takes the request from plant staff and delegates it.
- Procurement coordinator checks the budget and delegates the sourcing. Its context is cleared immediately after its first successful send.
- Sourcing agent (the remote agent at the end of the chain) searches supplier quotes and returns an allocation.
The request carried twelve checkable details, from a PO reference and cost center to a 60% single-supplier cap and the name of the approver. Each agent ran in its own container with its own model, credentials and disk, on the official A2A Python SDK (a2a-sdk 1.1.5, protocol 1.0). Claude Sonnet 4.6, OpenAI gpt-oss-120b and DeepSeek V3.2 rotated through the three positions, eight runs per arrangement.
On A2A alone, the ask and the results travel as A2A messages, and each agent keeps a private notes file. With records, the work is defined in three record types of the customer's own on AGLedger 1.8.0: a purchase request that requires every detail of the ask, a sourcing record that requires every constraint on the way down and each supplier's own price, date, terms and certification on the way up, and a recommendation that requires everything the approver needs. Each agent also checkpoints its working state there.
After the reset, nothing was handed back. Every agent had a durable store of the A2A tasks it had received and had to call it to learn what it was working on; everything else it had to find for itself.
Catching up on work already sent out
On A2A alone, the id of the delegated task exists only once the send succeeds, and the reset came right after. The way back is whatever the coordinator wrote down first, or asking the remote agent. 11 of 24 coordinators had written notes before the reset, and 6 of those picked the work back up; 2 of the rest found the task another way. 15 of 23 sent the sourcing work out again, which in a real procurement chain is a second sourcing run on the same order.
With records, the delegation is written down before the work is sent: the child record exists first, and the parent lists it. After the reset, 21 of 24 coordinators read their way back to it before sending anything, and 16 carried on with the running task. Only 8 of 24 sent the work out again.
| Coordinator model | Carried on, A2A alone | Carried on, with records |
|---|---|---|
| Claude Sonnet 4.6 | 5 of 8 | 8 of 8 |
| DeepSeek V3.2 | 3 of 8 | 8 of 8 |
| gpt-oss-120b | 0 of 7 | 0 of 8 |
The model decides the rest. With records, gpt-oss-120b found its child record and sent the task again anyway in 5 runs, and sent it again without looking in 3. To make a repeat harmless, derive the completion's Idempotency-Key from the record id, so AGLedger replays the first submission or refuses a different one instead of letting a second run replace it.
Checking status on long work
The loop only works if “not done” eventually becomes “done” or “failed”. With gpt-oss-120b as coordinator, three record runs ended up with a second sourcing child after the reset. The parent settles only when every child is finished, so each duplicate had to be dealt with. In one run the orchestrator asked the coordinator to cancel it. In another, neither child ever received a result, and the run sat at “not done” until the 30-minute cap.
A deadline on every delegated record is what ends that. In a separate test on EKS, a child record with a 45-second deadline that was never delivered went EXPIRED four seconds past it, and its parent moved straight to its owner's decision with the child marked failed. The next status check reads a failure the agent can act on, not another “not done”.
What else we looked at
The ask arrived intact in both arms. All twelve details reached the final deliverable in 23 of 24 runs either way. The record types asked for every fact while the writing agent still had it, and the agents filled every required field the first time in all but one write.
Guessed values were rare in both. A value with no source turned up in 4 A2A-alone runs (supplier dates overwritten with the deadline, made-up certificate numbers, a made-up PO reference) and in 1 record run, a wrong supplier date in a completion the schema accepted. A schema checks shape, not truth.
Writing it down costs tokens. The median record run billed 677K tokens across the three agents, against 67K on A2A alone, and took 216 seconds against 106. Most of that went on record calls and checkpointing: an agent loop pays for every earlier tool result again on every later model call, so verbose responses compound. Leaner responses and checkpoints only at the points an agent will catch up from both cut into it; neither was measured here.
Neither arm got the cheapest allocation every time. Some sourcing agents chose a costlier split that still met every constraint (2 runs on A2A alone, 3 with records). A required field carries a rule to the agent; it does not make the agent apply it. The record types here did not encode the share cap or the terms floor as gate rules, which the engine can run.
The numbers
| Runs | A2A alone | A2A + AGLedger | p |
|---|---|---|---|
| After the reset, found the delegated work before re-sending | 8/23 | 21/24 | < 0.001 |
| Carried on without re-sending | 8/23 | 16/24 | 0.04 |
| Sent the sourcing task again | 15/23 | 8/24 | 0.04 |
| Run finished | 24/24 | 23/24 | 1.0 |
| All 12 details in the final deliverable | 23/24 | 23/24 | 1.0 |
| All 12 in the coordinator's deliverable | 20/24 | 24/24 | 0.11 |
| Lowest-cost allocation delivered | 22/24 | 20/24 | 0.67 |
| A value guessed anywhere in the run | 4/24 | 1/24 | 0.35 |
| Median billed tokens per run | 67K | 677K | < 0.001 |
| Median seconds per run | 106 | 216 | < 0.001 |
p is from Fisher's exact test for counts and Mann-Whitney for medians. The recovery rows leave out one A2A-alone run whose coordinator died in the harness: a malformed tool call the model provider then refused to replay.
The calls for delegating over REST or A2A are in the A2A guide, and the checkpoint loop an agent catches up from is in the work context guide.
Sources & further reading
A2A v1.0 Specification (Linux Foundation): the “independent, potentially opaque AI agent systems” definition, and ListTasks
a2a-sdk: the official Python SDK the agents ran on (1.1.5)