For: Architects and engineers connecting AI agents to systems that create or change records.
The failure is in your knowledge, not necessarily in the operation#
Consider a hypothetical service team. An agent drafts an internal follow-up task, a person approves it, and a connector asks the CRM to create it. The CRM commits the task. Before the response arrives, the connection breaks. The agent sees an error and creates a second task with slightly different wording. Both tasks now exist, although the conversation contains only one approval.
This is not primarily a reasoning problem. The system has confused two questions: ‘Did the call return?’ and ‘Did the business effect occur?’ A stronger model may make the duplicate sound more plausible. A longer system prompt cannot recover a missing receipt.
The Amazon Builders’ Library explains the same ambiguity for ordinary distributed APIs: retrying a request after a lost response can create an extra resource. Agentic systems inherit that problem and add another route to duplication—the model can replan and express the same intent as a new tool call.
Keep effect state outside the conversation#
Use a durable operation record owned by the application, not a sentence in chat history. Separate the identity of the user’s goal, the intended effect and each network attempt. One goal can have several effects; one effect can have several attempts. A new model response must not silently create a new effect identity.
Propose
Normalize the intended effect and persist its operation identity.
Authorize
Check policy, required approval and current resource preconditions.
Dispatch
Claim the operation and send the exact authorized parameters.
Reconcile
Record verified completion, known rejection, or an unresolved outcome.
On a narrow screen, scroll the table sideways to see every column.
| State | What is known | Permitted next step |
|---|---|---|
| Proposed / awaiting approval | No attempt is allowed yet. | Validate, preview and obtain any required approval. |
| Ready | The action is approved, subject to current authorization and preconditions. | One executor claims and dispatches it. |
| In flight / unknown | An attempt may have reached the target. | Observe or replay the same operation only under a proven idempotency contract. |
| Succeeded | Authoritative evidence identifies the completed effect. | Report the result; do not create it again. |
| Rejected / not applied | The target provides evidence that the effect was not applied. | Correct or retry only if policy, approval and the operation contract still permit it. |
| Needs investigation | Available evidence cannot establish the outcome safely. | Stop further writes for this effect and give the named owner the evidence. |
The important property is not the choice of labels. It is that every transition has an owner and an evidence requirement. A timeout must not transition directly to ‘not applied’. Likewise, an HTTP acknowledgement may mean ‘queued’, not ‘completed’; the connector must understand the target’s actual completion contract.
Approval belongs to an action, not to the whole conversation#
In the CRM example, the preview should identify the customer record, task owner, due date and task contents. Store those normalized parameters with the operation. The executor should receive that object, not ask the model to reconstruct it after the person presses Approve.
Changing the customer, recipient, amount, scope or other material parameter means changing the effect. Revalidate it and obtain new approval where the policy requires one. Bind approval to the action version, approver, expiry and intended use. A digest can detect a changed object, but a hash by itself is not authorization.
This follows the distinction in OWASP’s transaction-authorization guidance: meaningful details must be verifiable, authorization must be enforced server-side, and modifications must not inherit approval for the original transaction. For an agent, the same principle applies to an internal write even when no money moves.
Also recheck current authority and resource state. Yesterday’s approval does not override today’s revoked permission or a record that has changed. Where the downstream system supports conditional writes, bind the change to the expected resource version. If it cannot enforce the necessary precondition, document the race and restrict the capability instead of implying that a preflight read makes the operation atomic.
Idempotency is a contract, not a random header#
An idempotency key identifies the same intended operation across attempts. Persist it before dispatch and reuse it unchanged. Do not generate it inside a retry loop, and do not let an agent invent a fresh key because the first request was inconvenient. Two intentionally separate but identical tasks need different operation identities; equal parameters alone do not prove duplicate intent.
The receiver must actually honour the key. Check its scope, retention window, payload-mismatch behaviour, concurrent-request handling and result semantics. Stripe’s idempotency documentation provides a useful concrete example: it describes stored responses, parameter comparison and key pruning after keys are at least 24 hours old. That is one provider’s contract, not a universal window for your agent’s memory or retry policy.
A durable workflow engine can preserve progress, but it does not automatically make an external effect idempotent. Temporal’s explanation of durable execution and idempotency distinguishes recorded workflow progress from activities and duplicate executions. Know which layer can retry the write: SDK, connector, worker and workflow retries can interact.
At your boundary, atomically claim an operation so two workers cannot independently dispatch it. Handle expired leases and late workers explicitly. A local database transaction cannot atomically commit a change in an unrelated SaaS system. If the receiver has neither a suitable idempotency mechanism nor an enforceable uniqueness constraint, your ledger helps explain uncertainty; it cannot promise exactly-once effects.
A missing search result is not always proof of absence#
Return to the lost CRM response. The connector should first look for the operation’s external reference or a target-supported unique business key, under an identity authorized to inspect that result. Finding one exact matching task with the expected fields can resolve the operation as succeeded. Finding a different payload or multiple candidates is a conflict to investigate—not an invitation to select whichever looks closest.
What if the lookup returns nothing? A search index may be delayed. The write may still be running. The recovery identity may not have permission to see the record. A request with a fresh operation ID could create a duplicate even after several empty searches. ‘Not found’ becomes ‘not applied’ only when the target’s semantics and timing justify that inference.
- If the target supports a still-valid idempotency contract: replay the same operation and parameters only as the contract permits.
- If a read can establish completion but not absence: use bounded observation, then escalate unresolved uncertainty.
- If the target offers an authoritative operation-status endpoint: use its terminal state and documented meaning.
- If neither safe replay nor reliable observation exists: stop automatic writes and route the evidence to a person. Consider keeping this capability draft-only.
Bound observation too. Give it a deadline, rate limit and owner; otherwise a supposedly safe recovery loop can consume the entire run budget. Tell the user ‘We are checking whether the task was created’ rather than claiming failure or success. Keep the operation reference available to support without exposing private payloads in the interface.
Stopping an agent does not undo an effect#
A stop control should prevent new dispatches and cancel pending work where possible. It cannot unsend a request already accepted by an external system. After stopping a run, continue the narrowly authorized observation needed to resolve in-flight operations, or hand that responsibility to a named operator. Otherwise the stop button can hide the very effects you need to inspect.
Compensation is a separate business action. Deleting a newly created CRM task may be appropriate before anyone acts on it, but not after a colleague has added work or the task triggered another workflow. A compensating change needs its own authorization, identity and evidence. Some effects have no meaningful inverse.
A recovery contract you can take to a design review#
Complete this for every state-changing tool before enabling autonomous retries. It is an architecture worksheet, not a production-ready configuration. If a field cannot be answered, record the missing capability and reduce the allowed automation until it can.
Use a tabletop test with an engineer from the downstream system. Ask them to show the documented semantics or a controlled demonstration for each recovery claim. ‘The API normally returns quickly’ is not evidence about what happens when a response is lost.
Run the five-interruption exercise#
For the hypothetical CRM task, pause execution at each boundary below. Write the durable record, next permitted transition and evidence that would unlock it. A good answer can say ‘needs investigation’; it cannot silently invent completion.
- After approval, before dispatch: change the task’s customer or revoke the initiating user’s access. Does execution still proceed?
- After the target commits, before the response arrives: can the same effect be recovered without a second task?
- After a worker loses its lease: let the old and replacement workers both wake up. Which mechanism prevents conflicting dispatches?
- After the idempotency window expires: resume an old run. Does it observe and escalate, or mistakenly create a fresh operation?
- After cancellation: deliver a late success response. Does the system record the real effect and decide separately whether compensation is appropriate?
This extends the bounded-execution and approval-gated write exercises in Agentic AI Architecture & Security. For the broader release decision, combine the recovery contract with an evaluation that measures actual outcomes and unsafe effects. Reliability is demonstrated by the failure paths your system can explain—not by how often a happy-path demo finishes.
Sources and further reading
Primary sources checked on . Worked examples and worksheets are Ardevant teaching material; sources do not imply endorsement.
- AWS Builders’ Library — Making retries safe with idempotent APIs
Explains lost-response ambiguity and explicit caller request identity. The agent-specific state model and worksheet here are Ardevant’s design guidance.
- Stripe API — Idempotent requests
A concrete provider contract for response reuse, parameter comparison and key retention; not a universal contract for other APIs.
- Temporal — Idempotency and durable execution
Clarifies workflow progress, activities and idempotency. No workflow product is required to use this article’s reasoning.