Designing Idempotent Side-Effect Contracts for AI Agents

Designing Idempotent Side-Effect Contracts for AI Agents

An AI agent calls send_refund. The payment provider accepts the request, but the response disappears during a network timeout. The runtime sees an exception and retries. The second request also succeeds.The trace now looks reassuring:send_refund attempt=1 timeout send_refund attempt=2 success agent run success The customer sees two refunds. This is a synthetic example, but the failure mode is ordinary distributed-systems behavior. The agent did not need a stranger prompt or a better model. It needed a tool contract that distinguished an unsuccessful response from an unsuccessful effect.That distinction matters anywhere an agent can change the world: sending a message, creating a ticket, issuing a credit, booking a trip, deleting a record, or triggering a deployment. If the runtime treats every exception as permission to call again, “self-healing” becomes “self-duplicating.”Developers often describe a tool with an input and output schema:type RefundInput = { paymentId: string; amountCents: number; reason: string; }; type RefundOutput = { refundId: string; status: "accepted"; }; That schema validates shape, not behavior. It does not tell the orchestrator:whether the tool is read-only;whether repeating it can create another effect;whether it supports an idempotency key;how to determine the outcome after a timeout;which failures are safe to retry; orwhether human approval applies to one attempt or one business operation.HTTP defines an idempotent method as one whose intended effect is the same after multiple identical requests, while also noting that a server may still log each request or create other non-idempotent side effects. That careful definition is useful for agent tools: the contract should be about the intended business effect, not merely about receiving one 200 response.The Model Context Protocol now includes tool annotations such as readOnlyHint, destructiveHint, and idempotentHint. Those annotations help a client present and route tools, but the MCP maintainers explicitly describe them as untrusted hints, not security or correctness guarantees. A runtime still has to enforce its own policy.Define a Side-Effect ContractI prefer to make retry behavior part of the tool definition rather than scatter it across prompts and catch blocks.type SideEffectClass = | "read_only" | "idempotent_write" | "deduplicated_write" | "non_repeatable_write"; type RetryPolicy = { maxAttempts: number; retryableCodes: readonly string[]; requiresReconciliationAfterUnknown: boolean; }; type ToolContract = { name: string; sideEffect: SideEffectClass; validateInput(input: unknown): I; execute(input: I, context: ToolExecutionContext): Promise; retry: RetryPolicy; reconcile?: ( operationId: string, signal: AbortSignal, ) => Promise>; }; The four classes force useful questions:read_only: another call should not change business state.idempotent_write: the target system guarantees one intended effect for equivalent calls.deduplicated_write: safety depends on a stable operation or idempotency key.non_repeatable_write: an automatic retry is forbidden unless reconciliation proves the first attempt did not commit.The names are less important than making the policy executable.Generate the Operation Identity Before the First AttemptAn idempotency key created inside each retry loop does not provide idempotency. Every attempt gets a new identity and therefore looks like a new operation.Persist the operation before calling the external system:type OperationState = | "prepared" | "in_flight" | "committed" | "rejected" | "outcome_unknown"; type OperationRecord = { operationId: string; runId: string; toolName: string; inputFingerprint: string; state: OperationState; providerReference?: string; attempts: number; }; async function prepareRefund( runId: string, input: RefundInput, ): Promise { return operations.insertOnce({ operationId: crypto.randomUUID(), runId, toolName: "send_refund", inputFingerprint: fingerprintValidatedInput(input), state: "prepared", attempts: 0, }); } fingerprintValidatedInput should canonicalize a schema-approved, secret-free representation. A raw JSON.stringify is not a safe universal canonicalizer: key ordering, unsupported values, and sensitive input all need deliberate handling. RFC 8785 defines a JSON Canonicalization Scheme when interoperable hashing or signing is required.Every attempt then reuses the same operationId:async function executeRefund( operation: OperationRecord, input: RefundInput, signal: AbortSignal, ): Promise { await operations.transition(operation.operationId, "in_flight"); try { const result = await payments.refund(input, { idempotencyKey: operation.operationId, signal, }); await operations.markCommitted(operation.operationId, result.refundId); return result; } catch (error) { if (isDefinitiveRejection(error)) { await operations.transition(operation.operationId, "rejected"); throw error; } await operations.transition(operation.operationId, "outcome_unknown"); throw new UnknownOutcomeError(operation.operationId, { cause: error }); } } The crucial line is not the idempotency header. It is the transition to outcome_unknown. A timeout tells you that observation failed. It does not tell you that the effect failed.Unknown Outcomes Need Reconciliation, Not HopeWhen the provider exposes an operation lookup, query it before allowing another write:async function recoverRefund( operationId: string, signal: AbortSignal, ): Promise { const result = await payments.findRefundByIdempotencyKey( operationId, { signal }, ); if (result.status === "found") { await operations.markCommitted(operationId, result.refundId); return "committed"; } if (result.status === "definitively_missing") { return "not_found"; } return "still_unknown"; } Only not_found can authorize a retry for a non-repeatable operation. still_unknown should pause, escalate, or schedule later reconciliation. It should not become a creative prompt asking the model what to do.If the target system has no lookup and no deduplication support, that is part of the tool's contract. The safe behavior may be to require a human to inspect the external system. Automation cannot manufacture a guarantee that the dependency does not provide.Approval Must Bind to the OperationSuppose a user approves a $500 refund, the first attempt times out, and the runtime generates a second operation ID. Has the user approved the second potential refund?No. Approval should bind to a canonical operation intent:type ApprovalGrant = { approvalId: string; operationId: string; intentHash: string; approvedBy: string; expiresAt: string; }; A retry with the same operation ID and unchanged intent can consume the same valid grant. A changed amount, recipient, destination, or operation identity requires a new approval. This prevents “retry” from becoming an accidental privilege expansion.Use an Outbox When Local State and Dispatch Must AgreeSometimes the agent records a decision in your database and publishes work to a queue. Writing the row and publishing the message as unrelated actions creates another ambiguity: the process can crash after either one.The transactional outbox pattern places the business record and an outbox event in the same database transaction. A separate dispatcher publishes the event and records delivery attempts. Consumers still need deduplication because brokers and dispatchers commonly provide at-least-once delivery.agent decision | v [database transaction] operation row + outbox row | v outbox dispatcher --may retry--> message broker | v consumer deduplicates by operationId This does not create magical exactly-once execution. It makes each boundary explicit and recoverable.Trace the Effect LifecycleOne tool.error=true field is not enough. Record the state machine without leaking raw payloads:operation.prepared id=op_73 tool=send_refund operation.attempted id=op_73 attempt=1 operation.outcome_unknown id=op_73 reason=deadline_exceeded operation.reconciled id=op_73 result=committed agent.completed id=run_12 outcome=refund_confirmed Useful evidence includes the operation ID, sanitized intent fingerprint, state transitions, attempt count, policy decision, approval reference, reconciliation method, and provider reference. Never put secrets, full payment details, or uncontrolled user content into trace attributes.For every side-effecting tool, I want written answers to these questions:What is the intended business effect?Is that effect naturally idempotent, deduplicated, or non-repeatable?Where is the stable operation identity created and stored?Which errors are definitive, retryable, or outcome-unknown?How does the runtime reconcile an unknown outcome?What changes invalidate approval?How are duplicate messages or callbacks handled?What evidence proves the final external state?The model can propose an action. The runtime owns the effect.A mature agent platform should never translate “I did not receive a response” directly into “do it again.” It should translate it into “the outcome is unknown—find out what happened.”ReferencesRFC 9110: Idempotent MethodsRFC 8785: JSON Canonicalization SchemeModel Context Protocol: Tool Annotations

Original Source

Read the full article at Hackernoon →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.