Error Handling: Retry and Idempotency
An AI agent that calls tools runs over a network — and networks are unreliable. APIs answer with 429 or 503, connections drop, models take three seconds one time and thirty the next. Error handling is the set of patterns an agent uses to absorb these disruptions instead of throwing away the whole task at the first hiccup: retries with backoff, idempotency, timeouts, and compensation.
That sounds like classic backend engineering, and it is. The difference with AI agents: they call tools autonomously and in loops, often without a human watching every single step. A failure a human would notice immediately can get repeated ten times unnoticed by an agent — with damage multiplied to match.
Why naive agents fail in production
A naive agent treats every tool call as guaranteed to succeed. When one fails, the agent either gives up entirely or blindly calls the tool again — with no idea whether the first attempt actually went through. Analyses of LLM API traffic show a noticeable share of all calls ending in an error, most of them caused by rate limits. Production agent systems are estimated to fail on transient issues in roughly ten to fifteen percent of calls.
Without countermeasures, these small hiccups cascade into real problems: duplicate tickets, duplicate emails, duplicate orders. The agent itself often has no idea it’s triggering the same action a second time — all it sees is “call failed, so try again.”
The core mistake
Retry without idempotency isn’t reliability — it’s a multiplier for side effects. Before you add retry logic, the action behind it has to be designed so that running it multiple times produces the same result as running it once.
Retries with backoff
A retry is the simplest reflex: if a call fails, try it again. Done naively — instantly, with no pause — this makes worse the exact problem that caused the failure in the first place. Hit a rate limit, and the agent just fires off more requests in a shorter time, deepening the throttling.
The standard countermeasure is exponential backoff with jitter: the wait time between attempts doubles after every failure, and a random scatter value stops many agents from retrying in the same rhythm at once. A common formula is base delay times two to the power of the attempt number, plus a random jitter value. In practice that often means: a first wait of one to two seconds, doubling after each attempt, and giving up after five to seven tries, surfacing the error instead.
Not every failure deserves a retry. A timeout or a 503 is usually transient — retrying makes sense there. A 400 caused by bad input stays exactly as bad on the second attempt. An agent should distinguish by failure type instead of treating every failure the same way.
Idempotency: the precondition for safe retries
Idempotency means: running an action multiple times has the same effect as running it once. A GET request is naturally idempotent — it changes nothing. A “create ticket” or “send email” call is not: every call creates a new ticket or sends a new message.
The standard pattern for making such actions retry-safe is the idempotency key. Before the first attempt, the agent generates a unique ID for the action and sends it along with every call. The receiving side remembers which IDs it has already processed — if the same ID comes in again, it simply returns the stored result instead of executing the action again. Major payment APIs like Stripe have done this for years, and the pattern has become the standard for state-changing agent tools.
One thing matters here: the ID has to be fixed before the call and stay the same across every retry — not be regenerated on each attempt. Otherwise the key loses its purpose.
Rule of thumb
Every tool that changes state — creating a ticket, sending an email, placing an order — gets an idempotency key before it gets any retry logic at all. That order isn’t negotiable.
Setting timeouts right
A missing timeout is just as dangerous as a missing retry — only in the opposite direction. Without a time limit, an agent can wait indefinitely for a response that never arrives, blocking the whole pipeline behind it. Every tool call needs an explicit time limit, after which the agent treats the attempt as failed and moves into retry or error logic.
The difficulty is getting the dose right: too short, and the agent aborts calls that would have actually succeeded — creating unnecessary retries and, at worst, unnecessary duplicates. Too long, and a single hanging request can stall the entire flow for minutes. A good starting point is the actual response-time distribution of the specific tool under load, not one blanket value for every call.
Compensation and rollback
Retry and timeout handle individual calls. But what happens when a multi-step flow fails partway through — say, step three of five fails after steps one and two have already triggered real side effects? This is where the saga pattern comes in: instead of one real transaction spanning every step, each step gets a matching compensating action that undoes its effect.
If step three fails, the agent calls the compensation for step two, then the one for step one — in reverse order, last started, first undone. A reserved payment gets cancelled, a created record gets deleted again, a sent notification gets replaced by a correction. Importantly, the compensation itself has to follow the same reliability patterns as the original action — a rollback can fail too, and needs its own retry logic.
No compensation, no production readiness
A multi-step agent workflow without a compensation path isn’t a finished system — it’s a prototype. The moment a flow produces real side effects across multiple steps, the question “what happens if step X fails?” needs an answer for every single step.
How it all fits together
These four patterns — retry, idempotency, timeout, compensation — aren’t an edge case for rare emergencies; they’re the baseline equipment of any production agent workflow. How multiple agents coordinate their work, and where error handling sits in the orchestration layer, is covered in Multi-Agent Orchestration. The API-side practice around rate limits and retry-friendly calls is covered in Working with the LLM API. And which frameworks already bring retry and state-management logic along is covered in Agent Frameworks at a Glance.
FAQ
- Retry is the act of trying a failed call again. Idempotency is the property of the underlying action that ensures running it multiple times causes no extra damage. Retry without idempotency creates duplicates; idempotency without retry alone brings no reliability — the two belong together.
- The wait time between attempts doubles after each failure, plus a random jitter value so many agents don’t retry in the same rhythm at once. A common setup is a starting wait of one to two seconds, doubling per attempt, and giving up after five to seven tries.
- Only state-changing calls strictly need one — anything that creates or changes data, or triggers an action in the real world. Pure read operations are usually naturally idempotent and don’t need an extra key.
- A timeout decides when a call counts as failed because it took too long. Retry decides what happens next — whether and how often to try again. Without a timeout, the agent never knows when it should even move to a retry.
- As soon as a flow consists of multiple steps with real side effects and a later step can fail after earlier ones have already taken effect. A single, self-contained call usually only needs retry and idempotency — a multi-step workflow additionally needs a compensation path for each step.
What is the difference between retry and idempotency?
How exactly does exponential backoff work?
Should every tool call get an idempotency key?
What is the difference between a timeout and a retry?
When do I need compensation logic instead of a simple retry?
Discover more
Bulk Content Engine: How Context and RAG Tags Make the Orchestrator Smarter
My Bulk Content Engine now pauses and resumes at any point, because the orchestrator maintains its own context — plus RAG tags per task.
EncyclopediaHuman-in-the-Loop — Approval Gates for AI Agents
What human-in-the-loop means in agent workflows: approval gates, intervention points before critical actions, and why they are mandatory for irreversible steps.
NewsExecution engine without an IDE: tickets from the dashboard to any repo
You write a ticket in the boostN dashboard and it runs on the right repository — no IDE needed. Multiple repos, in parallel, in seconds.