
Design safer AI agent retries with limits for attempts, time, cost and progress. Learn backoff, idempotency, handoff and Laravel implementation patterns.
An AI agent retry budget is a set of limits that decides how much additional work an automation may spend after failure. A useful budget caps attempts, elapsed time, cost and retries without measurable progress. When any limit is reached, the agent stops, records its state and hands the task to a queue, operator or recovery workflow instead of silently starting over.
That definition is broader than a queue worker's tries = 3. A traditional job usually repeats a known operation. An agent may call a model again, reinterpret the task, choose another tool, modify external state and consume more tokens on every loop. The third attempt is therefore not always the same operation for the third time. It can be a new branch with a new cost and a new failure mode.
After years of building PHP backends, queues and automation, I have learned to treat retries as a resource, not as free resilience. AI agents make that discipline more important. Their ability to improvise is useful while they are making progress. Without a budget, the same ability can turn a recoverable failure into an expensive, opaque loop.
Retry is not the same as recovery
A retry repeats work because the system believes another attempt has a reasonable chance of succeeding. Recovery restores a useful state after something went wrong. Sometimes a retry is part of recovery. Sometimes it only repeats the cause.
The distinction begins with the failure. A connection reset, a short provider outage or a rate limit with an explicit reset time may be temporary. A malformed payload, revoked permission, missing required field or violated business rule will not become valid because a worker waits two seconds and submits it again.
Agent failures add another category: the request may be technically successful while the result is not good enough to continue. The model returns valid JSON but omits required evidence. A browser tool loads a page but reaches the wrong account. A code agent changes files but fails the same test with the same diagnostic. Those are semantic failures. Retrying them unchanged is unlikely to create progress.
A retry policy should therefore answer two questions before it answers “how many times?”
- What evidence says this failure is temporary?
- What will be different on the next attempt?
If neither question has an answer, stopping is often the more reliable behavior.
Why AI agent retries are unusually risky
A conventional HTTP client can often define one operation, one response and a small set of retryable status codes. An agent loop has more moving parts.
- Output is probabilistic. Another model call may produce a better answer, the same mistake or an entirely different plan.
- Context grows. Tool results, error messages and prior attempts consume tokens and can make later calls slower or more confused.
- Actions have side effects. Sending a message, creating an invoice or updating a record cannot be treated like a failed read.
- Success is partly semantic. A 200 response does not prove the task is complete or correct.
- Retries can exist at several layers. The HTTP library, queue worker, model SDK and agent orchestrator may all retry independently.
This is why “let the agent try until it works” is not a production strategy. The loop needs explicit termination conditions and a definition of progress. JetBrains Research makes the same point in its discussion of agent loops: useful loops combine validation and repair with progress checks, retry budgets and clear termination conditions. The correct limit also depends on risk. A local formatting task can tolerate more autonomy than an operation that publishes, pays or deletes.
Use a retry envelope, not one maximum-attempt number
I model an agent's retry policy as a retry envelope. This is my practical extension of the retry-budget idea used in distributed systems. The envelope has four dimensions, and crossing any one of them stops automatic execution.
1. Attempt budget
This is the familiar limit: how many times may this operation or stage run? It protects against unbounded loops and gives queue infrastructure a simple hard stop. The important detail is scope. Count attempts for the logical operation, not separately inside every library that touches it.
There is no universal correct number. A read-only lookup might allow several attempts. An ambiguous payment write might allow none until its status has been reconciled. The number follows the operation's risk and the evidence that another attempt can help.
2. Time budget
A job can remain under its attempt limit while becoming uselessly late. If a response is only valuable within two minutes, a fourth attempt after ten minutes is not recovery. Record a deadline when the run starts and compare every proposed delay with the remaining time.
The deadline should include model latency, tool execution and scheduled backoff. A queue's job timeout is a final safety mechanism; it is not a substitute for a business deadline.
3. Cost budget
Each model call, search, browser session and external API request consumes something. Tokens are the obvious unit, but cost can also mean paid tool calls, compute time or scarce provider quota. Track the accumulated amount per run and, when useful, per tenant or workflow.
A fixed attempt count cannot control this well. One attempt with a large context and several tools may cost more than ten small validations. The next call should be authorized against estimated remaining cost before it starts, not merely recorded afterward.
4. Progress and side-effect budget
This is the agent-specific part. Define what meaningful progress looks like: a different failing test, fewer validation errors, a newly retrieved artifact, a changed state hash or a completed checkpoint. If consecutive iterations produce the same result, the loop is spending budget without reducing uncertainty.
Side effects make the limit stricter. A read can often be repeated. A write needs an idempotency key, reconciliation step or explicit approval. If the system cannot determine whether a previous write succeeded, the next action should usually be “check state,” not “write again.”
Together, the four limits prevent a common mistake: an agent appears compliant because it stayed under three attempts while it exceeded the acceptable time, cost or operational risk.
Classify the failure before spending the budget
A retry budget limits work; a failure classifier decides whether the work is worth attempting. I use categories like these:
| Failure type | Example | Default action |
|---|---|---|
| Transient transport | Connection reset, temporary DNS error | Retry with capped backoff and jitter |
| Rate limited | HTTP 429 with reset metadata | Wait as instructed if the deadline permits |
| Provider unavailable | HTTP 503 | Retry within both local and global budgets |
| Deterministic input | Schema validation fails | Repair input or stop; do not replay unchanged |
| Authorization or policy | Revoked token, denied action | Stop and escalate |
| Ambiguous write | Timeout after submitting a side effect | Reconcile by idempotency key before any repeat |
| No semantic progress | Same failing test and same patch twice | Stop or change strategy explicitly |
The classifier should use structured signals where possible: status codes, exception types, validation results and provider headers. Asking a model to decide whether its own last failure is retryable can be one input, but it should not override hard rules around permissions, deadlines or side effects.
This connects to handling partial failure in AI automation. A workflow can complete three stages and fail the fourth. Restarting the entire workflow wastes work and may duplicate actions. Resume from a durable checkpoint or execute a targeted compensating step.
Give retry ownership to one layer
Retries multiply when every layer tries to be helpful. AWS gives a sharp example: in a five-deep service stack, three retries at each layer can increase load on a failing database by 243 times. The arithmetic is simple: each layer can fan one request into three attempts, producing 3⁵ calls at the bottom.
An AI automation stack can create the same pattern:
- the HTTP client retries a model request;
- the model SDK retries the stream;
- the tool wrapper retries the whole tool call;
- the agent loop retries the step;
- the queue retries the job.
No individual setting looks extreme, but their product is. Choose one layer to own retries for the logical operation. Lower layers may handle a narrowly defined transport reconnect only when that behavior is visible to the owner. Otherwise, configure them for one attempt and let the orchestrator account for every repeat.
A system-wide budget also matters. Google SRE describes retry budgets as a way to prevent retries from overwhelming a service, giving an illustrative policy that allows only 60 retries per minute before requests fail. That number is not a recommendation for every system. The principle is: even if one job has budget left, the fleet may not.
Backoff and jitter are necessary, but they do not create progress
Exponential backoff spaces attempts further apart. Jitter adds randomness so that thousands of workers do not wake at the same instant. Both reduce synchronized pressure on a recovering dependency.
A typical capped schedule might resemble 1, 2, 4, 8 and then 15 seconds, with random variation around each delay. The cap prevents the wait from growing without bound. The deadline prevents the capped delay from continuing forever.
Backoff does not make a deterministic error retryable. It only changes when the same work happens. It also cannot repair a loop whose prompt, context and tools remain unchanged. For an agent retry, require one of these before spending another attempt:
- a transient dependency has had time to recover;
- the input or plan changed in a traceable way;
- new evidence became available;
- the next attempt uses a defined fallback with different capabilities;
- a human approved a risky continuation.
If none applies, backoff only makes the failure slower.
A retry budget cannot replace idempotency
When an operation has side effects, the system needs to answer a dangerous question: did the previous attempt fail before or after the action happened?
Suppose a tool submits an invoice and the connection times out before the response arrives. Repeating the request may create a duplicate. Refusing to continue may leave the workflow uncertain. The correct design is an idempotency key or an external operation identifier that lets the caller query the outcome and safely associate repeated requests with the same logical action.
I covered the broader pattern in idempotency in AI automation. The retry policy should reference that identity directly. Every attempt belongs to one logical operation, and its durable record stores the external key, known state and reconciliation result.
Some tools cannot offer idempotency. For those, reduce the automatic side-effect budget, add a read-after-timeout reconciliation path, or require approval. Human approval is a control boundary, not a failure of automation.
A practical Laravel implementation pattern
Laravel's queue features provide attempts, timeouts, backoff and failed-job handling. I still keep the agent's logical retry envelope in application data because one run can span several jobs and tools.
A durable automation_runs record can contain:
- status and current checkpoint;
- attempts used for each logical operation;
- started-at and deadline timestamps;
- tokens or estimated cost consumed;
- last progress fingerprint;
- idempotency and external operation keys;
- last classified failure and next eligible time;
- stop reason and handoff context.
The decision function can remain intentionally boring:
final class RetryDecision
{
public function for(Run $run, Failure $failure): Decision
{
if (! $failure->isRetryable()) {
return Decision::stop('non_retryable');
}
if ($failure->mayHaveCommittedSideEffect()) {
return Decision::reconcile();
}
if ($run->attemptsExhausted()
|| $run->deadlineExceeded()
|| $run->costBudgetExceeded()
|| $run->hasStoppedMakingProgress()) {
return Decision::handoff('retry_budget_exhausted');
}
return Decision::retryAfter(
Backoff::withJitter($run->attempts())
);
}
}
The exact classes are less important than centralizing the decision. The queue worker reports the failure and asks for a decision. It does not independently invent another loop. Use a database transaction or atomic update when reserving budget so that two workers cannot both believe the final attempt is available.
If the workflow already relies on Laravel queues for AI automation, the queue retry count remains a final execution guard. The application-level run record expresses the business policy across job boundaries.
Budget exhaustion is a designed state, not an exception to hide
A production automation needs a useful destination after retries stop. “Failed” is too little information. The handoff should preserve:
- the original objective and current checkpoint;
- the stop reason and which budget was exhausted;
- attempt count, elapsed time and cost consumed;
- the last useful output and validation result;
- known side effects and unresolved external operations;
- the safest available next action.
A dead-letter queue can hold the durable work item while an operator or recovery process decides what happens next. It should not be a graveyard that automatically replays the same message every night. I discuss that distinction in dead-letter queues for AI automation.
JetBrains recommends a similar handoff for bounded agent loops: record why the loop stopped, how many iterations ran and what state remains. A crucial rule follows: the agent must not silently restart itself or increase its own limit. The budget belongs to the system's policy, not to the component consuming it.
Observe retry quality, not only retry count
A dashboard that shows total retries is a start. It does not tell whether retries helped. I want to see:
- first-attempt success rate;
- recovery rate by attempt number and failure class;
- time and cost spent on successful versus exhausted runs;
- retries without a changed progress fingerprint;
- reconciliation outcomes for ambiguous writes;
- global retry volume by provider and operation;
- handoff age and dead-letter queue depth.
If the fourth attempt almost never recovers a workflow, remove it. If one provider error reliably clears after a documented reset interval, tune that path specifically. If no-progress stops are rising, investigate the prompt, validator or tool contract rather than raising the limit.
These signals belong in the same traces and logs described in observability for production AI agents. Each attempt should share a run ID and operation ID, while retaining its own attempt number and decision reason.
A retry policy checklist
- Define the logical operation and give retry ownership to one layer.
- Classify failures with structured signals before retrying.
- Set limits for attempts, deadline, cost and no-progress iterations.
- Apply a global budget so many healthy-looking jobs cannot overload one dependency.
- Use capped exponential backoff with jitter for temporary failures.
- Require idempotency or reconciliation before repeating side effects.
- Persist checkpoints and budget consumption outside process memory.
- Make budget reservation atomic across concurrent workers.
- Stop when the next attempt has no defined source of new information.
- Hand off with enough context for a human or recovery workflow to continue safely.
- Measure whether later attempts actually recover work.
Bounded persistence is reliable automation
Retries are valuable because temporary failures are real. Networks drop, providers throttle and models occasionally produce repairable output. The answer is not to avoid retries. It is to make their cost and authority explicit.
A maximum-attempt number protects against one kind of loop. A retry envelope protects the outcome: it limits how many times the agent acts, how long the result may take, how much it may cost and how long it may continue without evidence of progress.
The strongest automation is not the one that never gives up. It is the one that knows when another attempt is useful, when state must be reconciled and when the safest next step belongs to someone else.
A retry is justified by new probability of success, not by remaining patience.
Sources and further reading
- Google SRE Book: Addressing Cascading Failures
- Google SRE Book: Handling Overload
- AWS Builders' Library: Timeouts, retries, and backoff with jitter
- JetBrains: AI Agent Loops
- Better Stack: Exponential Backoff
Sources were checked on October 2, 2026. The four-dimensional retry envelope is a practical engineering framework, not a universal standard; appropriate limits depend on workload risk, provider behavior and service objectives. This article reflects experience with PHP backends, queues, automation and AI agents without identifying an employer or client. It was prepared with AI assistance and editorial review. Cover: network router photograph from Pixabay.
