Falling dominoes illustrating compensating actions in AI automation
← Back to Blog
October 2026·AI Engineering·12 min read

Compensating actions for AI automation when rollback is impossible

Compensating actions make partial AI automation failures recoverable by recording what changed, how to repair it and when humans must decide.

Compensating actions for AI automation are business operations that repair the effect of a completed step when a normal rollback is no longer possible. They do not rewind time. They move the system from a partially completed state into a new, valid and explainable state.

This distinction matters as soon as an AI workflow leaves one database transaction. An agent can create a CRM record, reserve inventory, send a message, publish a document and update an invoice in one run. If the fourth step fails, rolling back a row in PostgreSQL does not unsend the message or make an external API forget the earlier calls.

I treat every external write as a small promise. Before the workflow makes that promise, it should know how to confirm it, retry it, compensate for it or hand the decision to a human. That discipline is less impressive than an autonomous demo, but it is what lets automation survive contact with production.

Why a database rollback is not enough

A local transaction is still the best tool when all relevant changes live in one transactional store. Validate the input, change the rows and commit them together. If an exception occurs before commit, the database can restore the earlier state precisely.

AI automation rarely stays inside that boundary. A model call may finish before the database write. A queued job may run twice. A SaaS API may accept a request but time out before returning its identifier. A human may act on an email while the workflow is still executing. Each system has its own clock, failure modes and retention rules.

Consider an onboarding agent that performs five steps:

  1. creates a customer record;
  2. provisions an account in an external service;
  3. generates a contract;
  4. sends a welcome email;
  5. schedules a follow-up task.

If the task scheduler is unavailable after the email was sent, deleting the local customer row would make the situation worse. The customer already received evidence that onboarding succeeded. A useful recovery might keep the account, mark the workflow as incomplete, retry the follow-up task and notify an operator only if its deadline is at risk.

That is why compensation is not a synonym for rollback. A rollback restores a previous technical state. A compensating action applies domain rules to reach an acceptable business state after some effects have escaped the transaction.

Give every side effect a compensation contract

Before I let an agent execute a side effect, I want a short contract for that action. The contract forces the design discussion to happen before an incident rather than during one.

Contract fieldQuestion it answersExample
Operation identityHow do we recognize the same request?onboarding:842:provision:v1
Success evidenceWhat proves the provider accepted it?External account ID and response hash
Retry policyWhich failures are transient, and for how long?Three attempts within ten minutes
CompensationWhat safe action offsets the result?Suspend the provisioned account
Point of no returnWhen does automated reversal become unsafe?After the customer starts using the account
EscalationWho decides when rules are insufficient?Operations queue with deadline and context

The compensation should be explicit business language. “Delete whatever was created” is not a contract. “Suspend account X if it has no completed user activity; otherwise request an operator decision” is one. It states the target, precondition and fallback.

I also separate four outcomes that are often collapsed into one failed status:

  • Failed: the forward step did not complete and created no accepted effect.
  • Compensating: the workflow is actively applying repair actions.
  • Compensated: the repair completed and the resulting business state is valid.
  • Intervention required: automated repair is unsafe, ambiguous or exhausted.

Those states belong in the same explicit AI agent state machine as the happy path. Compensation is not exception cleanup hidden in a catch block. It is a first-class workflow with commands, transitions, deadlines and observable outcomes.

Record an action ledger before executing the action

A workflow cannot compensate reliably if it reconstructs history from logs. Application logs are useful evidence, but they are not a durable plan. I keep an action ledger: one record for every intended external effect, its stable operation key, current state, provider reference, request fingerprint and compensation data.

The important ordering is:

  1. persist the intended action and its compensation contract;
  2. commit that local record;
  3. dispatch the external command;
  4. record the observed outcome;
  5. advance the workflow only from accepted evidence.

This does not make a remote API call atomic with the database. It does make ambiguity visible. If a worker dies after the provider accepts the request but before the result is stored, the ledger still contains the operation key. Recovery can query the provider or retry with the same idempotency key instead of inventing a new action.

The ledger entry should contain enough data to compensate without depending on the model that planned the original step. Store the exact external resource ID, tenant, policy version, immutable artifact hash and relevant preconditions. Do not ask a model to rediscover which account it probably created from a transcript.

Workflow checkpoints capture resumable computation. The action ledger captures external commitments. They complement each other: a checkpoint says where reasoning can resume, while the ledger says what the outside world has already observed.

Compensation must be idempotent too

The repair path runs under the same unreliable conditions as the forward path. Workers crash, acknowledgements disappear and APIs time out. A compensation such as “issue refund” is dangerous if every retry creates another refund.

I give each compensating command its own operation key and state. Before execution, the handler checks whether the exact repair already completed. When the provider supports idempotency keys, it sends the stable key downstream. When it does not, the adapter needs a lookup or reconciliation step based on a provider reference.

final class CompensateProvisioning implements ShouldQueue
{
    public function __construct(public readonly string $actionId) {}

    public function handle(ActionLedger $ledger, AccountProvider $provider): void
    {
        $action = $ledger->lockForCompensation($this->actionId);

        if ($action->compensation_completed_at !== null) {
            return;
        }

        $result = $provider->suspend(
            accountId: $action->external_resource_id,
            idempotencyKey: $action->compensation_key,
        );

        $ledger->markCompensated($action, $result->reference);
    }
}

The database lock in this sketch reduces concurrent execution, but the provider idempotency key is what protects the remote boundary. The same principle applies to the forward action. Idempotency is the hidden contract that turns repeated delivery from a business incident into an ordinary implementation detail.

A practical Laravel workflow design

In a Laravel application, I keep the orchestrator intentionally boring. A relational database owns workflow and action states. Queue jobs carry stable IDs rather than the only copy of context. Provider adapters translate domain commands into external API calls. Policies decide whether to retry, compensate, continue with reduced functionality or pause for approval.

A typical forward transition looks like this:

  1. lock the workflow row and verify its expected version;
  2. create a pending action-ledger record and an outbox message in one database transaction;
  3. commit;
  4. publish the outbox message to the queue;
  5. let a worker execute the provider call using the stored operation key;
  6. write the result and request the next transition.

Laravel can dispatch queued work after database commit. That prevents a fast worker from loading a record that the originating transaction has not committed yet. For stronger delivery guarantees, I still prefer a transactional outbox: the workflow transition and intent to dispatch are committed together, then a separate publisher retries delivery.

The compensation flow uses the same infrastructure. It reads successful actions from the ledger, orders the required repairs according to domain policy and dispatches them. Reverse order is a useful default, not a universal rule. Revoking public access may be more urgent than deleting a draft created later. Some independent repairs can run in parallel.

Each repair has its own retry budget. Transient failures should not send a workflow directly to manual intervention, but infinite retries can hide a permanent policy or data problem. When the budget ends, the job moves to a visible recovery queue with the original action, attempted compensation, provider responses and next safe options.

Design around the point of no return

Some effects cannot be reversed at all. An email cannot be removed from a recipient's inbox. A public disclosure cannot be made unread. A payment may be refundable but still leave fees, exchange-rate differences or accounting entries. A generated legal document may already have been signed.

For these steps, “undo” should not appear in the design. The real choices are prevention, mitigation and escalation.

  • Validate early: resolve permissions, policy checks, required data and budget before irreversible work.
  • Stage artifacts: create a draft and immutable preview before publishing or sending.
  • Move risk to the end: execute compensable steps first and irreversible steps only after critical conditions pass.
  • Require approval: place a human boundary before high-impact actions rather than after damage.
  • Mitigate explicitly: send a correction, revoke access, suspend an account or create a case instead of pretending the original effect vanished.

This is where approval boundaries become part of reliability rather than a limitation on autonomy. A good workflow knows which decisions are cheap to automate and which require context, authority or accountability that the agent does not have.

The model can help classify an incident or draft a recovery option. It should not silently invent compensation policy at runtime. The allowed actions, thresholds and escalation owners must come from versioned application rules. Otherwise the recovery path is less predictable than the failure it is supposed to contain.

Test the broken middle, not only the happy ending

Compensation logic is easy to admire in a diagram and easy to neglect in code because it runs less often. I test it by stopping the workflow at every boundary: before the provider call, after acceptance but before local persistence, during result handling and halfway through compensation.

A useful test matrix includes:

  • duplicate delivery of every forward and compensating command;
  • provider success followed by a client timeout;
  • stale workflow versions and two workers claiming the same action;
  • compensation after another actor changed the resource;
  • expired credentials or permissions on the recovery path;
  • recovery after a process restart and after the queue is rebuilt;
  • retry exhaustion, dead-letter handling and manual completion;
  • auditing from original request to final compensated state.

The operational dashboard should answer business questions, not only worker questions. Which workflows are compensating now? How old is the oldest one? Which external effects remain unconfirmed? How often does compensation fail? Which action type creates the most intervention work?

Alerts need context an operator can act on: customer and workflow IDs, the last confirmed step, the exact ambiguous effect, attempted repairs, deadline, policy version and safe buttons. Sending a stack trace to a generic channel is not a recovery interface. A failed compensation belongs in the same deliberate process as a dead-letter queue: visible, owned and replayable after the cause is understood.

Reliable autonomy includes a way back

Compensating actions do not make distributed automation perfectly atomic. They make failure bounded and explainable. The system records its commitments, distinguishes ambiguity from failure, applies domain-specific repairs and stops when automated recovery would create more risk.

My review question for every production agent is simple: after each tool call succeeds, what new fact exists outside the workflow, and what can the system safely do if the next call never succeeds? If the answer is only “retry the whole run,” the automation is not ready for partial failure.

A reliable agent is not one that never gets stuck. It is one that can show exactly where it stopped, what already changed, what it is doing to repair the situation and when a human must take over. That is the practical value of compensation: not time travel, but controlled progress toward a valid state.

Sources were checked on October 9, 2026: Microsoft Azure Architecture Center on the Compensating Transaction pattern and Saga pattern; AWS Prescriptive Guidance on saga orchestration; and Laravel 13.x documentation on queued jobs and database transactions. The compensation contract, action-ledger structure and Laravel workflow are practical engineering guidance rather than a universal standard. This article reflects experience with PHP backends, queues, automation and AI agents without identifying an employer or client. It was prepared with AI assistance and editorial review. Cover: domino photograph from Pixabay.

Igor Gawrys
Igor Gawrys
AI Engineer & IT Consultant · Katowice, Poland