Control room panels representing explicit AI agent state machines
← Back to Blog
October 2026·AI Engineering·12 min read

Why production AI agents need explicit state machines

Design reliable AI agent state machines with durable checkpoints, safe transitions, reconciliation and a practical Laravel implementation.

An AI agent state machine is an explicit model of the states a workflow may occupy and the transitions allowed between them. The model can propose an action, but deterministic application code validates the evidence, authorizes the transition and stores it. That boundary makes an agent resumable, testable and safer when tools fail or side effects become uncertain.

The first version of an agent workflow often looks like a loop: give the model an objective, let it choose a tool, append the result and ask what to do next. That is enough for a demo. The loop becomes fragile when the process lasts longer than one request, waits for a person, writes to another system or resumes after a worker crash.

After more than seven years of building PHP backends, queues and automation, I have learned that a status field is not administrative detail. It is part of the safety model. An agent can be probabilistic inside a step. The system around it still needs an unambiguous answer to a boring question: what is allowed to happen next?

An agent loop is not a workflow contract

An agent loop answers a local question: should the model think, call a tool, observe the result or stop? A workflow contract answers a wider one: which durable state is the business process in, and what evidence permits it to move?

Those questions overlap, but they are not interchangeable. Consider an automation that researches a topic, drafts content, validates sources and waits for approval before publishing. The model may return done after drafting. That does not mean the workflow is complete. It may still need validation, approval and a side-effecting publish operation.

If the entire process is represented only by a message transcript, the application has to reconstruct business state from prose after every restart. That creates several problems:

  • a model can describe a step as complete without satisfying its validator;
  • two workers can read the same transcript and both continue;
  • a timeout can hide whether an external write happened;
  • a human approval can arrive after the workflow has already moved;
  • deploying new code can change how old conversational history is interpreted.

A durable workflow should not ask the model to remember the rules that protect it. Those rules belong in code and data. The transcript is evidence and context. It is not the source of authority.

Keep deterministic transitions around probabilistic work

The most useful design rule I know is simple: models may propose; application code must dispose.

The model can propose a plan, select a tool, classify a document or recommend that the task is ready for approval. A deterministic transition policy checks whether that proposal is permitted from the current state and whether the required evidence exists. Only then does the application commit the next state.

For example, a model can propose ready_for_validation. The application can require a stored draft, a source list and a schema-valid result before accepting that transition. A model can recommend publication, but application code can require an approval record, an idempotency key and a passed preflight check.

This pattern keeps intelligence where it helps and determinism where failure would be expensive. It also makes the workflow explainable. Instead of “the agent decided,” an operator can see: the agent proposed a transition, validator version 3 passed, approval 812 was present and transition rule 7 committed the change.

Microsoft's catalog of AI agent orchestration patterns distinguishes sequential, concurrent, handoff, reflection and human-in-the-loop arrangements. A state machine can sit around any of them. It does not dictate how reasoning works inside a state; it defines the durable boundaries between stages.

A practical state model for an AI workflow

I prefer state names that describe operational truth rather than model mood. thinking is vague. planning says which stage owns the run. waiting is incomplete. awaiting_approval says what event can resume it.

StateMeaningTypical next states
queuedThe run exists but no worker owns execution.planning, failed
planningThe agent is producing a bounded plan.executing, failed
executingApproved steps or tool calls are running.validating, reconciling, failed
validatingDeterministic checks assess the latest result.executing, awaiting_approval, completed, failed
awaiting_approvalAutomatic execution is paused for a named decision.executing, cancelled
reconcilingThe system is determining whether an ambiguous side effect occurred.executing, validating, failed
completedAcceptance criteria and required side effects are confirmed.none
failed or cancelledExecution stopped with a recorded reason.explicit recovery only

The table is a starting point, not a universal standard. A read-only research agent may need fewer states. A financial or publishing workflow may need more explicit approval, compensation and audit stages. What matters is that each state has one meaning, each transition has preconditions and terminal states cannot silently restart themselves.

I also separate workflow state from execution ownership. A worker lease, lock or queue reservation says who may work now. The state says what the process means. Combining both into processing = true makes crash recovery and concurrency harder than necessary.

Persist evidence, not only a status string

A useful checkpoint contains enough information to resume without guessing. At minimum, I want a durable run record to hold:

  • the current state and a monotonically increasing version;
  • the objective, current step and acceptance criteria version;
  • references to model inputs, outputs and tool results;
  • the transition requested, who requested it and why;
  • validator results and the policy version that evaluated them;
  • attempt count, deadline and remaining retry budget;
  • idempotency keys and external operation identifiers;
  • approval or cancellation records;
  • created, updated and next-eligible timestamps.

LangGraph's persistence model uses checkpoints associated with a thread so graph state can be inspected and resumed. Temporal approaches the same reliability problem through durable execution and recorded workflow history. Its overview of resilient agentic AI emphasizes long-running state, human interaction and recovery after infrastructure failure.

You do not have to adopt either framework to use the underlying design principle: persist decisions at meaningful boundaries. Do not rely on one long process staying alive. Do not serialize a live SDK object and call that durable state. Store the domain facts required to reconstruct the next safe action.

An append-only transition log is valuable alongside the current snapshot. The snapshot answers “where are we?” The log answers “how did we get here?” That distinction becomes essential when debugging an agent whose final output looks plausible but whose route through tools and approvals was wrong.

Side effects need reconciliation states

The hardest state is often not failed. It is unknown.

Suppose an agent sends a publish request and the connection times out. The external system may have accepted the request even though the worker saw no response. Moving directly back to executing can publish twice. Moving directly to failed can tell an operator nothing happened when it did.

That is why reconciling deserves first-class status. The next action is a read: query by an external operation ID, inspect the target state or search for the idempotency key. Only after reconciliation should the workflow validate success, retry safely or request human help.

I described the broader contract in idempotency for AI automation. A state machine makes that contract visible. A write transition must either be idempotent, reversible through a defined compensating action, or guarded by approval and reconciliation. “Try again” is not a state strategy.

Approval deserves the same precision. Human approval should be a control boundary with a named request, scope and expiration. It should not be a chat message the agent vaguely remembers. When approval arrives, the application checks that the run is still in the expected state and that its version matches the request.

A practical PHP state machine pattern

PHP does not need a special AI runtime to enforce state transitions. An enum, a transition policy and a transactional compare-and-swap update are enough for many systems.

enum RunState: string
{
    case Queued = 'queued';
    case Planning = 'planning';
    case Executing = 'executing';
    case Validating = 'validating';
    case AwaitingApproval = 'awaiting_approval';
    case Reconciling = 'reconciling';
    case Completed = 'completed';
    case Failed = 'failed';
}

final class TransitionPolicy
{
    private const ALLOWED = [
        'queued' => ['planning', 'failed'],
        'planning' => ['executing', 'failed'],
        'executing' => ['validating', 'reconciling', 'failed'],
        'validating' => [
            'executing', 'awaiting_approval', 'completed', 'failed',
        ],
        'awaiting_approval' => ['executing', 'failed'],
        'reconciling' => ['executing', 'validating', 'failed'],
        'completed' => [],
        'failed' => [],
    ];

    public function assertAllowed(RunState $from, RunState $to): void
    {
        if (! in_array($to->value, self::ALLOWED[$from->value], true)) {
            throw new DomainException("Invalid transition: {$from->value} -> {$to->value}");
        }
    }
}

The policy only proves that an edge exists. A transition service should also check evidence. Moving from validation to completion may require every validator to pass. Moving from approval back to execution may require an unexpired approval tied to the current run version. Moving out of reconciliation may require a confirmed external result.

When committing the transition, update the row only if its version is unchanged:

$updated = AutomationRun::query()
    ->whereKey($run->id)
    ->where('version', $run->version)
    ->where('state', $run->state->value)
    ->update([
        'state' => $next->value,
        'version' => $run->version + 1,
        'transition_reason' => $reason,
        'updated_at' => now(),
    ]);

if ($updated !== 1) {
    throw new ConcurrentTransition();
}

The transition log and an outbox event should be written in the same database transaction. A queue worker can then react to the committed event. Laravel's queue documentation covers retries, timeouts, unique jobs and dispatching after database commit. Those execution controls complement the state machine; they do not replace its domain rules.

If the workflow already uses Laravel queues for AI automation, let one layer own each responsibility. The queue owns delivery and worker execution. The transition service owns legal state changes. The validator owns acceptance criteria. The model owns proposals and content, not authorization.

Test transitions, not only model outputs

Agent evaluations often sample answer quality. That matters, but workflow correctness can be tested without calling a model at all.

  • Allowed-transition tests: every documented edge succeeds with the required evidence.
  • Forbidden-transition tests: terminal states cannot restart, approval cannot be skipped and stale requests cannot commit.
  • Concurrency tests: two workers attempt the same version; exactly one wins.
  • Resume tests: kill a worker after each checkpoint and verify the next worker continues from durable state.
  • Ambiguous-write tests: simulate a timeout after an external side effect and verify the run enters reconciliation rather than repeating the write.
  • Migration tests: load old run versions after a deployment and verify they can finish or stop safely.

Then evaluate the model inside individual states: does planning produce a valid plan, does tool selection respect policy and does validation detect unacceptable output? Separating these test layers makes failures easier to diagnose. A bad draft is different from an illegal transition, even when both prevent completion.

Record transition events in the same telemetry used for production AI agent observability. State, run version, proposed action, decision reason, validator result and external operation ID are far more useful than a log line that says “agent step completed.”

When a state machine is the wrong abstraction

Not every agent needs eight states and a transition service. A short, read-only task that completes inside one request may be clearer as a bounded loop. A computation with independent branches may be more naturally represented as a graph. Folarin Adebayo's comparison of loops, state machines and graphs for agents makes the trade-off explicit: start with the simplest control structure that matches the problem.

I reach for an explicit state machine when at least one of these is true:

  • the workflow survives process restarts or lasts longer than one request;
  • multiple workers or people can act on the same run;
  • the agent performs external side effects;
  • approval, deadlines or policy gates matter;
  • operators need to resume, cancel or explain the run;
  • a wrong transition is more dangerous than a mediocre model answer.

A graph and a state machine are not mutually exclusive. A graph can describe execution dependencies inside a stage, while the state machine defines the durable business lifecycle around it. The mistake is not choosing one over the other. The mistake is letting a framework's in-memory control flow become the only record of what the system has done.

A production checklist

  • Name states after operational facts, not vague model behavior.
  • Define allowed transitions and evidence requirements in application code.
  • Persist a versioned snapshot plus an append-only transition history.
  • Let models propose actions; do not let prose authorize durable state changes.
  • Use optimistic locking or another concurrency control for every transition.
  • Give ambiguous side effects a reconciliation path.
  • Represent approval as durable, scoped data tied to a run version.
  • Test forbidden transitions, crash recovery and concurrent workers.
  • Keep terminal states terminal unless an explicit recovery policy says otherwise.

The model is often the most visible part of an AI agent, but it is not the part that makes the system dependable. Reliability comes from the surrounding contracts: durable state, bounded retries, idempotent tools, observable decisions and explicit authority. A state machine gives those contracts somewhere concrete to live.

Sources were checked on October 6, 2026. The state model and PHP examples are practical engineering patterns rather than a universal standard; appropriate states and controls depend on workflow risk. This article reflects experience with PHP backends, queues, automation and AI agents without identifying an employer or client. It was prepared with AI assistance and editorial review. Cover: control-room photograph from Pixabay.

Igor Gawrys
Igor Gawrys
AI Engineer & IT Consultant · Katowice, Poland