
Checkpoints are the difference between restarting an AI agent and resuming it
AI agent checkpoints preserve verified progress so a failed workflow can resume safely instead of repeating model calls and side effects.
AI agent checkpoints are durable records of verified workflow progress. They let a new worker reconstruct what has already happened, what remains pending and which action is safe to perform next. Without them, recovery usually means restarting the agent from its original prompt and hoping repeated reasoning, tool calls and side effects produce the same result.
Those two operations sound similar in a status update. They are completely different in production. Restarting says, “run the task again.” Resuming says, “continue from a known boundary without pretending the earlier work never happened.”
I learned to care about that distinction through ordinary backend engineering long before every workflow was called an agent. Queue workers crash. Deployments interrupt processes. APIs time out after accepting a request. A person approves something hours after the original worker disappeared. The model adds uncertainty, cost and non-determinism, but the recovery problem is familiar: durable systems need a trustworthy place from which to continue.
Restarting repeats intent; resuming preserves progress
Imagine a research agent with six stages: collect sources, extract claims, draft an answer, validate citations, request approval and publish. It reaches approval after spending twelve minutes and making several paid model and search calls. The process then restarts during a deployment.
A restart begins with the original objective. The new run may find different sources, produce a different draft and request approval for a different artifact. If it cannot tell whether publication already happened, it may repeat the final write as well. Even when nothing dangerous occurs, the user waits again and the system pays twice.
A resume loads the last accepted checkpoint. It knows which source set was validated, which draft the approval refers to, which tools completed, which calls remain pending and whether a publish operation has an external identifier. It continues from the next eligible transition rather than regenerating the past.
This is why a transcript is not enough. A transcript can show that the model said, “The draft is ready.” It does not prove that the draft was stored, the citations passed validation or an approval request was created. Conversational history is useful context. A checkpoint is application-owned evidence.
The same distinction appears in established workflow systems. Microsoft Agent Framework documents checkpoints as saved workflow state from which execution can later resume. LangGraph associates graph-state checkpoints with a thread for interruption recovery, human-in-the-loop operation and fault tolerance. Temporal takes an event-history and deterministic replay approach. Their implementations differ, but they share a principle: recovery uses durable execution facts, not a fresh attempt to infer the past.
Create checkpoints at semantic boundaries
A checkpoint is valuable only when its meaning is clear. Saving every few seconds can produce many snapshots without telling the next worker whether a business action completed. Saving only at the end leaves too much work exposed. I prefer semantic boundaries: moments after the application can prove one unit of work is complete and before the next unit may change the outside world.
Useful boundaries include:
- after a plan passes schema and policy validation;
- after source documents are stored and their identifiers are fixed;
- after a model output passes deterministic acceptance checks;
- before requesting a human decision, with the exact artifact under review;
- immediately before an external side effect, with an idempotency key reserved;
- after that side effect is confirmed and its provider identifier is stored;
- when entering a retry delay, approval wait or reconciliation state.
These boundaries align naturally with an explicit AI agent state machine. The state machine defines legal transitions. The checkpoint stores the facts needed to resume one. Neither replaces the other. A state name without durable evidence is too thin; a pile of snapshots without transition rules is ambiguous.
I do not checkpoint every token streamed by a model. Partial text is often cheap to regenerate and difficult to validate. I checkpoint the accepted output of the step. For a long-running tool, I may save a provider job ID before polling so another worker can continue observing the same job instead of submitting a second one.
The right frequency follows the cost and irreversibility of lost progress. A read-only classification batch might checkpoint every 100 items. A payment, publication or outbound message needs a boundary around each individual write. “Every N seconds” is an implementation setting, not a recovery design.
What an AI agent checkpoint should contain
I treat a checkpoint as a versioned envelope, not a serialized framework object. Framework snapshots can be useful, but the application still needs a stable domain record it can inspect, migrate and explain.
| Field group | What to store | Why it matters |
|---|---|---|
| Identity | run ID, workflow type, checkpoint ID, parent checkpoint ID | Prevents state from one run being restored into another |
| Position | state, step, transition version, pending action | Defines where execution may continue |
| Inputs | objective reference, normalized inputs, prompt and policy versions | Makes the resumed step reproducible and auditable |
| Evidence | artifact IDs, hashes, validator results, tool-result references | Proves completed work instead of trusting prose |
| Side effects | idempotency keys, provider operation IDs, confirmation status | Prevents duplicate writes and enables reconciliation |
| Budgets | attempts, tokens, money, elapsed time and deadline remaining | Stops a restored run from receiving a fresh unlimited budget |
| Control | approval request ID, actor, lease owner, next eligible time | Preserves authority and scheduling boundaries |
| Compatibility | schema version, workflow code version, serializer version | Allows safe restoration after deployments |
Large prompts, documents and tool responses do not need to live inside the checkpoint row. Store immutable objects separately and reference them by ID and content hash. The envelope stays small while still proving exactly which input and output a transition used.
Hashing matters when an approval refers to an artifact. If a user approved draft hash abc123, a resumed agent must not publish a later draft under the old approval. The approval record should bind the run version, artifact hash, requested action, approver and expiration time.
Budgets must also survive restoration. If each worker reload gives an agent three more retries, then a crash loop silently turns a bounded workflow into an unlimited one. Preserve the retry budget, token spend and deadline as workflow state.
Checkpoint around side effects, not through them
The most dangerous recovery window is the gap between an external system accepting a write and the agent recording success. A worker sends a publish request. The provider creates the post. The response is lost. The worker dies before updating its database.
No checkpoint technology can make two independent systems commit atomically unless they participate in the same transaction, which external APIs usually do not. The practical answer is a small protocol:
- persist the intended action and reserve an idempotency key;
- commit a checkpoint that says the write is pending;
- call the external system with that key or a stable client reference;
- store the provider operation ID and confirmed result;
- commit the post-action checkpoint.
If the worker disappears between steps three and four, the next worker does not publish again immediately. It enters reconciliation and asks the provider whether the operation exists. This is the operational difference between resuming and repeating.
Some APIs support idempotency keys directly. Others let you query by a client-generated identifier. When neither is available, use the strongest observable fingerprint the target exposes and place human review at the ambiguity boundary. The approach follows the same contract described in idempotency for AI automation: duplicate delivery is normal, duplicate effect is a design failure.
A checkpoint cannot turn an irreversible action into a reversible one. It can record enough evidence to avoid blind repetition and route the run to a partial-failure or compensation path. That is already a major improvement over “the agent crashed, start it again.”
A practical checkpoint model in Laravel
For many PHP applications, a relational database, queue and object storage are enough. I would start with an automation_runs table for the current snapshot and an append-only automation_checkpoints table for history. Each checkpoint gets a monotonically increasing sequence and an immutable JSON payload.
final readonly class CheckpointEnvelope
{
public function __construct(
public string $runId,
public int $sequence,
public string $state,
public string $step,
public array $artifactRefs,
public array $validatorResults,
public array $externalOperations,
public array $remainingBudget,
public int $schemaVersion,
public string $workflowVersion,
) {}
}
The service that accepts a step result should write the checkpoint, update the current run version and insert an outbox message in one database transaction:
DB::transaction(function () use ($run, $envelope, $nextJob): void {
$updated = AutomationRun::query()
->whereKey($run->id)
->where('version', $run->version)
->update([
'state' => $envelope->state,
'step' => $envelope->step,
'version' => $run->version + 1,
'checkpoint_sequence' => $envelope->sequence,
]);
if ($updated !== 1) {
throw new ConcurrentCheckpointWrite();
}
AutomationCheckpoint::create([
'run_id' => $run->id,
'sequence' => $envelope->sequence,
'schema_version' => $envelope->schemaVersion,
'payload' => $envelope,
]);
OutboxMessage::create([
'topic' => 'automation.resume',
'payload' => $nextJob,
]);
});
The optimistic version check prevents two workers from creating competing “next” checkpoints from the same parent. The outbox ensures the queue message is derived from committed state. A dispatcher publishes it after the transaction. If delivery occurs twice, the new worker compares the expected checkpoint sequence before doing work.
This fits well with Laravel queues for AI automation. The queue transports commands and applies retry policy. The checkpoint store owns durable progress. The state machine owns legal transitions. The model produces bounded outputs. Keeping those responsibilities separate makes recovery testable without asking an LLM to remember infrastructure rules.
Restoration should be a validated procedure
Loading the newest JSON document and calling the next tool is not enough. A restore path should reject state it cannot understand or trust.
- Acquire ownership. Obtain a lease or versioned lock for the run so only one worker restores it.
- Load the latest accepted checkpoint. Ignore uncommitted temporary files and incomplete writes.
- Validate identity and integrity. Check the run ID, parent sequence, schema version and hashes of referenced artifacts.
- Apply migrations. Convert old checkpoint schemas through explicit, tested migrations. Never reinterpret them implicitly.
- Reconcile pending effects. Resolve any external operation whose outcome is unknown before scheduling another write.
- Re-evaluate time-sensitive policy. Confirm approvals have not expired, credentials remain valid and deadlines have not passed.
- Dispatch only the next eligible command. Include the expected run version so stale jobs become harmless no-ops.
Code versioning deserves special attention. A checkpoint created before a deployment may describe a step that no longer exists. Sometimes a migration can translate it. Sometimes the only safe choice is to pause the run for operator review. Silent fallback to the beginning is not compatibility; it is data loss disguised as recovery.
Test this path by terminating workers after every meaningful boundary. After a pre-write checkpoint, restoration should reconcile or perform one write. After a confirmed-write checkpoint, it should never repeat that write. After an approval checkpoint, it should continue only with approval tied to the current artifact. These tests reveal more about production reliability than another collection of ideal model responses.
Checkpointing creates operational responsibilities
Durability is not free. Checkpoints may contain customer data, model inputs, tool output and operational identifiers. Encrypt sensitive fields, apply tenant boundaries, restrict deserialization types and audit reads. Do not place credentials or access tokens in the envelope. Store references to secrets, then resolve them through the current credential system at execution time.
Retention also needs a policy. LangGraph's documentation explicitly warns that checkpoints can grow without bound. Keep enough history for recovery, audit and incident analysis, then compact or delete according to product and legal requirements. Terminal runs may retain the final checkpoint plus important transition evidence while expiring intermediate model context sooner.
Monitor checkpoint age, restore count and stalled pending operations. A run that has restored seven times is not healthy simply because it eventually completes. It may indicate a poison input, incompatible deployment or tool that repeatedly times out after success. Send those signals into the same event stream used for AI agent observability.
Finally, provide operators with deliberate controls: resume from latest, inspect evidence, cancel, retry a named step, or fork from an older checkpoint into a new run. Do not let “resume” secretly rewrite history. If an operator chooses an earlier point, create a new branch with a clear parent so the audit trail remains honest.
A production checklist for AI agent checkpoints
- Checkpoint accepted domain progress, not arbitrary wall-clock intervals.
- Store immutable artifact references and hashes instead of trusting transcript claims.
- Preserve retry, token, cost and time budgets across worker restarts.
- Bind approvals to a run version, action and artifact hash.
- Record an idempotency key before every external side effect.
- Reconcile unknown outcomes before repeating a write.
- Use optimistic concurrency so one parent checkpoint has one accepted successor.
- Version checkpoint schemas and test migrations across deployments.
- Exercise crash recovery after every important boundary.
- Encrypt, retain and delete checkpoint data according to its real sensitivity.
An AI agent does not become resilient because its process manager can launch it again. That only makes repetition automatic. Resilience begins when the application can prove what is complete, preserve the remaining authority and budget, and choose one safe next action after the original worker is gone.
That is the purpose of a checkpoint. It is not merely a saved conversation or a convenient framework feature. It is the durable contract that turns “try the agent again” into “continue the workflow from here.”
Sources were checked on October 6, 2026: Microsoft Agent Framework workflow checkpoints, LangGraph persistence documentation and Temporal's workflow execution overview. The checkpoint envelope and Laravel implementation are practical engineering patterns rather than a universal standard. This article reflects experience with PHP backends, queues, automation and AI agents without identifying an employer or client. It was prepared with AI assistance and editorial review. Cover: signpost photograph from Pixabay.
