Network router representing Redis in a production AI workflow
← Back to Blog
October 2026·AI Engineering·12 min read

Where Redis belongs in a production AI workflow - and where it does not

Redis can make AI workflows fast, coordinated and observable, but only when each key has a deliberate durability, ownership and recovery contract.

Redis in an AI workflow is most useful as a fast operational layer: it can hold short-lived state, coordinate workers, enforce limits, transport queued work and cache results that are safe to recompute. It becomes dangerous when convenience quietly turns it into the only record of approvals, money, published artifacts or other facts the business cannot afford to lose.

I have seen the same architecture discussion repeat in ordinary PHP systems and newer agent platforms. A team adds Redis for caching. Then it uses Redis for queues because Horizon makes that easy. Locks, rate limits, sessions and progress updates follow. Soon someone asks whether the workflow database can disappear too.

That is the moment to stop asking, “Can Redis store this?” It probably can. The better question is, “What happens if this key expires, is evicted, is delayed during failover or must be reconstructed six months later?” The answer determines whether Redis is the right owner of the data or merely the right acceleration layer.

The decision rule: latency, loss and reconstruction

I classify each piece of workflow data along three axes before choosing a store:

  1. Latency: does the workflow need this value in milliseconds on nearly every step?
  2. Loss tolerance: can the system continue safely if the value disappears or rolls back?
  3. Reconstruction: can the value be rebuilt deterministically from a durable source?

Redis is an excellent fit when latency matters, limited loss is acceptable and reconstruction is cheap. A cached model response, current progress percentage or short lease usually fits that profile. A signed approval does not. An invoice does not. The canonical history of which external side effects occurred does not.

Workflow dataRedis roleDurable ownerRecovery rule
Response cachePrimary hot copy with TTLNone, if safely recomputableRecompute on miss
Rate-limit countersOperational sourcePolicy configuration elsewhereFail closed or degrade deliberately
Worker leaseOperational source with expiryRun record in SQLReacquire after lease expiry
Queue payloadTransport referenceWorkflow/run tablesReload by stable ID
Progress updatesFast latest viewCheckpoint/event historyRebuild from accepted steps
Approval decisionOptional cached projectionRelational audit recordNever infer approval from cache
Generated documentOptional retrieval cacheObject storage plus metadataFetch immutable artifact

This is the article's core rule: put speed-sensitive, bounded and rebuildable workflow data in Redis; keep irreplaceable business evidence in a system designed to own it.

What belongs in Redis

Hot workflow state with a durable parent

An agent may need its current step, latest progress value, active tool-call ID or a small context window on almost every transition. Reading those values from Redis can reduce repeated database work and make a live interface feel immediate.

I still create a durable run record first. The Redis key is a projection such as workflow:{runId}:runtime, not the existence proof of the workflow. It contains a schema version, run version and expiry. If the key disappears, a worker can rebuild it from the latest accepted AI agent checkpoint. If it cannot be rebuilt, it was not really a cache.

TTL is part of the data model, not cleanup added later. A workflow that can pause for human approval for seven days should not use a six-hour runtime key unless restoration is intentional and tested. Every key family needs an owner, lifetime, maximum size and missing-key behavior.

Caches for deterministic or safely reusable work

AI workflows repeat expensive reads: document parsing, embedding generation, retrieval, provider capability lookup and sometimes model responses. Redis can remove repeated latency and cost when the cache key captures every input that affects correctness.

A good key includes the tenant, normalized input hash, model or algorithm version, prompt version and relevant policy version. Caching only by the user's question is a shortcut that can leak results across customers or return an answer produced under an obsolete policy.

$cacheKey = sprintf(
    'ai:summary:%s:%s:%s:%s',
    $tenantId,
    hash('sha256', $documentContents),
    $promptVersion,
    $modelVersion,
);

$summary = Cache::store('redis')->remember(
    $cacheKey,
    now()->addHours(12),
    fn () => $summarizer->summarize($documentContents),
);

I do not cache a model output merely because it was expensive. I cache it when reuse is allowed, the output passed validation, freshness requirements are known and a miss is safe. Sensitive prompts and responses need the same tenant isolation, encryption and retention thinking as any other customer data.

Rate limits, quotas and concurrency controls

Redis counters and atomic operations fit burst control well. They can protect a model provider quota, cap concurrent jobs per tenant and prevent one batch customer from starving interactive traffic. Locks can also coordinate short critical sections such as claiming a run or refreshing one shared cache entry.

The lock does not prove that a business action completed. It only controls concurrent access for a bounded time. The database still records the accepted state transition, operation ID and version. This distinction matters when a worker holds a lease, sends an external request and dies before recording the response.

In Laravel, unique jobs and WithoutOverlapping middleware can use a lock-capable cache, including Redis. They reduce duplicate execution pressure, but they do not replace idempotency at the side-effect boundary.

Queues, Streams and Pub/Sub are different promises

“We use Redis for messaging” is not a complete architecture statement. Redis-backed queues, Redis Streams and Redis Pub/Sub have different delivery and recovery behavior.

Laravel's Redis queue is a practical choice for background model calls and tool execution. Horizon adds worker configuration, throughput metrics, wait-time monitoring and failure visibility. I like this arrangement when the application already uses Laravel, operators understand Horizon and the durable workflow record exists outside the queue payload. The broader reliability contract is the same one I use for Laravel queues in AI automation.

The job should carry a stable run ID and expected version, not the only copy of a large prompt or document. A worker loads the current state, rejects stale work and writes the next transition transactionally. If a job is delivered more than once, deterministic infrastructure makes later delivery harmless. An explicit state machine decides whether that command is still legal.

Redis Streams are useful when the application needs an append-like sequence, consumer groups, pending-entry inspection and the ability to claim abandoned work. They can support coordination or an operational event feed. Their safety still depends on Redis persistence, replication, trimming and acknowledgement choices. “It is a Stream” does not automatically mean “it is our permanent audit log.”

Pub/Sub is for transient fan-out. Redis documents it as at-most-once delivery: a disconnected subscriber misses the message. That can be correct for “refresh the progress widget now.” It is wrong for “charge this customer,” “publish this article” or “record this approval.” If missing one message changes the business outcome, use a durable command or event mechanism.

Agent memory and vector retrieval need explicit boundaries

Redis now supports vector search and AI-oriented retrieval patterns, so it can serve low-latency semantic search, session memory and a retrieval layer close to other runtime data. That does not eliminate the need to define what “memory” means.

I separate at least four things that teams often merge under one label:

  • Conversation buffer: recent turns needed for the next response, normally bounded by time, tokens or count.
  • Working memory: accepted facts and intermediate artifacts for the current run.
  • Retrieval index: chunks and embeddings used to find relevant source material.
  • Business record: authoritative customer, policy, approval and outcome data.

Redis may own the bounded conversation buffer and retrieval index. Working memory may be a Redis projection over durable checkpoints. The business record usually belongs in a relational database or another governed system of record.

Embeddings are derived data. Store the source document ID, chunk version, embedding model and content hash with every vector. Then an index can be rebuilt after a model change or corruption. Without lineage, fast retrieval can return a fragment whose source and freshness nobody can explain.

What should not live only in Redis

The only copy of workflow truth

Redis offers RDB snapshots, append-only files and combinations of both. Those are real persistence options, not an excuse to call Redis “just a cache.” They also involve explicit durability trade-offs. Redis documents that periodic RDB snapshots can lose the latest minutes after a failure, while the common AOF policy that syncs every second may lose about one second of writes.

Those guarantees may be perfectly acceptable for some systems. The mistake is accepting them accidentally. If the workflow state determines whether an irreversible action may run again, choose persistence, replication, backups and recovery objectives deliberately. In many applications, keeping the canonical state machine in PostgreSQL or MySQL makes transactions, constraints and audit queries easier to reason about.

Approvals, audit history and authorization evidence

A production agent should be able to prove who approved which artifact, under which policy, at what time and for which action. That record should survive cache eviction and ordinary operational mistakes. Store it durably with an immutable artifact hash. Redis may accelerate the “is this run currently waiting?” view, but the workflow must never treat a missing or stale cache value as permission.

Large artifacts and unbounded transcripts

Documents, images, audio, full provider responses and long transcripts consume memory quickly. Put immutable blobs in object storage and metadata in the database. Redis should hold references, compact projections or deliberately bounded fragments.

An unbounded conversation key is not long-term memory. It is an eviction incident waiting for traffic. Summarize, checkpoint, archive and expire according to product needs rather than letting every token stay hot forever.

Irreplaceable events delivered through Pub/Sub

Pub/Sub is intentionally ephemeral. It is excellent for UI invalidation, live telemetry and best-effort notifications. It is not a durable handoff between “the customer approved” and “the agent may publish.” That transition needs a committed record and a recoverable command.

A practical Laravel architecture for Redis AI workflows

The architecture I reach for is deliberately split by responsibility:

  1. PostgreSQL or MySQL owns the workflow. Run state, accepted transitions, approvals, budgets, idempotency records and external operation IDs live here.
  2. Object storage owns artifacts. Source files, generated documents and large tool results are immutable objects addressed by ID and hash.
  3. Redis accelerates execution. Queues, Horizon, locks, rate limits, hot state, small caches and live progress use named key families with TTLs.
  4. A vector index supports retrieval. Redis can fill this role when its latency, filtering and operating model fit. Every vector retains source lineage.
  5. An outbox bridges commits to queues. The workflow transition and “dispatch next command” record commit in one database transaction. A publisher delivers the command afterward.

This creates useful failure behavior. If Redis is flushed, accepted workflow state remains. Queues can be reconstructed from the outbox and stalled runs. Hot state can be rebuilt from checkpoints. Vector data can be re-indexed from source objects. The incident is painful, but it does not force the application to guess whether a customer already approved or an external write already happened.

If the relational database is temporarily unavailable, I normally stop state-changing workflow transitions rather than continuing from Redis alone. The system may still serve cached reads or progress views, but it should not create an unrecorded branch of reality.

Production safeguards that make Redis boring

  • Name key families. Document their owner, value schema, TTL, cardinality, maximum size and deletion behavior.
  • Separate workloads. Do not let an unbounded semantic cache evict queue, lock or rate-limit data from the same instance. Use distinct instances or carefully isolated policies where consequences differ.
  • Set memory and eviction policies intentionally. Alert before saturation, not after keys disappear.
  • Measure queue health. Track wait time, oldest job, retries, failed jobs and per-tenant backlog rather than only Redis CPU. Feed those events into the same model used for AI agent observability.
  • Test restarts and failovers. Kill workers after model calls, before acknowledgements and around external writes.
  • Monitor cache quality. Hit rate alone is insufficient; watch stale reuse, cross-version collisions and recomputation cost.
  • Protect tenant boundaries. Include tenant identity in keys and vector filters, restrict access and never store credentials in workflow values.
  • Practice reconstruction. A rebuildable cache is only rebuildable after the recovery process has been tested.

Redis earns its place in a production AI workflow when it makes the system faster without becoming a hidden source of authority. It is excellent at holding what is hot, coordinating what is concurrent and discarding what is safely temporary. It should not be asked to erase the distinction between operational state and business truth.

The architecture test is simple: imagine Redis is empty after a bad morning. Can the team determine what happened, restore pending work and avoid repeating irreversible actions? If yes, Redis is probably serving the workflow. If no, the workflow is serving Redis.

Sources were checked on October 9, 2026: Redis documentation on persistence, Streams, Pub/Sub delivery semantics and AI and vector search; Laravel 13.x documentation on queues and Horizon. The storage matrix and reference architecture are practical engineering guidance rather than a universal standard. This article reflects experience with PHP backends, Redis, queues, automation and AI agents without identifying an employer or client. It was prepared with AI assistance and editorial review. Cover: network router photograph from Pixabay.

Igor Gawrys
Igor Gawrys
AI Engineer & IT Consultant · Katowice, Poland