
What production incidents teach that architecture diagrams cannot
Production incidents reveal temporal coupling, changing data shapes, uncertain state and recovery paths that static architecture diagrams cannot show.
Production incident lessons rarely fit inside the architecture diagram that existed before the outage. The diagram can correctly show the application, database, workers and external services while missing the conditions that made them fail together: a weekly trigger, a growing payload, a retry without idempotency, an undocumented operator decision or a recovery path nobody had tested.
I learned this during an incident that looked impossible on paper. A Laravel application at a previous employer crashed every Tuesday at 2 AM. The application diagram was not wrong. It showed the web tier, workers, database and scheduled jobs. Yet none of those boxes explained why Tuesday was different from Monday.
The answer existed across time and teams. A Monday marketing campaign changed the volume and shape of order metadata. A Tuesday export selected more data than accounting needed. The export built one large CSV string in memory. The operating system killed the workers when memory spiked.
An architecture diagram describes the intended structure of a system. A production incident reveals its conditional behavior: what happens when timing, data, capacity, ownership and recovery collide.
The diagram was accurate and still incomplete
It is easy to criticize an architecture diagram after an outage. The database icon did not show replication lag. The arrow to a provider did not mention a timeout. The worker box did not list its memory limit. Therefore, the diagram was “outdated.”
That conclusion is usually too simple. A useful diagram is a deliberate abstraction. If it included every timeout, payload field, cron expression, queue limit, deployment version and recovery command, nobody could read it. The problem is not that diagrams omit details. The problem is that teams sometimes treat one structural view as a complete operational model.
Before the Tuesday incident, the important design questions were visible: which service owned orders, where exports ran and how workers reached the database. During the incident, different questions mattered. What changed before the spike? Which process allocated memory? What was unique about this batch? Could we stop the export without affecting checkout? Which evidence would survive a restart?
The diagram answered “what connects to what.” Production forced us to ask “under which conditions does this connection become dangerous?” That is the gap incidents expose.
1. Dependencies have a clock
Most architecture diagrams represent spatial dependency. Service A calls service B. A worker reads queue C. An application writes database D. They are snapshots of components and connections.
Incidents reveal temporal dependency. A campaign runs Monday evening. Orders accumulate with unusually large metadata. An export starts Tuesday morning. A maintenance job overlaps with it. A certificate expires on the first day of the month. A cache warms only after deployment. A retry storm begins thirty seconds after the provider slows down.
These relationships are real architecture even though no direct arrow connects them.
In the Tuesday 2 AM production incident, the marketing campaign did not call the export job. The two systems were coupled through the data accumulated between them. The day and hour sounded irrational only because our mental model contained services but not the calendar.
After that experience, I became more suspicious of phrases such as “nothing changed.” Code may not have changed. Load distribution, input age, batch size, customer behavior, provider quotas and overlapping schedules may have changed. Time is an input even when it does not appear in a function signature.
A diagram used for operations needs a temporal companion: scheduled triggers, expected peaks, expiry boundaries, retention windows and retry intervals. It does not have to show every cron expression. It should show the few clocks capable of changing system behavior.
2. Data shape is part of the architecture
The original export had worked for years. That history created confidence, but it proved only that the implementation survived the data it had already seen.
The decisive change was not a new component. Order metadata became much larger after another feature began storing recommendation data. The export used SELECT *, so a field irrelevant to accounting crossed a boundary by accident. Repeated string concatenation then amplified the memory cost.
A box labeled “database” hides this entire class of risk. So does an arrow labeled “orders.” At runtime, the difference between 200 bytes and 50 KB per record can become more important than the choice between two database engines. Cardinality, payload size, skew and growth rate are architectural properties because they determine whether a design remains viable.
This applies beyond batch exports. A queue consumer may be fast enough for average messages but collapse on one tenant's large payloads. An AI workflow may handle short documents but exceed context or cost limits on a contract archive. A cache may be effective until one high-cardinality key space destroys its hit rate.
I do not put full schemas on high-level diagrams. I do annotate critical flows with operational facts: normal and worst credible volume, maximum payload, fan-out, retention, ordering requirements and the system that owns validation. These notes explain the weight carried by an arrow.
3. Component health is not workflow health
During a distributed incident, every dashboard can be green while the user-facing outcome is broken.
The API accepted the request. The database committed. The queue received a message. The worker returned success. The notification provider accepted an email. Each component reports a locally reasonable result, but the customer never sees the finished outcome because the record, event and message refer to different versions of reality.
This is why partial failure is the default in multi-step automation. A green HTTP response is stage evidence, not proof of a business outcome. A diagram can show the stages, but it cannot decide what “complete” means.
Incidents force that definition. Is an order complete when it is stored, when payment is confirmed, when inventory is reserved or when the customer can see it? Is a generated report complete when the model responds, when validation passes or when the intended person can open it?
The answer belongs in the domain model and the observability model. I want a durable operation identifier, explicit states and invariants that connect technical progress to user-visible completion. Component metrics remain necessary. They become useful when they can be joined into the story of one outcome.
4. Recovery paths are architecture
Design reviews spend most of their time on the forward path. A request enters, services collaborate, data is stored and a response leaves. During an incident, the reverse questions dominate.
- Can this deployment be rolled back without rolling back the data?
- Can a worker resume from a checkpoint, or will it repeat external side effects?
- Can one subsystem be disabled while the core service remains available?
- Can traffic move to another region without creating two writers?
- Can an operator prove whether an uncertain request reached the provider?
If the team cannot answer those questions, the recovery path is not merely missing documentation. It may be missing from the architecture.
This is where rollback, idempotency, graceful degradation, compensating actions and manual repair tools matter. They are often drawn as notes after the “real” design. Production shows that they are part of the real design.
I now ask for a recovery path beside every important success path. The answer can be simple. Stop intake, drain the queue, restore a snapshot or switch a feature flag. What matters is that the path has been chosen, its assumptions are visible and somebody can execute it safely.
The same principle guides how I design systems for reversible failures: the system does not need to prevent every error, but it must limit uncertainty and preserve a safe next action.
5. Ownership appears under pressure
An architecture diagram usually assigns technical ownership implicitly. This team owns the service. That team owns the database. A vendor owns the external API.
An incident reveals operational ownership. Who is allowed to stop the batch? Who can decide that degraded service is acceptable? Who contacts the provider? Who tells support what customers should expect? Who verifies recovery after the graphs return to normal?
These are not management details outside the system. Slow or ambiguous decisions extend outages. A technically correct failover is useless when nobody knows who may trigger it. A complete rollback script is not a recovery plan when only one unavailable engineer knows the credentials.
Microsoft's Azure Well-Architected guidance separates the architectural and procedural sides of incident response, but treats both as necessary. It recommends defined roles, escalation structures and procedures for detection, triage, containment and recovery. That matches my experience: the system's runtime behavior includes the humans making decisions around it.
After an incident, I want ownership attached to the operational boundary, not only the repository. The person who maintains code may not own the business decision to pause a workflow or replay external actions.
Read the evidence the diagram cannot hold
The architecture diagram is a hypothesis about how the system is organized. Runtime evidence shows how that hypothesis behaves under real conditions.
| The diagram shows | The incident reveals | Evidence to preserve |
|---|---|---|
| A service dependency | Latency, timeout and retry interaction | Traces, attempt count, timeout reason |
| A database read | Payload size, skew and consistency needs | Row count, bytes, query plan, replica lag |
| A queue | Backpressure and poison-message behavior | Queue age, depth, redelivery and dead-letter reason |
| A scheduled job | Overlap with traffic and upstream business events | Trigger time, input window and resource peak |
| A failover path | Whether people can actually execute it | Decision owner, runbook result and recovery time |
Google's SRE troubleshooting guidance describes diagnosis as an iterative comparison between observations and theories about system behavior. Metrics, logs, traces and exposed state help confirm or reject hypotheses. The diagram contributes theory. Telemetry contributes observation. Neither is sufficient alone.
This is also why I prefer events over dashboards as the foundation of observability. A dashboard summarizes what someone expected to ask. Preserved events allow the next investigator to ask a question nobody predicted.
Add five operational overlays after every incident
I do not respond to an outage by turning one clean diagram into an unreadable forensic map. I keep the structural view and add small operational overlays for the path that failed.
- Trigger. Record what activated the failure condition. This may be a deployment, schedule, traffic pattern, payload class, expiry or operator action. In the Tuesday incident, the trigger chain began with a Monday campaign and ended with a scheduled export.
- Failure boundary. Mark where the system lost a guarantee. A timeout may create uncertainty about an external side effect. A database commit followed by a failed publish creates split state. An in-memory batch crosses a resource boundary. Name the guarantee, not only the component.
- Durable state. Show what survives a crash and what exists only in memory. Include checkpoints, operation records, outbox entries and external identifiers. This tells responders where recovery can safely begin.
- Evidence. Link the signals that prove progress or failure: SLI, log event, trace span, queue age, audit entry or output checksum. If a critical arrow has no observable evidence, the diagram has identified an instrumentation gap.
- Recovery. Add the available containment and restoration actions, their owner and their preconditions. A recovery arrow should say whether it retries, reverses, bypasses, degrades or requires manual repair.
These overlays can live in the postmortem, runbook or a second diagram. They do not need to become permanent clutter on the context view. Their purpose is to connect the stable system model with operational reality.
I also attach an expiry condition to assumptions. A maximum batch size without a monitoring query will become fiction. A recovery path that has not been exercised since three major deployments ago is a hope. Operational annotations should point to measurable evidence or a test cadence.
This method changes architecture review from “do these boxes look right?” to “which conditions can invalidate the promises between them?” That is a more useful conversation before the next incident.
A postmortem should change the model
A postmortem is not complete because the document contains a root cause and action items. It is complete when the organization has improved the model it uses to operate the system.
Google's SRE guidance defines a postmortem as a record of impact, mitigation, contributing causes and preventive actions. It emphasizes blameless learning because engineers make decisions using the information and tools available at the time. That principle matters for architecture too. “The engineer should have known” usually means the necessary fact was absent from the shared model.
Useful action items change at least one of four things:
- System behavior: reduce the likelihood or blast radius of recurrence.
- Detection: expose the condition earlier through better evidence.
- Recovery: make containment or restoration faster and safer.
- Understanding: update the diagram, runbook, ownership or limits that shaped the response.
In the Tuesday incident, changing the query and streaming the export fixed behavior. Monitoring execution time, memory peak and output size improved detection. Documenting the cross-team trigger improved understanding. Each action addressed a different part of the failure.
A diagram update alone would not have prevented recurrence. A code fix alone would have left the same blind spot for other jobs. The value came from changing both the system and the model.
What architecture diagrams still do well
None of this makes architecture diagrams obsolete. I still use them to establish scope, communicate boundaries, review dependencies, onboard engineers and discuss proposed changes before code exists.
A good diagram compresses complexity. It lets a team notice that one service has too many synchronous dependencies, that sensitive data crosses a trust boundary or that a supposedly isolated workflow shares a database with the critical path.
The mistake is asking the diagram to prove reliability. A map can show the roads without proving that a particular vehicle will survive winter traffic. Reliability depends on runtime conditions, operating practices and tested recovery.
Treat the diagram as an index into deeper evidence. A critical component should lead to its owner. An important flow should lead to its SLI and runbook. A state boundary should lead to its consistency and recovery rules. The diagram stays readable while the operational model remains accessible.
Questions for the next architecture review
When reviewing an existing system, I now ask questions that are difficult to answer from boxes and arrows alone:
- Which scheduled or business events can change the load profile?
- What payload or cardinality assumption would break this path?
- Where can one step succeed while the next step fails?
- What state survives if the process dies at each boundary?
- Which side effects are safe to retry, and how is sameness identified?
- What user-visible outcome defines completion?
- Which signal proves that outcome rather than one component's health?
- How can the team contain failure without taking down healthy capabilities?
- Who may make the recovery decision, and can they access the tools?
- When was the recovery path last tested?
If the team cannot answer all ten, that does not automatically mean the architecture is bad. It identifies uncertainty. The next step is to decide which uncertainty carries enough risk to deserve instrumentation, testing or redesign.
Production is the second architecture review
A design review asks whether the system should work. Production eventually asks under which conditions it stops working, how anyone will know and what happens next.
The incident is not wiser than the architect. It simply has access to evidence the original design did not: real timing, real data, real dependency behavior and real human decisions under pressure.
Use that evidence. Keep the clean diagram, but connect it to the clocks, limits, states, signals and recovery paths that make the system operable. Then the architecture becomes more than a picture of components. It becomes a shared model of responsibility.
Architecture diagrams show where the system is. Production incidents show where its promises end.
Sources and further reading
- Google SRE Book: Effective Troubleshooting
- Google SRE Book: Postmortem Culture - Learning from Failure
- Microsoft Azure Well-Architected: Incident response architecture strategies
Sources were checked on September 29, 2026. The production example is an anonymized incident previously documented on this site; no employer or customer is identified. This article was prepared with AI assistance and editorial review, and its framework reflects practical engineering judgment rather than a universal architecture standard. Cover: network router photograph from Pixabay.
