THIRTY YEARS OF PATTERNS

PATTERN 02

Your last critical incident is trying to teach you someting.

Your last critical incident is trying to teach you someting.

Your last critical incident is trying to teach you someting.

After participating in hundreds of critical incidents over my career, I stopped seeing outages as isolated events. I began seeing patterns. Over time, I realized the organizations that became more reliable weren't the ones with fewer incidents—they were the ones that learned the most from them.

After participating in hundreds of critical incidents over my career, I stopped seeing outages as isolated events. I began seeing patterns. Over time, I realized the organizations that became more reliable weren't the ones with fewer incidents—they were the ones that learned the most from them.

4 MIN READ

JULY 2026

The Pattern

The Pattern

Like many, I originally believed the purpose of a critical incident was straightforward: restore service as quickly as possible. When customers can't use your product, nothing else matters.

But after participating in hundreds of incidents over the years, something changed. I stopped seeing individual outages and started seeing patterns.

Different organizations. Different technologies. Different architectures. Yet many of the same themes surfaced again and again. Sometimes it was a connection pool exhausted under load because retry behavior hadn't been fully considered. Other times, teams hesitated to enable automated scaling or failover because they weren't yet confident in how those capabilities would behave under real-world conditions. In nearly every case, the technology was only part of the story. What interested me wasn't the specific failure—it was the engineering decisions, operational tradeoffs, and organizational habits that revealed it.

The technologies changed. The patterns often didn't.

At first, I thought those patterns were telling me something about the applications. Over time, I realized the applications were only one part of the story. Every incident reflected a series of engineering and operational decisions that had accumulated over months or years. Some decisions had stood the test of time. Others had quietly become constraints that no one had revisited.

The outage wasn't the beginning of the story. It was simply the moment the story became visible.

Like many, I originally believed the purpose of a critical incident was straightforward: restore service as quickly as possible. When customers can't use your product, nothing else matters.

But after participating in hundreds of incidents over the years, something changed. I stopped seeing individual outages and started seeing patterns.

Different organizations. Different technologies. Different architectures. Yet many of the same themes surfaced again and again. Sometimes it was a connection pool exhausted under load because retry behavior hadn't been fully considered. Other times, teams hesitated to enable automated scaling or failover because they weren't yet confident in how those capabilities would behave under real-world conditions. In nearly every case, the technology was only part of the story. What interested me wasn't the specific failure—it was the engineering decisions, operational tradeoffs, and organizational habits that revealed it.

The technologies changed. The patterns often didn't.

At first, I thought those patterns were telling me something about the applications. Over time, I realized the applications were only one part of the story. Every incident reflected a series of engineering and operational decisions that had accumulated over months or years. Some decisions had stood the test of time. Others had quietly become constraints that no one had revisited.

The outage wasn't the beginning of the story. It was simply the moment the story became visible.

Every incident reflected a series of engineering and operational decisions that had accumulated over months or years

Every incident reflected a series of engineering and operational decisions that had accumulated over months or years

Impact

Impact

The organizations that impressed me most weren't the ones with the fewest incidents.

They were the ones whose incidents kept changing.

They rarely experienced the same major failure twice because every critical incident became an opportunity to revisit assumptions, refine engineering practices, and strengthen operational discipline. They understood that reliability isn't achieved through perfect technology. It's built through continuous learning.

That changed the way I approached incident reviews. I became less interested in assigning ownership to a single failure and more interested in understanding the chain of decisions that made that failure possible. Sometimes the answer lived in architecture. Other times it was found in engineering practices, operational processes, testing strategies, or simply assumptions that had never been challenged as the organization evolved.

Over time, I stopped asking, "How do we prevent this incident from happening again?"

I started asking, "What does this incident reveal about the way we build and operate software?"

That question almost always led to a better conversation.

The organizations that impressed me most weren't the ones with the fewest incidents.

They were the ones whose incidents kept changing.

They rarely experienced the same major failure twice because every critical incident became an opportunity to revisit assumptions, refine engineering practices, and strengthen operational discipline. They understood that reliability isn't achieved through perfect technology. It's built through continuous learning.

That changed the way I approached incident reviews. I became less interested in assigning ownership to a single failure and more interested in understanding the chain of decisions that made that failure possible. Sometimes the answer lived in architecture. Other times it was found in engineering practices, operational processes, testing strategies, or simply assumptions that had never been challenged as the organization evolved.

Over time, I stopped asking, "How do we prevent this incident from happening again?"

I started asking, "What does this incident reveal about the way we build and operate software?"

That question almost always led to a better conversation.

Supporting Framework

Supporting Framework

Re-evaluate Before Recreate framework

The Signal

The Signal

Every critical incident tells two stories. One explains why the application failed. The other reveals how the organization built, operated, and evolved it. That's usually the more valuable lesson.

If this pattern reflects challenges you’re seeing in your organization, I’d enjoy continuing the conversation.

If this pattern reflects challenges you’re seeing in your organization, I’d enjoy continuing the conversation.

If this pattern reflects challenges you’re seeing in your organization, I’d enjoy continuing the conversation.

Start the Conversation →

High Signal Advisory

© 2026 High Signal Advisory, LLC

High Signal Advisory

© 2026 High Signal Advisory, LLC