Reliability

Incident Review Without the Blame Theatre

You can tell within ninety seconds whether an incident review will produce anything. It comes down to the first question asked in the room.

Every organization says it runs blameless postmortems. Far fewer actually do, and the gap is not in the template. It is in what the most senior person in the room asks first.

If the first question is a variant of who deployed this, the meeting is over before it starts. Everyone present recalibrates toward self-protection, the account narrows to what is defensible, and the useful information — the near-misses, the thing that felt wrong on Tuesday, the alert everyone had learned to ignore — never surfaces.

You do not get honest incident data from people who are managing their exposure. And you cannot fix what you were never told.

What blame costs, concretely

It is not primarily a culture concern. It is an information concern.

The most valuable content in any incident is the set of small signals that preceded it. Those live in individual engineers' heads, and they are volunteered or not depending entirely on whether volunteering feels safe. A blameful review does not just fail to surface them once — it teaches everyone present not to surface them next time either.

So the second incident is diagnosed with less information than the first. That is the compounding cost.

The question that changes the room

"What made this possible?"

Not what caused it — causation invites a single culprit, human or technical. What made it possible invites the conditions: the test that did not exist, the review rushed because a date nobody challenged was three days away, the runbook last accurate eighteen months ago, the alert that had been noisy for weeks.

Every one of those is fixable. None of them is a person.

Things I insist on

Where blamelessness stops

Blameless does not mean consequence-free. If someone repeatedly bypasses controls, that is a performance conversation — held separately, with their manager, not in a review with fifteen people. Conflating the two is how organizations end up with neither honest reviews nor real accountability.

Severity is a decision, not a measurement

One practical thing that improves reviews more than the reviews themselves: agree severity definitions in advance, in terms of customer impact, and apply them consistently.

Without that, severity gets negotiated during the incident by whoever is most senior or most anxious, and your incident data becomes uncomparable across quarters. You cannot tell whether reliability improved, because the yardstick moved.

The organizations I have seen handle this well are not the ones with the most sophisticated tooling. They are the ones where an engineer will say, in a room with their director present, "I saw something odd on Tuesday and did not chase it." That sentence is worth more than any dashboard, and you only get it if you have made it free to say.

← All writing Next: Choosing React Native for a Commerce App, Honestly →