Sooner or later something will go badly wrong. A server dies, the product breaks in production, customer data goes missing, or a release takes the site down.
The technical fix is only half the job. The other half is communication, and that is usually the half that damages the relationship.
You know this rule applies when someone reaches out about a critical issue. It might be the CEO, a client, or a Product Owner. Something they rely on is broken, and they are looking to you.
They may sound stressed, frustrated, or upset. That is a fair reaction when something important is broken, so take it as a signal of how much this matters to them.
The severity is set by the impact on them. It is not set by how hard the fix looks to you, so don't spend time deciding whether it really counts. If it is urgent to them, treat it as urgent.
Most incidents go pear-shaped for one reason. It is rarely that the fix took too long. It is that the person who raised it was left wondering whether anyone cared.
This rule covers the human side of an incident: what you say and when. For the technical steps of an infrastructure outage, see Outage - Do you have an unplanned outage process?
If your team has an on-call roster or agreed response times, follow those. Otherwise, the clock starts when you see it. Once the message is in front of you, act on it straight away rather than reading it and coming back to it later.
π Call, don't chat
Chat is the wrong tool for an incident. It is slow, it is easy to misread, and a long back and forth reads as a lack of urgency.
Send one quick line so they know the message has landed, then call.
"Seen it, calling you now."
β Figure: Good example - Takes 2 seconds, and it stops them chasing other people while your phone is ringing
Pick up the phone even if you are nowhere near your computer. You can often still help verbally.
π€ If you can't help, still point the way
If you truly can't help, say so straight away so nobody is left waiting on you. Then keep going. "Sorry, I can't help" on its own leaves the person exactly where they started, and it reads as though you have handed the problem back.
Even when you are the wrong person, you can still move things forward:
"I can't get to a computer for the next 2 hours, so I can't look at the database myself. Rick knows this system best, so it's worth trying him. I'm not sure he's around, so I'll call Kiki at the same time and let you know either way."
β Figure: Good example - Wrong person, but still helpful. It gives a name to try, takes on part of the chasing, and promises to come back
This costs you very little, and it is the difference between the stakeholder feeling passed over and feeling helped.
π Use the 3 A's
Follow Communication - Do you know the 3 A's for receiving feedback/criticism?:
Acknowledging is not just confirming the facts. Agree that it is a problem and that it is bad. The person needs to hear that you are on the same side of it, not that you have simply logged what they said.
"Yep, I can see the errors."
β Figure: Bad example - Confirms the facts but says nothing about whether it matters, so it reads as detached
"Yes, the site is down, I can see it too. That's a problem, we need to get this sorted tonight."
β Figure: Good example - Agrees it is bad, so the person knows you are taking it as seriously as they are
β° Be specific about your availability
Never leave someone guessing when you will be back. A vague answer forces the stakeholder to either wait and hope, or chase you.
"I'll look when I'm home."
β Figure: Bad example - Vague. There is no timeframe and no fallback, so the stakeholder can't tell whether to wait or escalate
"I'll be home in 30 minutes and will look into it."
β Figure: Good example - Specific. The stakeholder knows exactly when to expect the next update
"I won't be home for the next 2 hours. Are you able to get hold of Tom or Rick?"
β Figure: Good example - Pairs a clear timeframe with a name to try, so the gap in your availability doesn't stall the incident
π₯ Bring in a second person early
A second person is not a luxury during an incident. They halve the stress. They spot the thing you are staring straight past. Best of all, one of you can keep the stakeholder informed while the other investigates.
That split protects the person doing the fixing. Troubleshooting needs concentration, and a constant stream of "any update?" calls and messages breaks it every few minutes. Losing your place in a set of debugging steps is frustrating, and it makes the fix take longer.
A second person becomes a shield. They absorb the questions, keep everyone informed, and let you work uninterrupted. Just make sure the stakeholder knows who to contact so they aren't left guessing.
"Rick is on the phone with me now and is taking over updates. He'll message you every 20 minutes so I can stay heads down on the fix. Call him if you need anything."
β Figure: Good example - Everyone knows who to talk to, so the person fixing it gets a clear run
Call someone from the team if you can. If nobody from the team is free, grab any co-worker who is.
π¨ Escalate early if progress has stalled
It is always better to involve people too early than to sit blocked. Every quiet minute burns the stakeholder's confidence. Escalate as soon as you hit a dead end, and tell the stakeholder you are doing it.
"I've hit a wall. I need more OpenRouter credits and I don't have access to the account. I've called Chris and I'm trying Kiki now as a backup."
β Figure: Good example - The blocker, the person who can unblock it, and the fallback, all in one message
π Give regular status updates, even when there is no progress
Silence is the single biggest cause of "they don't care about this". Short updates every 15-30 minutes keep everyone calm during a major incident.
If you need a longer block of time to investigate, say so upfront so nobody is left waiting.
"I need 30 minutes to get home, then about an hour to dig into the logs. I'll update you at 8pm either way."
β Figure: Good example - Sets expectations for a quiet period, so silence doesn't get read as neglect
π Surface blockers loudly
If you are waiting on another person, chase them actively. Call, message, use whatever works. If you still can't reach them, don't go quiet. Tell everyone that you tried, what is blocking you, and what happens next.
"Sent a message to Kiki. If I don't hear back in 15 minutes I'll call him, then try Rob."
β Figure: Good example - Shows the chase is active and timeboxed, not a passive wait
π₯ Never downplay the urgency
A statement that pushes the problem into the future tells the stakeholder you have mentally clocked off.
"They can look at it on Monday."
β Figure: Bad example - Signals the incident isn't urgent to you, even when it is critical to them
"I'm rolling back to the last good release now. If that doesn't fix it, I'll call Rob tonight to help me restore the database."
β Figure: Good example - Explains what is happening right now and what the next option is
π Make the incident the highest priority on the next business day
An incident over a weekend is not over when the service comes back up. On the next business day:
Nobody can see you working. From the outside, an hour of careful troubleshooting looks exactly the same as an hour of doing nothing.
The fix earns you the outcome. The updates earn you the trust.