Skip to main content
GullySystem

Incident Management

A written process for what happens after an alert fires — who is notified, who takes charge, how customers are told, and how the issue is reviewed afterwards — so an outage is handled by a process, not a scramble.

What Happens After the Alert Fires

An alert on its own only tells you something is wrong. Incident management is the agreed sequence that follows it — who picks it up, who decides what to tell customers, and who confirms it is actually resolved rather than just quiet again.

What We Define

Severity Levels

A shared definition of what counts as critical, major or minor, so everyone reacts with the same urgency to the same kind of problem.

Incident Roles

Who leads the response and who handles communication, so two people are not both trying to fix the system while nobody updates anyone.

Communication Templates

Prepared wording for updating customers and internal staff, so the first message during an incident is not written from scratch under pressure.

Escalation Path

A clear next step when the first responder cannot resolve the issue alone, instead of it sitting unassigned.

During and After an Incident

A Live Incident Channel

One place where everyone responding communicates, rather than status scattered across calls and messages.

A Timeline Kept as It Happens

What was tried and when, recorded during the incident rather than reconstructed from memory afterwards.

Post-Incident Review

A short, blame-free review of what happened and why, held after the system is stable again.

Actions Tracked to Closure

Fixes identified in the review are assigned and followed up, so the same incident does not repeat for the same reason.

How This Differs From Monitoring

Monitoring detects a problem and raises an alert. Incident management is the defined response once people need to act on that alert together — the two are set up as a pair, but they solve different parts of the problem.

FAQ

Frequently asked questions

Do we need a formal process if we are a small team?

A lighter version still helps — even a one-page severity definition and a named first responder prevents confusion during the first ten minutes of an outage, which is usually when the most time is lost.

What drives the cost of building this process?

How many systems and teams the process needs to cover, and how much of it — severity definitions, escalation paths — already exists informally versus needing to be written from scratch.

What drives how long agreeing the roles and steps takes?

Mostly how much agreement is needed across the people who would be involved in an incident, since the process only works if they have actually agreed to their roles in it.

Do we need an existing on-call rotation for this to work?

No, we can help set one up as part of this if you do not already have one, sized to how many people you have available to share the responsibility.

Who owns the incident process and runbook afterwards?

You do. The severity definitions, roles, templates and runbook are handed over as your own documentation.

What should you have ready for us?

A list of your critical systems, who is realistically available to respond to an incident, and examples of past incidents if any exist.

How are post-incident reviews kept from turning into blame?

The review focuses on what the process allowed to happen and what would have caught it sooner, not on which individual made a mistake, and that framing is agreed with your team before the first review happens.

Talk to us

Tell us what you need.

Send a short brief and one of our engineers will come back to you — usually the same day.

  • No obligation
  • We reply the same working day
  • Your details stay private

Your details are private and secure. Protected by reCAPTCHA.