Many conversations we have about improving incident management naturally turn into conversations about systems and tools (which ITSM platform, what correlation engine, how much automation).
Those are real questions and they matter, but they also tend to take time, money, and configuration work.
Meanwhile the same three gaps show up in team after team, and none of them require anything you don’t already own. They’re process items. A team that isn’t ready for a platform migration can still close all three inside a week, and the effect shows up in the numbers you already report.
Here they are, in the order an incident encounters them.
1. Establish priority levels with definitions
If you’re dealing with incident flow in any volume where everything can’t be handled immediately, you need a priority field. Many teams use some kind of system here, but far fewer have definitions behind it that make it work in practice.
Without them, priority gets assigned by “feel,” which means it varies by engineer, by shift, and by how the last few hours have gone.
The same event gets logged as a P2 on Tuesday and a P3 on Thursday (we use P1s for the highest priority incidents, with P2s, P3s, and P4s following that). Once that’s happening, everything downstream becomes unreliable. Your SLA reporting measures how people felt more than what occurred, and your escalation rules fire against a field nobody really trusts.
The version we use has four levels, and it’s a reasonable starting point for anyone who doesn’t have their own:
P1 is a complete service outage with no redundancy. The customer is hard down and stays that way until it’s resolved.
P2 is severe impairment, or a redundant component failing. They’re still running, but with real limitations. A redundant link down, or a primary path with bad enough latency to hurt.
P3 is partial impairment. Noticeable, affecting service quality, but operations continue. Intermittent connectivity, degraded performance.
P4 is informational or scheduled maintenance. No immediate impact, but somebody should look at it before it becomes something.
Behind those definitions are two questions worth making explicit for your team, because they’re what the definitions are really testing.
Impact: how many people or systems does this affect?
Urgency: what does it cost per minute that it continues?
An outage hitting one workstation and an outage hitting a trading floor circuit are not the same event, and a priority scheme that can’t tell them apart will send your engineers running at the same speed for both.
The other thing definitions fix is the “3 a.m. problem” that exists in organization where systems and infrastructure needs to be up 24x7. When a piece of personal equipment gets powered off at the end of the day and generates a critical alarm, somebody’s phone rings for nothing. Do that often enough and the cost stops being wasted attention and starts being people quitting. Backhaul going down deserves a phone call. (A laptop going to sleep doesn’t.)
We suggest all teams write these definitions for themselves, put them where the team can see them, and spend fifteen minutes in a shift meeting walking through five recent tickets to check whether they were logged consistently. That’s the whole exercise.
2. Build the notification template, and specify what has to be known before it goes out
The complaint we hear more than any other from clients and from teams assessing their own performance, isn’t about resolution speed. It’s about not knowing what’s happening in full.
Teams tend to pour their improvement effort into detecting and characterizing incidents faster, and then leave the communication that follows entirely to whoever happens to be on shift. So, the first notification goes out as a sentence that says something is wrong somewhere, and the next forty minutes get spent answering questions that a better first message would have preempted.
A documented notification process has four pieces, and you can draft all of them pretty quickly:
Who gets contacted, by priority level and by affected service.
How they get contacted. Some stakeholders want a phone call, some want SMS, some want an email they can file.
The technical questions that have to be answered before anyone sends anything.
The template itself, with blanks to fill in.
The third one is the piece is the one we find missing the most in support operations, and it’s where most of the value is a lot of the time.
What is affected, what isn’t, when did it start, what’s being done right now, who’s doing it, and when will the next update come.
An engineer who has to answer those six questions before hitting send produces a message that’s worth receiving. Without them, the notification is a fire alarm with no address on it.
Calibration is worth a moment too, because more communication isn’t automatically better.
Notify too often and the important message gets buried in the routine ones.
Notify too rarely and stakeholders operate blind and start calling.
The answer varies by client and by priority. A P3 usually needs occasional updates. Some teams want every notification because they sort and route internally. And a simple time threshold cuts a lot of noise on its own: if the incident resolves inside a few minutes, don’t send anything. Nobody needs a notification about a port that bounced and came back.
One more thing that lands well and costs basically nothing: Tell people why you’re not doing something. If a carrier is repairing a fiber cut and can’t move faster, and you’re checking in hourly because hourly is the actual cadence of useful information, say that. Stakeholders escalate when they think nothing is happening. Explaining the decision usually stops the escalation.
3. Put a clock on escalations
The third gap is the one that produces some of the worst individual outcomes because it’s invisible while it’s happening.
Here’s what it usually looks like:
An engineer picks up a ticket, works it, and gets stuck. But nobody’s tracking that they’re stuck.
There’s no rule that says when a ticket has to move, so it moves when the engineer decides to ask for help, and people are generally slower to do that than they should be.
The ticket sits. The customer waits. And the post-incident review finds two hours that nobody can account for.
The fix is a time limit paired with judgment rather than replacing it. Set a maximum for how long a ticket can sit at a given tier before it has to escalate, varying by priority. When the clock runs out, it moves, and moving isn’t a failure or a mark against anyone. Then have leads watch for patterns, because someone hitting the limit repeatedly on the same category of incident is telling you exactly where your training gap is.
We find that two things make this work rather than just shifting the pile upward.
Define what an escalation has to contain. What’s been tried, what was ruled out, what the current hypothesis is, and what data has already been gathered. Most escalation friction is a content problem, not a routing problem, and a receiving engineer who has to reconstruct the last ninety minutes from scratch loses most of what the escalation was supposed to save.
Make sure the information the engineer needs arrives with the alarm in the first place. A technician who has to go hunting for the device location, the customer, and the service it supports before they can start is spending their clock on lookup rather than diagnosis.
What to watch after
None of these improvements are particularly complicated, which is exactly why it goes undone. It’s also easy to verify! Pull thirty tickets from the month before you make the changes and thirty from the month after, and check three things: whether the priority assignments look consistent against your written definitions, how long it took for the first notification to go out, and how much time tickets spent at one tier before moving.
If those three move, everything downstream moves with them.
We have a few guide that go quite a bit deeper into this, including the parts that do need tooling:
If you'd rather have someone look at your operation directly and tell you which fixes will pay off fastest in your environment, that's what our NOC Operations Consulting practice does. Every engagement opens with an assessment built on interviews, data collection, and technical analysis, and closes with recommendations we help you put in place. Contact us and we'll set up a discovery call.
📄 Also, be sure to read our free white paper that digs into the NOC and ITOps side of this a little deeper: The NOC Improvement Playbook — 10 Common Problems We See and Solve in Our Consulting Engagements
About INOC, a service of Xerox IT Solutions
INOC is an ISO 27001:2022 certified 24×7 NOC and an award-winning global provider of NOC Lifecycle Solutions®, including NOC support, optimization, design, and build services for enterprises, communications service providers, and OEMs. INOC solutions significantly improve the support provided to partners’ and clients’ customers and end users.
INOC assesses internal NOC operations to improve efficiency and shorten response times, and provides best practices consulting to optimize, design, and build NOC operations, frameworks, and procedures. Proactive 24×7 NOC support is provided with several options, including North America, EU, or APAC only or global integrated NOCs. INOC’s 24×7 staff provides a hands-on approach to incident resolution for technology infrastructure support.
Learn more about our NOC support and NOC operations consulting services. Get in touch to start the conversation. We’d love to talk NOC.




