Part 1 of our mini-series on common NOC problems covered the five issues we find on the path an alarm takes from firing to fixed: fragmented monitoring, missing correlation, scattered incident management, runbooks nobody trusts, and manual work where automation belongs.
Those are the “visible” kinds of fixes we advise on. They’re also the ones that come back if you don’t also fix the elements we’re covering today.
We’ve watched teams do good work on all five of those first problems and then lose most of it inside a year because the conditions that produced the problems were never addressed.
The five below are those conditions. They might be less “satisfying” to fix, but they’re central to a support operation working well, and they’re what separates one that improves from one that improves temporarily before sliding back into stress.
If anything here feels like a familiar operational problem you’re ready to solve, learn more about how we can help and get in touch.
These are excerpted from our NOC Improvement Playbook on the common problem we see and solve for in NOC operations. It’s free to download!
6. A CMDB that doesn't match the network
The configuration management database is a key ingredient of any NOC that’s looking after many devices and services.
It brings order to what’s otherwise chaos: a central repository designed to consolidate and streamline information about IT assets and services, offering a unified and comprehensive view of the entire IT landscape.
At a global tech provider we worked wit, the configuration management database they had in place rarely matched the actual infrastructure as the environment evolved faster than the CMDB was getting updated. During incidents, engineers burned time cross-referencing systems and manually validating configurations, which routinely pushed resolution out by hours on exactly the incidents where that hurt most.
The fiber-optic provider we brought up in Part 1 had it worse. Staff regularly couldn’t find current documentation for their own network configurations, so engineers troubleshot in the dark with very little reliable data about the equipment that had just failed. One fiber hut failure that should have closed inside an hour ran multi-hour because nobody could locate updated configuration information for the affected gear.
In both cases the root cause was the same pair: no regular updates, and no automated tooling keeping the record in sync with reality.
The fix is also a pair, and one half without the other doesn’t hold.
Automated discovery comes first, using something like ServiceNow Discovery or Device42 to continuously scan and update the record as servers appear, routers get reconfigured, and firmware changes. That work typically takes CMDB accuracy from under 50% to near 100% within months.
Then make the CMDB update a mandatory step in change management, so every approved change includes updating the record, whether automatically or through a designated team.
Two things we’d also add from experience. Scope it before you start. Accuracy on the several hundred assets that carry customer-facing services beats partial accuracy across everything you own, and trying to boil the ocean is the most common reason a CMDB project stalls out.
And measure it. Pull a random sample every month, walk it against the real device, and track the percentage that matched. A CMDB nobody audits reverts to fiction on a predictable schedule.
Read our CMDB guide for a deeper dive on developing a NOC-focused system here.
7. Changes the NOC can’t see
This one produces some of the most frustrating outages we’ve assessed because the team is troubleshooting hard and looking in entirely the wrong place!
The banking platform’s operations center attended its Change Advisory Board meetings but never recorded the resulting changes in their ITSM system. That gap led to unauthorized changes and preventable outages, and we watched NOC staff spend time diagnosing problems that were obviously caused by work somebody had done that morning.
The tech services provider held CAB meetings too, but hadn’t connected their monitoring systems to their change management tooling. Changes went in without impact analysis. In one particularly expensive outage, the team spent hours on a diagnosis that an undocumented network change would have answered immediately.
The fiber provider we worked with was the riskiest of the three. Infrastructure changes were made ad hoc with no review at all, and we documented several outages caused by unauthorized reconfigurations that basic change control would have caught.
We did quite a bit of tool- and process-based solutioning for all three:
For the banking platform, we automated ServiceNow workflows to capture and track every change, including emergency modifications made outside the normal CAB path.
At the technology services provider we integrated MicroFocus OBM with their ITSM platform for automatic change tracking, which removed the manual documentation lag and made real impact analysis possible before changes went live.
For the fiber provider we built the process from scratch inside ServiceNow: defined roles, approval chains, change windows, and maintenance schedules.
The single highest-value piece of that work is smaller than that, though.
Surface the change record inside the incident ticket. An engineer opening a P1 should see what changed on that device or in that service in the last 48 hours without leaving the ticket to go ask someone. That one integration removes more wasted troubleshooting time than most of what else you’ll build.
And design for emergency changes first. Every organization has a formal path and an urgent path, and the formal path is rarely where the process fails.
8. No NOC performance data (so no argument)
A shortage of real performance measurement shows up in every NOC we’ve assessed and in nearly every organization that comes to us about outsourcing part or all of their monitoring and service deck.
The fiber provider didn’t have formal KPIs for incident response times, staff utilization, or change effectiveness. So leadership had no visibility into how fast incidents were being resolved or how the team was actually spending its hours. It’s cliche, but you really can’t improve what you can’t measure.
The tech provider tracked basic availability, but the more granular numbers like first-level resolution and mean time to resolution and utilization, were collected by hand or not at all. During post-incident reviews we watched them struggle to identify anything to improve, because they had no detail to find patterns in.
We generally introduce automated reporting through something like Tableau or PowerBI integrated with the ITSM platform, which lets a NOC track MTTA, MTTR, FLR, and utilization without anyone assembling a spreadsheet.
At the global technology provider, one of the more useful findings fell straight out of that: FLR rates varied considerably across teams inside the same NOC. Team-level reporting told leadership exactly where to direct training and resources, which an aggregate number would have hidden completely.
Two practical notes on how to start here:
Pick a small set and baseline it before you change anything, because improvement you can’t compare against a starting point isn’t an argument you can take to a budget meeting. Make FLR one of them.
Report speed alongside quality: rework rates, escalation counts, customer satisfaction. Resolution time on its own tells you how fast the team closed tickets, not whether the tickets stayed closed.
9. Skill concentrated in people you can’t afford to lose
Staffing a NOC for continuous coverage, three shifts a day and 365 days a year, is extremely difficult to plan and expensive to sustain for everyone but the biggest operations (and it’s not easy then, either). If you under-invest in too small of a team, it will absorb too much until it burns out. Over-invest and you drain the budget without touching the operational problems that created the workload.
At the fiber provider, only four people covered 24/7, with no cross-training program. When their primary network troubleshooter wasn’t available, resolution times doubled or tripled. That’s not a staffing shortage, exactly. It’s a single point of failure that’s bound to cause problems.
The tech provider showed us the other end of it: 90% staff turnover in the previous year. Training consisted mostly of shadowing colleagues who were already overloaded, so new hires struggled with routine incidents, and the ones who did develop expertise left for better opportunities. Knowledge came in, matured, and walked out continuously. This is characteristic of any stressful support operation.
The tactical fix is cross-training and scenario-based drills. The diagnostic question we’d start with is blunter: who can you not afford to lose? If a name comes to mind immediately, that’s your finding, and the cross-training plan should be built specifically against that person’s knowledge rather than run as a general program.
At the tech provider, we designated a dedicated subject matter expert to lead training and keep the material current, particularly around newer technology like cloud services. At the fiber provider, we introduced regular cross-training and pushed hard for documented career paths with both technical and management tracks, and turnover improved once people could see somewhere to go.
Don’t skip the soft skills, either. Most NOC training programs are entirely technical, and at the technology provider the gap showed during major incidents, where poor communication and unclear escalation handling stretched resolution times and frustrated internal teams and customers alike. During a serious incident the bottleneck is often communication rather than diagnosis.
Read our full guide to staffing a 24/7/365 NOC. There’s a lot that goes into it.
10. A NOC cut off from the teams it depends on
The last one is structural. When the NOC operates separately from development, security, and second-level support, the cost surfaces exactly when cross-functional work matters most.
At the banking platform, the operations center worked in isolation from development and infrastructure. So during application outages, NOC staff tracked down developers through email chains and phone calls while the service stayed down.
The technology services provider had poor handoff between the NOC and second-level support. Escalations went up without adequate context, which produced long back-and-forth exchanges while the outage continued.
At a large digital services provider, the split was between the NOC and security. DDoS attempts and attempted breaches were detected by the security team, and the NOC often didn’t hear about them until the incident had already progressed.
The underlying cause is the same in all three: different teams, different tools, separate communication channels, and priorities that don’t line up. The NOC can’t function as the hub of incident management when it’s structurally a spoke.
This is why we try to design ITOps to centralize the NOC.
At the technology services provider, that looked like a unified ITSM platform connecting the NOC to every support tier, so escalations carry full context automatically.
At the banking platform we connected monitoring to the development team’s tooling, so application issues routed to the right people with diagnostic data attached, and stood up dedicated incident channels in Microsoft Teams for live collaboration during critical events.
For the security gap, integrating the SIEM with the NOC’s monitoring meant the NOC saw potential threats as they were detected and could respond alongside the security team.
Before you build any of that, write down what an escalation has to contain. Most escalation friction is a content problem rather than a routing problem, and a defined template fixes more of it than an integration does.
A few questions worth answering
If you want a quick read on where your own operation sits, these are among the more revealing questions from our playbook’s self-assessment:
Can you state the exact percentage of the alerts your NOC receives that are noise?
Can your team immediately determine which customers are affected by any infrastructure incident?
Do you have accurate numbers for your first-level resolution rate?
Do you have documented evidence that your escalation procedures actually work 24x7?
Can you accurately trace service dependencies across your infrastructure?
Several “no” or “not sure” answers is the normal result, including at organizations that believe they’re performing well!
Our NOC Operations Consulting practice is here to answer those questions with data rather than impressions. We open with an assessment built around interviews, data collection, and technical analysis of your environment, and end it with recommendations we help you implement rather than hand over.
Contact us to schedule a discovery call, and we’ll tell you what we’re seeing. If you missed part 1, read it here.
📄 The full white paper, The NOC Improvement Playbook: 10 Common Problems We See and Solve in Our Consulting Engagements, covers all ten problems along with several more we see regularly, and closes with a self-assessment you can run against your own operation.
About INOC, a service of Xerox IT Solutions
INOC is an ISO 27001:2022 certified 24×7 NOC and an award-winning global provider of NOC Lifecycle Solutions®, including NOC support, optimization, design, and build services for enterprises, communications service providers, and OEMs. INOC solutions significantly improve the support provided to partners’ and clients’ customers and end users.
INOC assesses internal NOC operations to improve efficiency and shorten response times, and provides best practices consulting to optimize, design, and build NOC operations, frameworks, and procedures. Proactive 24×7 NOC support is provided with several options, including North America, EU, or APAC only or global integrated NOCs. INOC’s 24×7 staff provides a hands-on approach to incident resolution for technology infrastructure support.
Learn more about our NOC support and NOC operations consulting services. Get in touch to start the conversation. We’d love to talk NOC.





