Restoring service fast and finding out why it broke are both necessary — and ITIL 4 deliberately keeps them as two different practices with two different clocks.
On a live outage, the instinct to fix the root cause immediately is strong, and it's exactly the instinct ITIL 4 asks teams to resist. Incident management exists to restore normal service operation as quickly as possible, minimizing business impact — speed is the entire objective, even if the fix is a workaround that doesn't touch the underlying cause. Problem management exists to identify, document, and eliminate the root causes of incidents (and to reduce the likelihood and impact of future ones) — thoroughness is the objective, even if it takes far longer than any single incident could tolerate. Conflating the two either slows down incident response chasing root cause under pressure, or lets recurring root causes go uninvestigated because each recurrence gets treated as "just another incident" to work around again.
ITIL 4 links the two practices through a specific artifact: the known error record, created once problem management has identified a root cause but before a permanent fix is deployed. Once that known error exists, the next incident caused by the same underlying issue can be resolved far faster — the incident team applies the documented workaround immediately instead of re-diagnosing from scratch, restoring service in minutes while the permanent structural fix (which might require a code change, a hardware replacement, or a vendor patch, all going through formal change management) proceeds on its own, slower timeline in parallel. This is the mechanism that lets incident response stay fast even while problem management stays thorough — each recurrence gets progressively cheaper to handle without ever forcing the live incident to wait on the deep investigation.
ITIL 4 explicitly frames problem management as able to run proactively, independent of any single incident — analyzing trends across the incident log to spot patterns (like a service degrading slightly every Monday morning under batch-job load) before those patterns ever escalate into a full outage that anyone would call an "incident." Reactive problem management (triggered by an incident or cluster of incidents, as in the diagram above) is only one entry point; proactive problem management is a standing, ongoing practice — reviewing incident data, capacity trends, and monitoring alerts specifically to find and fix root causes before they generate an incident at all. Waiting for incident management to "fail" before engaging problem management throws away most of that proactive value.
Explains the ITIL 4 distinction between incident management (restoring service quickly, even via workaround) and problem management (finding and eliminating root causes to prevent recurrence), and how the known error record hands work off between the two practices.
Both practices deal with things going wrong, both generate tickets in the same service desk tool, and a single underlying defect can generate records in both logs simultaneously — making it easy to treat them as one activity with two names. ITIL 4 keeps them as separate practices specifically because they optimize for opposite things: speed of restoration versus thoroughness of root-cause elimination, and mixing those objectives in one workflow tends to degrade both.
Incident management: the practice of minimizing the negative impact of incidents by restoring normal service operation as quickly as possible. An incident is an unplanned interruption to a service or reduction in its quality. Success is measured primarily by time-to-restore against agreed SLA targets, and a workaround (a temporary fix that doesn't address the underlying cause) is a fully valid and often preferred resolution.
Problem management: the practice of reducing the likelihood and impact of incidents by identifying actual and potential causes, and managing workarounds and known errors. A problem is a cause, or potential cause, of one or more incidents. Root cause analysis (RCA) techniques — the "5 Whys," fishbone/Ishikawa diagrams, fault tree analysis — are used to trace from the symptom back to the structural cause, and the outcome is either a permanent fix (via formal change management) or a documented, faster workaround captured as a known error.
Engineering organizations running critical infrastructure — SCADA historians, PLM systems, CAD license servers, manufacturing execution systems — depend on this separation to avoid two failure modes: incident responders who won't restore service until they fully understand why it broke (extending downtime unnecessarily), and organizations that never invest in problem management at all, so the same outage recurs indefinitely because each occurrence gets logged, worked around, and closed as "resolved" without anyone ever being tasked with preventing the next one.
No — problem management typically engages only when an incident is high-impact, recurring, or its cause is unclear. Low-impact, one-off incidents with an obvious, already-understood cause usually don't warrant a separate root-cause investigation.
A known error is a problem that has a documented root cause and a documented workaround, but does not yet have a permanent fix implemented. It lets the service desk resolve future occurrences quickly using the documented workaround, without duplicating the root-cause investigation each time.
They are often staffed by overlapping but distinct teams — incident management is commonly owned by a service desk or NOC team optimized for fast response, while problem management is often owned by senior engineers, SREs, or a dedicated problem manager role with the time and authority to conduct a multi-day root-cause investigation.
No — a permanent fix normally must go through formal change management (risk assessment, approval, scheduled deployment window), since it modifies the production system structurally, whereas an incident workaround is applied immediately under incident management's own authority specifically because it does not make a lasting structural change.
Try our Engineering Management, Business & Professional Skills Studio
More calculators, simulators, and guides for this discipline.
Data Analysis Studio
✨ Premium ContentA structured, paid professional training program for this discipline.