← Engineering Management, Business & Professional Skills Studio
Concept Explainer · ITIL 4 Service Management

Incident Management vs Problem Management

Restoring service fast and finding out why it broke are both necessary — and ITIL 4 deliberately keeps them as two different practices with two different clocks.

On a live outage, the instinct to fix the root cause immediately is strong, and it's exactly the instinct ITIL 4 asks teams to resist. Incident management exists to restore normal service operation as quickly as possible, minimizing business impact — speed is the entire objective, even if the fix is a workaround that doesn't touch the underlying cause. Problem management exists to identify, document, and eliminate the root causes of incidents (and to reduce the likelihood and impact of future ones) — thoroughness is the objective, even if it takes far longer than any single incident could tolerate. Conflating the two either slows down incident response chasing root cause under pressure, or lets recurring root causes go uninvestigated because each recurrence gets treated as "just another incident" to work around again.

Incident management: restore service, now

Reactive · Fast
OUTAGEdetected 09:14TICKET LOGGEDpriority: P1WORKAROUNDrestart the failing serviceSERVICE UP09:26 — 12 minCLOSEDSuccess measured by time-to-restore (12 minutes) — why the service crashed is not answered hereSame crash recurs three more times this month — root cause still unknown
Primary goal
Restore service, minimize impact
A workaround that gets users back up is a complete success, even without a root cause finding.
Key metric
MTTR (mean time to restore)
Tracked per ticket, reported against SLA targets by priority level.

Problem management: eliminate the root cause

Proactive · Thorough
4 recurring incidents this month, all logged against the same crashing servicePROBLEM RECORD OPENEDlinked to all 4 incidentsROOT CAUSE ANALYSIS (RCA)memory leak found in v2.3 — takes 6 daysKNOWN ERROR + FIXvia formal change request
Primary goal
Eliminate or reduce recurrence
Success is fewer future incidents, not a fast individual restore.
Key metric
Reduction in recurring incident volume
Tracked over weeks/months, not per-ticket — measures whether the fix actually worked.
Why this works

A known error record is the handoff between the two practices

ITIL 4 links the two practices through a specific artifact: the known error record, created once problem management has identified a root cause but before a permanent fix is deployed. Once that known error exists, the next incident caused by the same underlying issue can be resolved far faster — the incident team applies the documented workaround immediately instead of re-diagnosing from scratch, restoring service in minutes while the permanent structural fix (which might require a code change, a hardware replacement, or a vendor patch, all going through formal change management) proceeds on its own, slower timeline in parallel. This is the mechanism that lets incident response stay fast even while problem management stays thorough — each recurrence gets progressively cheaper to handle without ever forcing the live incident to wait on the deep investigation.

Common misconception
"Problem management only starts after incident management fails to find a quick fix."

ITIL 4 explicitly frames problem management as able to run proactively, independent of any single incident — analyzing trends across the incident log to spot patterns (like a service degrading slightly every Monday morning under batch-job load) before those patterns ever escalate into a full outage that anyone would call an "incident." Reactive problem management (triggered by an incident or cluster of incidents, as in the diagram above) is only one entry point; proactive problem management is a standing, ongoing practice — reviewing incident data, capacity trends, and monitoring alerts specifically to find and fix root causes before they generate an incident at all. Waiting for incident management to "fail" before engaging problem management throws away most of that proactive value.

Related Concept Explainers
Risk vs Issue
Read next →
Leading vs Lagging Indicators
Read next →

Incident Management vs Problem Management — Concept Explainer

Explains the ITIL 4 distinction between incident management (restoring service quickly, even via workaround) and problem management (finding and eliminating root causes to prevent recurrence), and how the known error record hands work off between the two practices.

Why This Is Commonly Confused

Both practices deal with things going wrong, both generate tickets in the same service desk tool, and a single underlying defect can generate records in both logs simultaneously — making it easy to treat them as one activity with two names. ITIL 4 keeps them as separate practices specifically because they optimize for opposite things: speed of restoration versus thoroughness of root-cause elimination, and mixing those objectives in one workflow tends to degrade both.

The ITIL 4 Definitions

Incident management: the practice of minimizing the negative impact of incidents by restoring normal service operation as quickly as possible. An incident is an unplanned interruption to a service or reduction in its quality. Success is measured primarily by time-to-restore against agreed SLA targets, and a workaround (a temporary fix that doesn't address the underlying cause) is a fully valid and often preferred resolution.

Problem management: the practice of reducing the likelihood and impact of incidents by identifying actual and potential causes, and managing workarounds and known errors. A problem is a cause, or potential cause, of one or more incidents. Root cause analysis (RCA) techniques — the "5 Whys," fishbone/Ishikawa diagrams, fault tree analysis — are used to trace from the symptom back to the structural cause, and the outcome is either a permanent fix (via formal change management) or a documented, faster workaround captured as a known error.

Where This Matters in Engineering and IT Operations Practice

Engineering organizations running critical infrastructure — SCADA historians, PLM systems, CAD license servers, manufacturing execution systems — depend on this separation to avoid two failure modes: incident responders who won't restore service until they fully understand why it broke (extending downtime unnecessarily), and organizations that never invest in problem management at all, so the same outage recurs indefinitely because each occurrence gets logged, worked around, and closed as "resolved" without anyone ever being tasked with preventing the next one.

Frequently asked questions

Does every incident get its own problem record?

No — problem management typically engages only when an incident is high-impact, recurring, or its cause is unclear. Low-impact, one-off incidents with an obvious, already-understood cause usually don't warrant a separate root-cause investigation.

What is a "known error" exactly?

A known error is a problem that has a documented root cause and a documented workaround, but does not yet have a permanent fix implemented. It lets the service desk resolve future occurrences quickly using the documented workaround, without duplicating the root-cause investigation each time.

Who owns problem management versus incident management in an organization?

They are often staffed by overlapping but distinct teams — incident management is commonly owned by a service desk or NOC team optimized for fast response, while problem management is often owned by senior engineers, SREs, or a dedicated problem manager role with the time and authority to conduct a multi-day root-cause investigation.

Is a permanent fix from problem management deployed through the same process as an incident workaround?

No — a permanent fix normally must go through formal change management (risk assessment, approval, scheduled deployment window), since it modifies the production system structurally, whereas an incident workaround is applied immediately under incident management's own authority specifically because it does not make a lasting structural change.

🎓

Try our Engineering Management, Business & Professional Skills Studio

More calculators, simulators, and guides for this discipline.

🎓

Data Analysis Studio

Premium Content

A structured, paid professional training program for this discipline.

Related tools & guides

Business & Professional Skills StudioData Analysis Studio