Retail chain · report · 2026-10-01 10:11 UTC

Confidential

This is a sample. The log is a year of tickets of a fictional retail chain, built like real life: background failures, requests in the same queue and a few planted deviations.
Upload your own log at opslab.consulting - the report will look the same, but about your team.

What was checked

This report is built only on these sources:

Source What it held
Incident journal An export from your ticket system for the period 364 days, 1728 records. Fields: time registered, time closed, service or node, priority, description or cause.
List of critical services As named by you: Store tills, Payment gateway, Online shop.
Metric exports Not uploaded. We need CPU, memory and link load, packet transfer delay, packet loss ratio, delay variation, backup success and certificate expiry - a measurement time and a value column for each one. (see the check map below)
Change log Not uploaded. We need the date and time of each change, its type (standard or emergency), the service it touched, and the outcome. (see the check map below)
First-line support data Not uploaded. We need the resolving group or support line for each ticket, an escalation flag, and a reopen flag. (see the check map below)

Summary

Across 364.2 days, incidents produced 1668.40 hours of elapsed downtime. Typical restoration took 54 minutes, but a few long incidents absorb much of the recovery effort.

Incident volume has been higher since early September, moving from about 3.23 to 4.91 incidents per day. ERP has incidents in the last three days of almost every month. The next expected window is 2026-01-29. 2026-01-31 if nothing changes.

  1. ERP has incidents at every month-end - Month-end is the busiest trading and accounting period, so a repeated ERP incident hits the closing process.
  2. Incident rate changed upward in early September - This is a sustained increase, not a short spike, so teams are carrying about half as many more incidents each day.
  3. Store network incidents cascade into sales services - A single link problem is seen as multiple tickets across the checkout path and sales-floor connectivity, obscuring the real scope.

Check map

The map holds every kind of check we offer. What fills it depends on the measures you can still add.

Legend: measuredneeds your field or your targetneeds another data sourceavailable in the extended review

Does the infrastructure keep the promised level - 1/4
Total downtimeDetails

What this is. Total time your services were unavailable over the log period.

How we count it. We add up the duration of all incidents, but merge overlapping intervals: two incidents running at the same time from 10:00 to 11:00 count as one hour of downtime, not two. Without merging, the number is inflated more the larger the infrastructure is.

Where the definition comes from. Our calculation. The concept of downtime comes from availability in the ITIL 4 glossary: the ability of a service to perform its agreed function when required.

What we need from you. Already calculated - see the number in the row.

If the row is amber: we need a field with the incident closing time, otherwise the duration cannot be computed.

69d 12h 24m
Availability of critical servicesDetails

What this is. The share of observation time the service was up. This is the number people call "three nines".

How we count it. Uptime / (Uptime + Downtime), calculated separately for each service you named. An incident that hit three services counts in full for each of the three: from that one service's point of view, it was down the whole time.

Where the definition comes from. Formula - Google SRE Book, ch. 3 "Embracing Risk", verbatim. The concept of availability - ITIL 4 glossary.

What we need from you. Mark the critical services in the list and state the target level (e.g. 99.9%). We calculate right after you answer - no need to re-upload the report.

Critical services The names come from the "service" field of your journal, as they are. If that field holds devices or device models, availability is counted for the device, not for the service it carries.

choose your critical services and your target level
Error budget spentDetails

What this is. A target like 99.9% allows a certain amount of downtime by definition. That allowed time is the budget. This indicator shows how much of the budget you have already spent.

How we count it. Allowed downtime = period length x (1 - target). Spent = actual downtime / allowed downtime x 100%. We use the actual period - exactly as many days as are in your log, with no rounding to a month. A value above 100% means the budget is used up and the target was missed for the period.

Where the definition comes from. Google SRE Book, ch. 3: the budget is the gap between the target and the actual uptime - the remaining "unreliability" for the period.

What we need from you. One number - the target availability. Without a target the indicator does not exist: the budget is counted from a promise, not from the data.

set your target - without it the measure does not exist
Actual recovery time against the targetDetails

What this is. RTO (Recovery Time Objective) is the deadline for a service to come back after an outage. This indicator shows how often you met that deadline.

How we count it. The share of incidents no longer than your target, plus the worst case for the period - the longest incident. The worst case matters more than the average: a continuity plan is tested on a bad day, not on a normal one.

Where the definition comes from. ISO 22300:2021 (vocabulary for ISO 22301): RTO is "the period of time following an incident within which a product, service or activity is resumed, or resources are recovered". The standard is paid, so we have not verified the clause number and do not cite one.

What we need from you. Target recovery time in hours. If it differs by service, give the target for the most critical one.

set your target - without it the measure does not exist
How fast you repair - 1/6
Time to restore a serviceDetails

What this is. How long it takes from an incident being logged to being closed.

How we count it. We give three numbers at once: median, mean, and P95. The median is a normal day: half the incidents are fixed faster. The mean is pulled up by rare, long cases. P95 is the "tail": only one incident in twenty takes longer. A single average without the tail hides exactly the cases that make the business call.

Where the definition comes from. ITIL 4 calls this MTRS, mean time to restore service: "a measure of how quickly a service is restored after a failure". The reliability vocabulary IEC 60050-192 calls the same number MTTR (term 192-07-23).

What we need from you. Already calculated.

If the row is amber: we need a field with the incident closing time.

54 min / 1h 28m / 4h 41m - median / mean / P95
Time to detectDetails

What this is. How long an incident ran before you knew about it. This is a blind spot: the service is already down, but there is no ticket yet.

How we count it. Actual incident start minus the time it was logged in the system. This cannot be derived from the logging time alone under any assumption: the log knows when the problem was reported, not when it began.

Where the definition comes from. No standard exists. MTTD is an established incident-management practice, and we say so plainly instead of dressing practice up as a norm.

What we need from you. A target will not fix this - we need an export with a separate "actual start" field (systems name it differently: start time, occurred at, time of occurrence). Upload such an export via the same link, and the row will be calculated.

needs the field "real start time, separate from the time of registration" filled in your database
Time to respondDetails

What this is. How long a ticket waited for someone to pick it up. This shows how the queue and the on-call shift are performing, not how hard the fault was.

How we count it. Time work started minus the time the ticket was logged.

Where the definition comes from. No standard exists. MTTA is a service-desk practice.

What we need from you. We need an export with a "work started" field (picked up, assigned, acknowledged - whatever your system calls it).

needs the field "time accepted by an engineer" filled in your database
Share closed on timeDetails

What this is. What share of incidents met the deadlines agreed with the business.

How we count it. Calculated separately per priority: how many incidents of that priority closed within your target. Priorities cannot be combined into one number - their targets differ, and a combined figure would hide overruns on the critical ones. We only calculate this when priority is filled in for at least 80% of records; below that threshold the number reflects log data quality, not service performance.

Where the definition comes from. ITIL 4: service level - metrics of expected or achieved quality; SLA - a documented agreement on required services and the expected level of service.

What we need from you. Target times per priority, in minutes (e.g. critical - 60, high - 240, normal - 1440).

If priority is not tracked in the log, or is filled in for less than 80% of records: we first need an export with this field - a target has nothing to apply to without priority.

set your target - without it the measure does not exist
Share solved the first timeDetails

What this is. What share of incidents did not reopen after being closed. A direct measure of repair quality: was the ticket closed, or was the problem solved.

How we count it. The share of incidents with no reopening marker.

Where the definition comes from. No standard exists - this is service-desk practice.

What we need from you. We need an export with a "reopened" field, or a status-change history - from the history we can reconstruct the fact of reopening ourselves.

needs the field "reopened" filled in your database
Share solved on the first lineDetails

What this is. What share of incidents were closed without escalating to specialists. Shows how much work moves upward and how loaded the expensive hands are.

How we count it. The share of incidents closed by the same group that took them in.

Where the definition comes from. No standard exists. ITIL 4 defines only escalation and service desk; the metric itself is not in the glossary.

What we need from you. We need an export with a "resolver group" or "support line" field.

needs the field "resolver group" filled in your database
How often and what exactly breaks - 4/8
How often incidents happen and how often a service failsDetails

What this is. How often something breaks in the whole company and how often each of the most frequent services fails.

How we count it. The period divided by the number of incidents: for the company - the overall rate, for a service - the time between failures. We count service failures, not hardware reliability.

Where the definition comes from. ITIL 4, MTBF: "a measure of how frequently a service or other configuration item fails". IEC 60050-192, term 192-05-13, mean operating time between failures.

What we need from you. Already calculated.

1369 incidents, on average one every 6h 23m. Most often: Store tills - once every 21h 29m; Online shop - once every 2d 05h 38m; Mobile app - once every 2d 23h 04m
Downtime concentration by serviceDetails

What this is. Shows whether your downtime is concentrated in a few points or spread evenly. If concentrated, work on a short list gives a disproportionately large effect.

How we count it. We rank services by their contribution to downtime and check how much of it the top ones account for. We also calculate an inequality score between 0 and 1 (named in Appendix A for readers who want the formula): 0 means "all services contribute equally", closer to 1 means "almost everything falls on a few".

Where the definition comes from. Our calculation. No borrowed norms exist: we compare your infrastructure with itself, not with someone else's average.

What we need from you. Already calculated.

If the row is amber: we need a "service or node" field - without it downtime cannot be attributed. If grey-blue: this one comes with the extended review.

3 of 13 services carry 53% of all downtime
Returns to the same objectDetails

What this is. What share of failures came back on the same object - server, CI, store - within a week after the fix, and whether that is more than chance gives.

How we count it. The key is the object together with its system: "the till in store 12", not "tills". Tickets of one outage on an object (opened while the previous one is open, or within an hour after it closed) count as one failure. A return is a failure within 7 days after the end of the previous one. Next to it: the share chance would give if all objects of a system failed alike. An object is named when it fails more often than the other objects of its system, or its failures come in series more often than the others' rhythm gives; both tests are corrected for the number of objects. The same object is not yet the same cause: in ITIL 4 a repeat points to a problem, "a cause, or potential cause, of one or more incidents".

Where the definition comes from. Our calculation; interpretation - ITIL 4, problem.

What we need from you. Already calculated.

If the row is amber: we need an object column - server, CI or store - filled in at least one ticket of five. If grey-blue: available in the extended review.

34% of failures came back on the same object within 7 days after the fix; if objects failed alike, about 33%
A metric before failuresDetails

What this is. Which monitoring metric - CPU load, memory, traffic, disk - moves away from usual before failures, and about how long before them.

How we count it. We put the journal and the metrics export on one time axis. "Usual" for a metric is its level at the same hour of the day, workdays and weekends apart. We look at windows 15 min - 1 h, 1-6 h and 6-24 h before the ticket and compare how often the metric was above or below usual before failures and at the same hours on other days. A finding needs this to happen clearly more often before failures, on different days, corrected for the number of series. A ticket opened by an alert on the same metric is not a precursor.

Where the definition comes from. Our calculation; tested on artificial months and years of monitoring with a known answer.

What we need from you. Already calculated.

If the row is grey: send a monitoring metrics export for the same period together with the journal. If amber: the common period or its failures are too few - it needs 14 days and 8 failures or more.

needs another data source: a monitoring metrics export for the same period (Zabbix, Grafana, Prometheus)
Breakdown by priorityDetails

What this is. How incidents are distributed across importance classes.

How we count it. We calculate the share of each priority, but only if the field is filled in for at least 80% of records. Below that, the breakdown reflects log-filling discipline, not the real picture of breakages, and we do not print it. Look at two things. First, the share of the top classes: if critical is more than a quarter of the flow, priority is set by how loud the reporter is, not by business impact, and the urgent lane stops being urgent. Second, compare it with "how fast you repair": critical incidents repaired no faster than normal ones mean the priority exists in the log but not in the shift's work.

Where the definition comes from. ITIL 4, incident: an unplanned interruption to a service or a reduction in its quality. The standard does not define priority scales - we use yours.

What we need from you. Already calculated.

If the row is amber: we need an export with a "priority" field, or for it to be filled in.

P3 - MEDIUM - 667 (48%); P4 - LOW - 393 (28%); P2 - HIGH - 270 (20%); P1 - CRITICAL - 53 (4%)
Share of incidents with a known root causeDetails

What this is. What share of incidents were traced to a root cause, rather than just closed after the service was restored.

How we count it. The share of records with a root cause filled in, or linked to a problem record.

Where the definition comes from. ITIL 4: problem - the cause, or potential cause, of incidents; known error - a problem that has been analysed but not resolved (a workaround is not required for this).

What we need from you. This field is not tracked in your log. That fact alone is already a finding: repeat incidents are closed without root-cause analysis, so they will return. To see a number, we need an export with a "root cause" field, or a link between incidents and problems.

the field is not kept - so repeat incidents are closed without anyone finding the cause
Share of emergency changesDetails

What this is. What share of changes are made in "needed it yesterday" mode, without the usual approval. A high share means the planned process cannot keep up with reality.

How we count it. The share of emergency changes out of all changes in the period.

Where the definition comes from. ITIL 4: emergency change - a change that "must be implemented as soon as possible"; standard change - a pre-authorised, low-risk change.

What we need from you. A field in the incident log will not fix this: the denominator here is changes, not incidents. We need an export from the change log for the same period. It will also refine the "incidents with a change trace" row - currently that row is a lower-bound estimate from description text.

needs another data source: a change log
Share of incidents caused by suppliersDetails

What this is. What share of breakages came from outside your perimeter - a provider, vendor, or contractor. This is a number for a conversation with the supplier, not with your team.

How we count it. The share of incidents with an external party named as at fault.

Where the definition comes from. ITIL 4, supplier management practice.

What we need from you. We need an export with an "at-fault party" or "supplier" field.

needs the field "party at fault" filled in your database
Capacity and health
This is where the state of your resources belongs: load, network delay and loss, backup success, certificate expiry, out-of-support components. An incident journal does not hold them - we need a separate metrics export for the same period: the time of each measurement and one value column per measure.

In total: 6 measures computed, 10 wait for your decision or a filled field, 8 need other data.

Every line marked "needs your field or your target" is not our limit - it is the name of a field in your database. Send us the field "time accepted by an engineer" and you will get response time separately from repair time. Then you will see where you lose more: while the dispatcher looks for a free engineer, or while the engineer repairs.

What we did not get

Below is what would close the most rows of the map. Answering commits you to nothing.

We got no metrics export. Do you collect resource load, loss, latency or response time?

The records name no cause. Do you keep repeating causes separately, as problems?

There was no change log in what you sent. Do you keep one, with a date and time for every change?

Your step and advice

Step 2 of 5 - Recording

Outages are recorded. That history already shows how long recovery takes, where downtime piles up and what outages to expect next.

The step is counted from the log only. Functions the log does not show are not counted until you tell us about them, so this is a lower bound. Confirmed by data: 1 of 1 key functions up to step 2.

What to add to the export

Check that the export has a planned recovery time or a "target missed" mark. Deadlines for the post-incident review do not count.
The share of outages closed on time and the remaining error budget become visible in the data.
Check that the export has the time the outage was picked up and the line or group.
Reaction time and the share solved at the first line become visible in the data.

Key functions

StepFunctionITIL 4 practiceStatus
2 Every outage is recorded Incident management confirmed by data
3 Main services have owners Service level management confirmed by data
3 Priority is set by impact and urgency Incident management confirmed by data
3 Recovery times are promised Service level management not seen in the data
3 An on-call rota and a hand-over order Service desk not seen in the data
4 Repeats are investigated down to the cause Problem management not seen in the data
4 Changes are recorded with a type and a result Change enablement not seen in the data
4 Outages are known before the consumer notices Monitoring and event management not seen in the data
5 Improvements have an owner and a measured effect Continual improvement not seen in the data
5 Measures are reviewed with the consumer Measurement and reporting not seen in the data
5 Service dependencies on equipment are known Service configuration management not seen in the data
What this is based on

The assessment scheme follows formal process assessment (ISO/IEC 33000, the CMMI appraisal method): what people say and what the work records show. An answer makes a function "declared", the export makes it "confirmed", a mismatch makes it "disputed".

A function is confirmed by data when the log covers at least 90 days and 30 records, the start, end and object of an outage are filled in 95% of records, the function's own field (service, priority) in 80%, and no single value stands on more than 95% of records.

Practices and their names come from ITIL 4. The ladder rests on ITIL 4, CMMI and ISO/IEC 33000; it is not a certification and not an assessment by an accredited assessor.

What the steps mean

Workload profile

What we looked at. 1728 journal records over 364 days.

What we found. Out of 1728 records, 320 (18%) are service requests by the column “Issue Type”, not failures; 45 of these requests have no type, we recognised them by the words of the description, so this part is an estimate; another 25 are changes, problems or planned works. We took all these records out of every calculation in this report: below we work with 1369 incidents, which is 3.8 per day. In total, services were down or worked worse than normal for 69d 12h 24m. This is the sum over all objects, not the time when everything was down at once.

What the log holds: failures and the rest

A bar is how many log records of each kind. We took the orange ones out of every calculation in the report: they are not failures, and their completion time does not go into the time to restore.

Log rows we took out

TicketStartType or description
INC0250002025-01-01 08:54:09Service request
INC0250022025-01-01 12:51:31Service request
INC0250122025-01-03 11:41:38Service request
INC0250622025-01-17 10:02:51Unlock my account after holiday
INC0250762025-01-20 16:48:40Please install the e-signature software

What it means. Counted all together, there are 25% more failures than really happened, and the repair time mixes with the time to fulfil requests. A management report on all tickets shows more outages than there were. About downtime: if we simply added the durations one after another, we would get 83d 06h 19m. The difference of 13d 17h 54m is incidents that ran at the same time. Counting them twice inflates the downtime.

What to do. Count failure numbers only on records with the incident type in the column “Issue Type”, and make sure the type is always filled.

Check it yourself.* Filter the export by the column “Issue Type”: you should get 300 records typed as a request, change or problem.

How fast services are restored

What we looked at. Time from registration to restoration for 1369 incidents. This is calendar time: nights and weekends count in full. 14 tickets are not closed: their end is unknown, so they are left out. Priority 1-2 among them: 2, the oldest open since 22.12.2025. If this is an outage going on now, it is larger than any number above.

What we found. Half of the incidents close faster than 54 min. The average is 1h 28m. The worst 5% run longer than 4h 41m; the longest was 19h 30m. Those 69 incidents hold 24% of all downtime.

Time to restore - how many incidents take how long

Half of the incidents are fixed within 54 min, the average is 1h 28m. The 5% longest (right of 4h 41m) hold 24% of downtime.

Largest incidents and their share of downtime

The 3 largest incidents account for 2.4% of total downtime - most downtime builds up from many ordinary incidents, not from a few outages.

Log rows: the longest incidents

TicketStartServiceObjectLength
INC0253112025-03-15 22:10ERPERP-0119h 30m
INC0264982025-11-26 06:05Warehouse systemWMS-0115h 00m
INC0257882025-07-08 14:42Online shopWEB-0113h 15m
INC0258072025-07-13 13:43Store Wi-FiStore #3311h 48m
INC0253152025-03-16 19:09BackupBACKUP-0110h 20m

What it means. The average (1h 28m) is above the median (54 min) - repairs almost always look like this: a few long cases pull the average up. But there is no separate class of heavy incidents: the 5% longest hold only 24% of downtime, most of it comes from ordinary repairs. Report the median - it shows a typical incident.

What to do. Create a separate procedure for incidents that are not closed within 4h 41m. Not "tighten control", but exactly this: after that time, escalate to an engineer who is allowed to call the vendor. Today such incidents stay in the common queue.

Check it yourself.* Open the last three incidents longer than 4h 41m and look at how much time went into diagnosis and how much into the repair itself. If more than half went into diagnosis, the bottleneck is in monitoring, not in the hands of your engineers.

Problem areas

What we looked at. How incidents and downtime spread across 13 objects, and repeats.

What we found. 3 objects out of 13 hold 53% of all downtime. The worst is "Store tills": 407 incidents, 19d 14h 50m of downtime. With an even spread it would be 23%. For "Store tills", 77% of the downtime comes from priority 3-4 tickets. That is ticket time, not necessarily business downtime: check whether the service was really down. Tickets closed by a timer or a queue clean-up, and open tickets, are not in the downtime (details in section 2). 34% of failures on objects (435 of 1296) came back on the same object of the same system within 7 days after the fix. If every object of each system failed alike, it would be about 33%. Tickets of one outage on one object count as one failure: 87 such tickets. "ERP-01" in "ERP": 49 failures in the period, a typical object of this system has 22: it fails more often than the others, and chance does not give that. "Store #17" in "Store tills": 29 failures in the period, a typical object of this system has 9: it fails more often than the others, and chance does not give that.

Where the downtime sits

A bar is a share of all downtime; the dashed line is what each would hold if downtime were shared evenly. Orange: the ones this section names.

Objects where failures come back

A row is an object, a mark is a failure. A red dot came within 7 days after the previous fix, a dark tick after a break. On the right: failures on the object against a typical object of the same system, or how many came back.

Log rows: the longest incidents of “Store tills”

TicketStartServiceObjectLength
INC0260592025-09-10 09:52Store tillsStore #98h 08m
INC0261622025-09-28 15:39Store tillsStore #347h 58m
INC0262362025-10-09 11:08Store tillsStore #407h 33m
INC0260982025-09-18 15:41Store tillsStore #65h 38m
INC0264002025-11-08 08:34Store tillsStore #54h 55m

Log rows: failures on “ERP-01”

TicketStartServiceObjectLength
INC0263522025-10-31 16:15ERPERP-012h 21m
INC0265062025-11-27 10:25ERPERP-014h 56m
INC0265092025-11-27 16:21ERPERP-011h 36m
INC0265162025-11-28 12:40ERPERP-013h 34m
INC0267202025-12-31 13:15ERPERP-011h 20m

What it means. A few objects hold the downtime, and more than chance would give. People and reviews sent there win back the most time. The named objects fail again and again. The same object is not yet the same cause, and the journal does not show the cause. Possible causes: the cause was not removed - it was fixed by a restart or a workaround; the object is worn or overloaded; the object has more equipment or people than its neighbours; one outage is logged again later. What tells them apart: the description and the closing notes of these tickets, a record in the problem log, and how much equipment the object has.

What to do. Start with 3 objects, not 13. For "Store tills", review its incidents together, in one review, and look for a common cause instead of the cause of each one. Review the failures of "ERP-01" together, in one review: compare the descriptions and what was done at closing. One cause, fixed by a restart - open a problem and look for the cause. Different causes, and the object is bigger than its neighbours - compare it with objects of the same size. The same outage logged again - agree to link a repeat ticket to the first one.

Check it yourself.* Open any three incidents from "Store tills" from the last month. If they have different causes, we were wrong and the object is simply overloaded. If the cause is the same, you have a problem that nobody owns. Send the journal again in a quarter: "ERP-01" should have fewer failures and fewer returns within a week.

Incident sources

What we looked at. Which systems fail together: after a failure of one, a ticket for another opens within 15 minutes more often than chance allows.

What we found. After a failure of "Store network links", tickets for other systems open within 15 minutes: "Store Wi-Fi" 14 times out of 71, "Store tills" 11 times out of 71, "Payment gateway" 9 times out of 71. By chance, given these systems' own rhythm, it would happen at most 0.7 times over the whole period. 80 tickets out of 1383 opened in the first 15 minutes of an outage already under way: they continue it, or register the same outage again. The number of failures and the time between them in this report are counted in tickets, so they are overstated by about 6%; in the list of the largest incidents such an outage counts as one.

Failures that pull tickets on other systems

A bar is how many times a ticket of the second system opened within 15 minutes after the first one failed; the grey tick is what chance gives with that system's own rhythm. The arrow shows which system is first; a double arrow means neither is regularly first.

Log rows: “Store network links” and right after it “Store Wi-Fi”

TicketStartServiceObjectLength
INC0265752025-12-08 06:20Store network linksStore #442h 41m
INC0265772025-12-08 06:31Store Wi-FiStore #442h 33m
INC0267062025-12-29 21:00Store network linksStore #3013 min
INC0267082025-12-29 21:10Store Wi-FiStore #3017 min

What it means. Tickets for "Store Wi-Fi" in the first minutes after a failure of "Store network links" are more likely a consequence than a separate problem. There are three usual reasons: "Store Wi-Fi" works through "Store network links"; both depend on a third thing - power, a shared node, a site; or two teams registered one outage. The object column (store, node, site) tells them apart: if a pair of tickets has the same object, it is one outage.

What to do. Link tickets for "Store Wi-Fi" opened in the first 15 minutes after a failure of "Store network links" to its ticket as child tickets. Then failures stop being counted twice, and people fix one cause. If "Store Wi-Fi" works through "Store network links", "Store network links" needs the backup.

Check it yourself.* Open the last three outages of "Store network links" and the tickets for "Store Wi-Fi" in the next 15 minutes: is it the same object, and did they close together with the outage?

Periods of higher risk

What we looked at. How incidents spread across hours and days of the week, and how the rate changed over time.

What we found. By time of the week: weekdays 9 to 18 hold 46% of incidents (27% of the hours of a week), weekday evenings and nights 28%, weekends 27%. Thu 02:00-03:00: 34 incidents, where your usual rhythm would give about 4. Failures in this hour came back in 31 of 52 weeks. Most of them are Warehouse system (28 of 34). ERP: failures on the last three days of the month - in 11 of 12 months. Failure days on these dates: 22, where the usual rhythm would give about 7. Last time: 30.12.2025. If nothing changes, the next one by the cycle is 29.01.2026-31.01.2026. The incident rate went up: from 3.2 a day at the start of the period to 5.1 at the end. These data cannot tell a step from a slow rise. If it was a step, it came between 13.08.2025 and 24.09.2025, most likely around 01.09.2025: it was 3.2 a day, it became 4.9. If it was a slow rise, there is no single date. The new level has held for 122 days now.

Incidents by weekday and hour

The darker the cell, the more incidents. Weekdays 9 to 18 (dotted frame) hold 46%, evenings, nights and weekends 54%. The red frame is an hour that breaks your usual rhythm; the number is its incidents.

Failures that come back on a cycle

Each circle is a day of the cycle: red when the system failed that day, hollow when it passed quietly. Ticks are failures on other days. The hollow red circle after the end of the export is the next day by the cycle, if nothing changes.

Log rows: incidents in the marked hour

TicketStartServiceObjectLength
INC0263892025-11-06 02:21Warehouse systemWMS-011h 59m
INC0264672025-11-20 02:15Warehouse systemWMS-011h 03m
INC0265052025-11-27 02:23Warehouse systemWMS-012h 04m
INC0266332025-12-18 02:04Warehouse systemWMS-011h 34m
INC0266342025-12-18 02:25Payment gatewayPAY-011h 19m

Log rows: failures of “ERP” on the cycle days

TicketStartServiceObjectLength
INC0261742025-09-29 16:54ERPERP-011h 54m
INC0263382025-10-29 19:44ERPERP-0321 min
INC0265162025-11-28 12:40ERPERP-013h 34m
INC0267112025-12-30 09:47ERPERP-014h 28m
INC0267162025-12-30 15:04ERPERP-014h 14m

What it means. A failure tied to an hour is almost always tied to a schedule: a night job (exchange, export, backup), a release window, planned works of a supplier, or a check that writes what it found at its own time. Random failures do not gather like this. A failure that comes back at equal intervals is almost always tied to a schedule or to something that fills up: a job every few days, month-end closing and reports, disk space or memory that runs out by the same time, a certificate or password that expires. Random failures do not repeat like this. The journal does not say what happened: it has no cause field. Options: more users or sites, a release or a migration, a new contractor, new logging rules - for example, monitoring started to open tickets by itself. The change log for those weeks and a "registered by" column (person or monitoring) will tell them apart.

What to do. Find what runs at this time: the job schedule, the release window, supplier works. If it is a job, move it or add a check before it starts; if it is releases, add a check right after the rollout. Compare these dates with your calendar of month-end closing, reports and payments: what loads the system on these days. Give it resources in advance and put someone on duty who knows the system. Find what changed in those weeks. If it is business growth, recount the on-call shift for the new level. If it is a release or a contractor, return the question to whoever made the change.

Check it yourself.* Open the tickets of Thu 02:00-03:00: if they share one system and similar text, it is one repeating cause. The time is taken as it is in the export. If your system writes time in UTC, shift the hours to your time zone. Open the tickets of ERP on the cycle dates (29.09.2025, 29.10.2025, 28.11.2025, 30.12.2025): if the text is the same, it is one cause. If the system was restarted between failures, the cycle may be the time it takes a resource to run out. Look at what you did from 13.08.2025 to 24.09.2025: releases, switch-overs, new contracts, monitoring set-up.

Forecasts for volume and pace

What we looked at. The distribution of the longest outages and the current rate at which incidents appear.

What we found. 2.7% of incidents last longer than 6 h - about one in 37. At your rate that is about 3 times a month. Outages longer than 6 h among critical and high incidents (priority 1-2): 3.4% of them, about once every 2 months. Tickets still open in the export: 14, of them open longer than 6 h: 13. Their end is not known yet, so they are not in the count of long incidents. Priority 1-2 among them: 2, the oldest open since 22.12.2025. Check whether an outage is going on now or the ticket was not closed. If the current rate holds, in the next 7 days expect around 36 incidents (likely range 22-53), over 30 days - around 156 (128-185). This estimate is built from your own incident rate over the observed period, using the last 60 days: the rate changed over the period.

Incidents a day by week, and the forecast

The rate went from 3.2 to 4.9 a day; a step or a slow rise - the data cannot tell, the shaded band is where a step could be. Next 30 days: about 156 incidents (128-185).

What it means. Such a case happens about once every 2 months. It is your own result, not an industry norm: it is counted from your own incidents.

What to do. Write and rehearse once the order of actions for an outage longer than 6 h: who decides to switch to the reserve, who talks to customers, who talks to the regulator.

Check it yourself.* Ask the engineer on duty what they will do if a service does not come back within 3 h. If the answer starts with "I will call my manager", there is no plan.

Level of this report

What we looked at. The completeness and consistency of the journal itself.

What we found. Data quality grade - B (out of A, B, C, D). Field completeness:

What it means. The fields time registered, time closed, service or node, priority, description or cause are filled well - so everything we said based on them is reliable. Records with no closing time were excluded from the duration calculation, not given an average value.

What to do. Make the cause field mandatory at closing, but with a short list of values instead of free text. Free text is filled badly, a list of seven items is filled well. In a quarter you will get the share of incidents with a known root cause - a measure there is nothing to compute today.

Check it yourself.* Open your ticket system and look at which fields are mandatory at closing. If the cause is not mandatory, that is why it is not in this report.

FieldFilled
Time registered100%
Time closed99%
Service or node100%
Priority100%
Description or cause100%

Possible measures

30 days - what gives a result at once

1. Prepare for the ERP month-end incident window
Before 2026-01-29..2026-01-31, keep the ERP support team on standby and run a pre-window health check of month-end jobs on ERP-01, freezing any change on those days.
2. Identify what changed in the September rate shift
Ask service managers to compare the change log and release calendar between 2025-08-13 and 2025-09-24 with the jump in incident volume, and reverse or tune the change that matches.
3. Investigate the store network cascade
Have the network team trace the 14 Store network links to Store Wi-Fi paired tickets and run a root-cause analysis to confirm whether one link incident caused the later sales-channel tickets, starting with the earliest case.

The findings this plan is built on

R1 - ERP has incidents at every month-end
Month-end is the busiest trading and accounting period, so a repeated ERP incident hits the closing process.
(2025-03-15 22:10:00)
services: ERP
ERP recorded incidents in the last three days of the month in 11 of the past 12 months, with 22 incident days against a usual ~7. If nothing changes, the next expected window is 2026-01-29..2026-01-31. This is likely tied to month-end closing jobs or the end-of-month change window rather than an occasional hardware problem.
How to verify: Count ERP incidents by day of month for the past year in your ticket tool.
R2 - Incident rate changed upward in early September
This is a sustained increase, not a short spike, so teams are carrying about half as many more incidents each day.
The weekly incident volume changed around 2025-09-01, between 2025-08-13 and 2025-09-24: the rate moved from about 3.23 to 4.91 incidents per day (+52.2%) and the higher level has held for 122 days. The records cannot tell whether this was a sharp step or a slow build-up; it is likely a change in operations or environment around that window.
How to verify: Compare weekly incident counts before and after 2025-09-01 in your ticket tool.
R3 - Store network incidents cascade into sales services
A single link problem is seen as multiple tickets across the checkout path and sales-floor connectivity, obscuring the real scope.
services: Store network links, Store Wi-Fi, Store tills, Payment gateway
Store network links → Store Wi-Fi
When store network links opened an incident, a ticket for Store Wi-Fi followed within 15 minutes in 14 of 71 cases where chance gives about 0.2, Store tills in 11 of 71 cases, and the payment gateway in 9 of 71 cases. The link side always came first; the reverse direction appeared only 3 times for tills. This is consistent with a shared connection or one incident logged by several teams, not a proven single cause.
How to verify: Review the paired Store network links and Store Wi-Fi tickets and compare closing notes.
R4 - A small share of incidents run longer than six hours
A handful of long open tickets keep consuming capacity, including two high or critical items.
(2025-03-15 22:10:00) (2025-11-26 06:05:00) (2025-07-08 14:42:00)
services: ERP, Warehouse system, Online shop
About 2.7% of single incidents last longer than 6 hours, about 3 such cases a month, and about 23.1% last longer than 2 hours, about 26 cases a month. These long cases are not more frequent than ordinary operations would produce. Right now, 13 of 14 tickets without an end have been open longer than 6 hours, including 2 high or critical tickets, the oldest since 2025-12-22.
How to verify: List tickets open longer than six hours and check their start and resolution.
R5 - Warehouse system has incidents at 02:00 on Thursdays
A predictable weekly incident window disrupts the night shift and can be prevented or staffed in advance.
services: Warehouse system
On Thursday between 02:00 and 03:00, 34 incidents were recorded against a usual ~4, with this pattern present in 31 of 52 weeks. The Warehouse system accounts for 28 of those 34 incidents. This is likely a scheduled night job or maintenance window that repeatedly disturbs the warehouse system.
How to verify: Pull Warehouse system tickets for Thursdays 02:00-03:00 over the past two months.

Appendix A. Methods used

We left method names out of the main text on purpose, to keep it light.

What in the reportHow it was computed
Total downtimeUnion of overlapping intervals before summing - without it parallel incidents are counted twice
Time to restore a serviceMedian, mean and 95th percentile of durations. ITIL 4 calls this measure MTRS (mean time to restore service); in the reliability vocabulary IEC 60050-192 the same number is called MTTR
Time between failuresFor the whole company - the period divided by the number of incidents: how often something breaks. For a service - the period divided by its incidents: this is the time between failures (ITIL 4: MTBF). We count service failures, not hardware reliability
Downtime concentrationSumming downtime by object. Separately, the Gini coefficient: how differently the incidents themselves cost. Yours is 0.53: 0 means every incident costs the same, above 0.6 means a few incidents give most of the downtime
Returns to the same objectThe same object of the same system within 7 days after the fix; tickets of one outage are one failure; compared with chance where the objects of a system fail alike
Share of service requestsBy the ticket type column where it exists and is filled; otherwise by normal service phrases in the description (grant access, reset a password, create an account) with no words about a breakdown next to them. The records found are taken out of every calculation; what is found by words is an estimate from below
Links between systemsA ticket of another system within 15 minutes after a failure, against that system's own rhythm (level over the weeks around, weekday, hour); corrected for the number of pairs tested (247); a link must hold without its two busiest days; days of common storms left out
Rate change on 01.09.2025Weekly means before and after every possible date, corrected for the week-to-week spread; signal strength 5.4 against a threshold of 3.3 (ordinary years cross it in about 3-5% of cases)
Probability of a long outageGeneralized Pareto distribution over threshold exceedances (Peaks-Over-Threshold). Shape parameter ξ = +0.12, 95% confidence interval [-0.07, +0.27]; 137 exceedances of the 3.4 h threshold
Rate forecastMean rate of the current regime, no trend carried forward; the prediction interval follows the spread of your weeks and months (Poisson or, when failures come in waves, negative binomial)
Data quality gradeCompleteness of mandatory fields, consistency of times, share of excluded records

The numbers in this report were checked by a second independent pass of a different model; no divergence was found.

Lorenz Curve - Downtime Concentration

The curve pulls clearly away from the straight line - confirmation that incidents cost very differently: 41.5% of them hold 80% of all downtime (Gini=0.534). A straight line would mean every incident costs the same.

Brief Glossary of Methods Used

p-valueProbability of observing results at least as extreme as measured, assuming no real effect. p < 0.05 is statistically significant (< 5 % chance it is random).
95% CI95 % Confidence Interval - the range that contains the true value in 95 out of 100 repeated experiments. Wider CI = more uncertainty.
Gini coefficientMeasures how unevenly downtime is distributed. 0 = all incidents equal; 1 = one incident caused all downtime. Gini > 0.6 = strong concentration.

Appendix B. Terms

The wording follows the official ITIL 4 glossary.

TermMeaning
EventAny change of state that has significance for the management of a service
IncidentAn unplanned interruption to a service or reduction in the quality of a service
ProblemA cause, or potential cause, of one or more incidents
Known errorA problem that has been analysed but has not been resolved
WorkaroundA solution that reduces or removes the impact of an incident while a full resolution is not yet available
ChangeThe addition, modification, or removal of anything that could affect services
Service requestA request from a user for an action agreed as a normal part of service delivery. Not an incident
Time to restore a serviceHow quickly a service is restored after a failure

* These tasks can be done more precisely and to a higher standard with our help.