Retail chain · report · 2026-10-01 10:11 UTC
Confidential
This report is built only on these sources:
| Source | What it held |
|---|---|
| Incident journal | An export from your ticket system for the period 364 days, 1728 records. Fields: time registered, time closed, service or node, priority, description or cause. |
| List of critical services | As named by you: Store tills, Payment gateway, Online shop. |
| Metric exports | Not uploaded. We need CPU, memory and link load, packet transfer delay, packet loss ratio, delay variation, backup success and certificate expiry - a measurement time and a value column for each one. (see the check map below) |
| Change log | Not uploaded. We need the date and time of each change, its type (standard or emergency), the service it touched, and the outcome. (see the check map below) |
| First-line support data | Not uploaded. We need the resolving group or support line for each ticket, an escalation flag, and a reopen flag. (see the check map below) |
Across 364.2 days, incidents produced 1668.40 hours of elapsed downtime. Typical restoration took 54 minutes, but a few long incidents absorb much of the recovery effort.
Incident volume has been higher since early September, moving from about 3.23 to 4.91 incidents per day. ERP has incidents in the last three days of almost every month. The next expected window is 2026-01-29. 2026-01-31 if nothing changes.
The map holds every kind of check we offer. What fills it depends on the measures you can still add.
Legend: measuredneeds your field or your targetneeds another data sourceavailable in the extended review
| Does the infrastructure keep the promised level - 1/4 | |
Total downtimeDetailsWhat this is. Total time your services were unavailable over the log period. How we count it. We add up the duration of all incidents, but merge overlapping intervals: two incidents running at the same time from 10:00 to 11:00 count as one hour of downtime, not two. Without merging, the number is inflated more the larger the infrastructure is. Where the definition comes from. Our calculation. The concept of downtime comes from availability in the ITIL 4 glossary: the ability of a service to perform its agreed function when required. What we need from you. Already calculated - see the number in the row. If the row is amber: we need a field with the incident closing time, otherwise the duration cannot be computed. |
69d 12h 24m |
Availability of critical servicesDetailsWhat this is. The share of observation time the service was up. This is the number people call "three nines". How we count it. Uptime / (Uptime + Downtime), calculated separately for each service you named. An incident that hit three services counts in full for each of the three: from that one service's point of view, it was down the whole time. Where the definition comes from. Formula - Google SRE Book, ch. 3 "Embracing Risk", verbatim. The concept of availability - ITIL 4 glossary. What we need from you. Mark the critical services in the list and state the target level (e.g. 99.9%). We calculate right after you answer - no need to re-upload the report. |
choose your critical services and your target level |
Error budget spentDetailsWhat this is. A target like 99.9% allows a certain amount of downtime by definition. That allowed time is the budget. This indicator shows how much of the budget you have already spent. How we count it. Allowed downtime = period length x (1 - target). Spent = actual downtime / allowed downtime x 100%. We use the actual period - exactly as many days as are in your log, with no rounding to a month. A value above 100% means the budget is used up and the target was missed for the period. Where the definition comes from. Google SRE Book, ch. 3: the budget is the gap between the target and the actual uptime - the remaining "unreliability" for the period. What we need from you. One number - the target availability. Without a target the indicator does not exist: the budget is counted from a promise, not from the data. |
set your target - without it the measure does not exist |
Actual recovery time against the targetDetailsWhat this is. RTO (Recovery Time Objective) is the deadline for a service to come back after an outage. This indicator shows how often you met that deadline. How we count it. The share of incidents no longer than your target, plus the worst case for the period - the longest incident. The worst case matters more than the average: a continuity plan is tested on a bad day, not on a normal one. Where the definition comes from. ISO 22300:2021 (vocabulary for ISO 22301): RTO is "the period of time following an incident within which a product, service or activity is resumed, or resources are recovered". The standard is paid, so we have not verified the clause number and do not cite one. What we need from you. Target recovery time in hours. If it differs by service, give the target for the most critical one. |
set your target - without it the measure does not exist |
| How fast you repair - 1/6 | |
Time to restore a serviceDetailsWhat this is. How long it takes from an incident being logged to being closed. How we count it. We give three numbers at once: median, mean, and P95. The median is a normal day: half the incidents are fixed faster. The mean is pulled up by rare, long cases. P95 is the "tail": only one incident in twenty takes longer. A single average without the tail hides exactly the cases that make the business call. Where the definition comes from. ITIL 4 calls this MTRS, mean time to restore service: "a measure of how quickly a service is restored after a failure". The reliability vocabulary IEC 60050-192 calls the same number MTTR (term 192-07-23). What we need from you. Already calculated. If the row is amber: we need a field with the incident closing time. |
54 min / 1h 28m / 4h 41m - median / mean / P95 |
Time to detectDetailsWhat this is. How long an incident ran before you knew about it. This is a blind spot: the service is already down, but there is no ticket yet. How we count it. Actual incident start minus the time it was logged in the system. This cannot be derived from the logging time alone under any assumption: the log knows when the problem was reported, not when it began. Where the definition comes from. No standard exists. MTTD is an established incident-management practice, and we say so plainly instead of dressing practice up as a norm. What we need from you. A target will not fix this - we need an export with a separate "actual start" field (systems name it differently: start time, occurred at, time of occurrence). Upload such an export via the same link, and the row will be calculated. |
needs the field "real start time, separate from the time of registration" filled in your database |
Time to respondDetailsWhat this is. How long a ticket waited for someone to pick it up. This shows how the queue and the on-call shift are performing, not how hard the fault was. How we count it. Time work started minus the time the ticket was logged. Where the definition comes from. No standard exists. MTTA is a service-desk practice. What we need from you. We need an export with a "work started" field (picked up, assigned, acknowledged - whatever your system calls it). |
needs the field "time accepted by an engineer" filled in your database |
Share closed on timeDetailsWhat this is. What share of incidents met the deadlines agreed with the business. How we count it. Calculated separately per priority: how many incidents of that priority closed within your target. Priorities cannot be combined into one number - their targets differ, and a combined figure would hide overruns on the critical ones. We only calculate this when priority is filled in for at least 80% of records; below that threshold the number reflects log data quality, not service performance. Where the definition comes from. ITIL 4: service level - metrics of expected or achieved quality; SLA - a documented agreement on required services and the expected level of service. What we need from you. Target times per priority, in minutes (e.g. critical - 60, high - 240, normal - 1440). If priority is not tracked in the log, or is filled in for less than 80% of records: we first need an export with this field - a target has nothing to apply to without priority. |
set your target - without it the measure does not exist |
Share solved the first timeDetailsWhat this is. What share of incidents did not reopen after being closed. A direct measure of repair quality: was the ticket closed, or was the problem solved. How we count it. The share of incidents with no reopening marker. Where the definition comes from. No standard exists - this is service-desk practice. What we need from you. We need an export with a "reopened" field, or a status-change history - from the history we can reconstruct the fact of reopening ourselves. |
needs the field "reopened" filled in your database |
Share solved on the first lineDetailsWhat this is. What share of incidents were closed without escalating to specialists. Shows how much work moves upward and how loaded the expensive hands are. How we count it. The share of incidents closed by the same group that took them in. Where the definition comes from. No standard exists. ITIL 4 defines only escalation and service desk; the metric itself is not in the glossary. What we need from you. We need an export with a "resolver group" or "support line" field. |
needs the field "resolver group" filled in your database |
| How often and what exactly breaks - 4/8 | |
How often incidents happen and how often a service failsDetailsWhat this is. How often something breaks in the whole company and how often each of the most frequent services fails. How we count it. The period divided by the number of incidents: for the company - the overall rate, for a service - the time between failures. We count service failures, not hardware reliability. Where the definition comes from. ITIL 4, MTBF: "a measure of how frequently a service or other configuration item fails". IEC 60050-192, term 192-05-13, mean operating time between failures. What we need from you. Already calculated. |
1369 incidents, on average one every 6h 23m. Most often: Store tills - once every 21h 29m; Online shop - once every 2d 05h 38m; Mobile app - once every 2d 23h 04m |
Downtime concentration by serviceDetailsWhat this is. Shows whether your downtime is concentrated in a few points or spread evenly. If concentrated, work on a short list gives a disproportionately large effect. How we count it. We rank services by their contribution to downtime and check how much of it the top ones account for. We also calculate an inequality score between 0 and 1 (named in Appendix A for readers who want the formula): 0 means "all services contribute equally", closer to 1 means "almost everything falls on a few". Where the definition comes from. Our calculation. No borrowed norms exist: we compare your infrastructure with itself, not with someone else's average. What we need from you. Already calculated. If the row is amber: we need a "service or node" field - without it downtime cannot be attributed. If grey-blue: this one comes with the extended review. |
3 of 13 services carry 53% of all downtime |
Returns to the same objectDetailsWhat this is. What share of failures came back on the same object - server, CI, store - within a week after the fix, and whether that is more than chance gives. How we count it. The key is the object together with its system: "the till in store 12", not "tills". Tickets of one outage on an object (opened while the previous one is open, or within an hour after it closed) count as one failure. A return is a failure within 7 days after the end of the previous one. Next to it: the share chance would give if all objects of a system failed alike. An object is named when it fails more often than the other objects of its system, or its failures come in series more often than the others' rhythm gives; both tests are corrected for the number of objects. The same object is not yet the same cause: in ITIL 4 a repeat points to a problem, "a cause, or potential cause, of one or more incidents". Where the definition comes from. Our calculation; interpretation - ITIL 4, problem. What we need from you. Already calculated. If the row is amber: we need an object column - server, CI or store - filled in at least one ticket of five. If grey-blue: available in the extended review. |
34% of failures came back on the same object within 7 days after the fix; if objects failed alike, about 33% |
A metric before failuresDetailsWhat this is. Which monitoring metric - CPU load, memory, traffic, disk - moves away from usual before failures, and about how long before them. How we count it. We put the journal and the metrics export on one time axis. "Usual" for a metric is its level at the same hour of the day, workdays and weekends apart. We look at windows 15 min - 1 h, 1-6 h and 6-24 h before the ticket and compare how often the metric was above or below usual before failures and at the same hours on other days. A finding needs this to happen clearly more often before failures, on different days, corrected for the number of series. A ticket opened by an alert on the same metric is not a precursor. Where the definition comes from. Our calculation; tested on artificial months and years of monitoring with a known answer. What we need from you. Already calculated. If the row is grey: send a monitoring metrics export for the same period together with the journal. If amber: the common period or its failures are too few - it needs 14 days and 8 failures or more. |
needs another data source: a monitoring metrics export for the same period (Zabbix, Grafana, Prometheus) |
Breakdown by priorityDetailsWhat this is. How incidents are distributed across importance classes. How we count it. We calculate the share of each priority, but only if the field is filled in for at least 80% of records. Below that, the breakdown reflects log-filling discipline, not the real picture of breakages, and we do not print it. Look at two things. First, the share of the top classes: if critical is more than a quarter of the flow, priority is set by how loud the reporter is, not by business impact, and the urgent lane stops being urgent. Second, compare it with "how fast you repair": critical incidents repaired no faster than normal ones mean the priority exists in the log but not in the shift's work. Where the definition comes from. ITIL 4, incident: an unplanned interruption to a service or a reduction in its quality. The standard does not define priority scales - we use yours. What we need from you. Already calculated. If the row is amber: we need an export with a "priority" field, or for it to be filled in. |
P3 - MEDIUM - 667 (48%); P4 - LOW - 393 (28%); P2 - HIGH - 270 (20%); P1 - CRITICAL - 53 (4%) |
Share of incidents with a known root causeDetailsWhat this is. What share of incidents were traced to a root cause, rather than just closed after the service was restored. How we count it. The share of records with a root cause filled in, or linked to a problem record. Where the definition comes from. ITIL 4: problem - the cause, or potential cause, of incidents; known error - a problem that has been analysed but not resolved (a workaround is not required for this). What we need from you. This field is not tracked in your log. That fact alone is already a finding: repeat incidents are closed without root-cause analysis, so they will return. To see a number, we need an export with a "root cause" field, or a link between incidents and problems. |
the field is not kept - so repeat incidents are closed without anyone finding the cause |
Share of emergency changesDetailsWhat this is. What share of changes are made in "needed it yesterday" mode, without the usual approval. A high share means the planned process cannot keep up with reality. How we count it. The share of emergency changes out of all changes in the period. Where the definition comes from. ITIL 4: emergency change - a change that "must be implemented as soon as possible"; standard change - a pre-authorised, low-risk change. What we need from you. A field in the incident log will not fix this: the denominator here is changes, not incidents. We need an export from the change log for the same period. It will also refine the "incidents with a change trace" row - currently that row is a lower-bound estimate from description text. |
needs another data source: a change log |
Share of incidents caused by suppliersDetailsWhat this is. What share of breakages came from outside your perimeter - a provider, vendor, or contractor. This is a number for a conversation with the supplier, not with your team. How we count it. The share of incidents with an external party named as at fault. Where the definition comes from. ITIL 4, supplier management practice. What we need from you. We need an export with an "at-fault party" or "supplier" field. |
needs the field "party at fault" filled in your database |
| Capacity and health | |
| This is where the state of your resources belongs: load, network delay and loss, backup success, certificate expiry, out-of-support components. An incident journal does not hold them - we need a separate metrics export for the same period: the time of each measurement and one value column per measure. | |
In total: 6 measures computed, 10 wait for your decision or a filled field, 8 need other data.
Every line marked "needs your field or your target" is not our limit - it is the name of a field in your database. Send us the field "time accepted by an engineer" and you will get response time separately from repair time. Then you will see where you lose more: while the dispatcher looks for a free engineer, or while the engineer repairs.
What we did not get
Below is what would close the most rows of the map. Answering commits you to nothing.
We got no metrics export. Do you collect resource load, loss, latency or response time?
The records name no cause. Do you keep repeating causes separately, as problems?
There was no change log in what you sent. Do you keep one, with a date and time for every change?
Step 2 of 5 - Recording
Outages are recorded. That history already shows how long recovery takes, where downtime piles up and what outages to expect next.
The step is counted from the log only. Functions the log does not show are not counted until you tell us about them, so this is a lower bound. Confirmed by data: 1 of 1 key functions up to step 2.
| Step | Function | ITIL 4 practice | Status |
|---|---|---|---|
| 2 | Every outage is recorded | Incident management | confirmed by data |
| 3 | Main services have owners | Service level management | confirmed by data |
| 3 | Priority is set by impact and urgency | Incident management | confirmed by data |
| 3 | Recovery times are promised | Service level management | not seen in the data |
| 3 | An on-call rota and a hand-over order | Service desk | not seen in the data |
| 4 | Repeats are investigated down to the cause | Problem management | not seen in the data |
| 4 | Changes are recorded with a type and a result | Change enablement | not seen in the data |
| 4 | Outages are known before the consumer notices | Monitoring and event management | not seen in the data |
| 5 | Improvements have an owner and a measured effect | Continual improvement | not seen in the data |
| 5 | Measures are reviewed with the consumer | Measurement and reporting | not seen in the data |
| 5 | Service dependencies on equipment are known | Service configuration management | not seen in the data |
The assessment scheme follows formal process assessment (ISO/IEC 33000, the CMMI appraisal method): what people say and what the work records show. An answer makes a function "declared", the export makes it "confirmed", a mismatch makes it "disputed".
A function is confirmed by data when the log covers at least 90 days and 30 records, the start, end and object of an outage are filled in 95% of records, the function's own field (service, priority) in 80%, and no single value stands on more than 95% of records.
Practices and their names come from ITIL 4. The ladder rests on ITIL 4, CMMI and ISO/IEC 33000; it is not a certification and not an assessment by an accredited assessor.
What we looked at. 1728 journal records over 364 days.
What we found. Out of 1728 records, 320 (18%) are service requests by the column “Issue Type”, not failures; 45 of these requests have no type, we recognised them by the words of the description, so this part is an estimate; another 25 are changes, problems or planned works. We took all these records out of every calculation in this report: below we work with 1369 incidents, which is 3.8 per day. In total, services were down or worked worse than normal for 69d 12h 24m. This is the sum over all objects, not the time when everything was down at once.
A bar is how many log records of each kind. We took the orange ones out of every calculation in the report: they are not failures, and their completion time does not go into the time to restore.
Log rows we took out
| Ticket | Start | Type or description |
|---|---|---|
| INC025000 | 2025-01-01 08:54:09 | Service request |
| INC025002 | 2025-01-01 12:51:31 | Service request |
| INC025012 | 2025-01-03 11:41:38 | Service request |
| INC025062 | 2025-01-17 10:02:51 | Unlock my account after holiday |
| INC025076 | 2025-01-20 16:48:40 | Please install the e-signature software |
What it means. Counted all together, there are 25% more failures than really happened, and the repair time mixes with the time to fulfil requests. A management report on all tickets shows more outages than there were. About downtime: if we simply added the durations one after another, we would get 83d 06h 19m. The difference of 13d 17h 54m is incidents that ran at the same time. Counting them twice inflates the downtime.
What to do. Count failure numbers only on records with the incident type in the column “Issue Type”, and make sure the type is always filled.
Check it yourself.* Filter the export by the column “Issue Type”: you should get 300 records typed as a request, change or problem.
What we looked at. Time from registration to restoration for 1369 incidents. This is calendar time: nights and weekends count in full. 14 tickets are not closed: their end is unknown, so they are left out. Priority 1-2 among them: 2, the oldest open since 22.12.2025. If this is an outage going on now, it is larger than any number above.
What we found. Half of the incidents close faster than 54 min. The average is 1h 28m. The worst 5% run longer than 4h 41m; the longest was 19h 30m. Those 69 incidents hold 24% of all downtime.
Half of the incidents are fixed within 54 min, the average is 1h 28m. The 5% longest (right of 4h 41m) hold 24% of downtime.
The 3 largest incidents account for 2.4% of total downtime - most downtime builds up from many ordinary incidents, not from a few outages.
Log rows: the longest incidents
| Ticket | Start | Service | Object | Length |
|---|---|---|---|---|
| INC025311 | 2025-03-15 22:10 | ERP | ERP-01 | 19h 30m |
| INC026498 | 2025-11-26 06:05 | Warehouse system | WMS-01 | 15h 00m |
| INC025788 | 2025-07-08 14:42 | Online shop | WEB-01 | 13h 15m |
| INC025807 | 2025-07-13 13:43 | Store Wi-Fi | Store #33 | 11h 48m |
| INC025315 | 2025-03-16 19:09 | Backup | BACKUP-01 | 10h 20m |
What it means. The average (1h 28m) is above the median (54 min) - repairs almost always look like this: a few long cases pull the average up. But there is no separate class of heavy incidents: the 5% longest hold only 24% of downtime, most of it comes from ordinary repairs. Report the median - it shows a typical incident.
What to do. Create a separate procedure for incidents that are not closed within 4h 41m. Not "tighten control", but exactly this: after that time, escalate to an engineer who is allowed to call the vendor. Today such incidents stay in the common queue.
Check it yourself.* Open the last three incidents longer than 4h 41m and look at how much time went into diagnosis and how much into the repair itself. If more than half went into diagnosis, the bottleneck is in monitoring, not in the hands of your engineers.
What we looked at. How incidents and downtime spread across 13 objects, and repeats.
What we found. 3 objects out of 13 hold 53% of all downtime. The worst is "Store tills": 407 incidents, 19d 14h 50m of downtime. With an even spread it would be 23%. For "Store tills", 77% of the downtime comes from priority 3-4 tickets. That is ticket time, not necessarily business downtime: check whether the service was really down. Tickets closed by a timer or a queue clean-up, and open tickets, are not in the downtime (details in section 2). 34% of failures on objects (435 of 1296) came back on the same object of the same system within 7 days after the fix. If every object of each system failed alike, it would be about 33%. Tickets of one outage on one object count as one failure: 87 such tickets. "ERP-01" in "ERP": 49 failures in the period, a typical object of this system has 22: it fails more often than the others, and chance does not give that. "Store #17" in "Store tills": 29 failures in the period, a typical object of this system has 9: it fails more often than the others, and chance does not give that.
A bar is a share of all downtime; the dashed line is what each would hold if downtime were shared evenly. Orange: the ones this section names.
A row is an object, a mark is a failure. A red dot came within 7 days after the previous fix, a dark tick after a break. On the right: failures on the object against a typical object of the same system, or how many came back.
Log rows: the longest incidents of “Store tills”
| Ticket | Start | Service | Object | Length |
|---|---|---|---|---|
| INC026059 | 2025-09-10 09:52 | Store tills | Store #9 | 8h 08m |
| INC026162 | 2025-09-28 15:39 | Store tills | Store #34 | 7h 58m |
| INC026236 | 2025-10-09 11:08 | Store tills | Store #40 | 7h 33m |
| INC026098 | 2025-09-18 15:41 | Store tills | Store #6 | 5h 38m |
| INC026400 | 2025-11-08 08:34 | Store tills | Store #5 | 4h 55m |
Log rows: failures on “ERP-01”
| Ticket | Start | Service | Object | Length |
|---|---|---|---|---|
| INC026352 | 2025-10-31 16:15 | ERP | ERP-01 | 2h 21m |
| INC026506 | 2025-11-27 10:25 | ERP | ERP-01 | 4h 56m |
| INC026509 | 2025-11-27 16:21 | ERP | ERP-01 | 1h 36m |
| INC026516 | 2025-11-28 12:40 | ERP | ERP-01 | 3h 34m |
| INC026720 | 2025-12-31 13:15 | ERP | ERP-01 | 1h 20m |
What it means. A few objects hold the downtime, and more than chance would give. People and reviews sent there win back the most time. The named objects fail again and again. The same object is not yet the same cause, and the journal does not show the cause. Possible causes: the cause was not removed - it was fixed by a restart or a workaround; the object is worn or overloaded; the object has more equipment or people than its neighbours; one outage is logged again later. What tells them apart: the description and the closing notes of these tickets, a record in the problem log, and how much equipment the object has.
What to do. Start with 3 objects, not 13. For "Store tills", review its incidents together, in one review, and look for a common cause instead of the cause of each one. Review the failures of "ERP-01" together, in one review: compare the descriptions and what was done at closing. One cause, fixed by a restart - open a problem and look for the cause. Different causes, and the object is bigger than its neighbours - compare it with objects of the same size. The same outage logged again - agree to link a repeat ticket to the first one.
Check it yourself.* Open any three incidents from "Store tills" from the last month. If they have different causes, we were wrong and the object is simply overloaded. If the cause is the same, you have a problem that nobody owns. Send the journal again in a quarter: "ERP-01" should have fewer failures and fewer returns within a week.
What we looked at. Which systems fail together: after a failure of one, a ticket for another opens within 15 minutes more often than chance allows.
What we found. After a failure of "Store network links", tickets for other systems open within 15 minutes: "Store Wi-Fi" 14 times out of 71, "Store tills" 11 times out of 71, "Payment gateway" 9 times out of 71. By chance, given these systems' own rhythm, it would happen at most 0.7 times over the whole period. 80 tickets out of 1383 opened in the first 15 minutes of an outage already under way: they continue it, or register the same outage again. The number of failures and the time between them in this report are counted in tickets, so they are overstated by about 6%; in the list of the largest incidents such an outage counts as one.
A bar is how many times a ticket of the second system opened within 15 minutes after the first one failed; the grey tick is what chance gives with that system's own rhythm. The arrow shows which system is first; a double arrow means neither is regularly first.
Log rows: “Store network links” and right after it “Store Wi-Fi”
| Ticket | Start | Service | Object | Length |
|---|---|---|---|---|
| INC026575 | 2025-12-08 06:20 | Store network links | Store #44 | 2h 41m |
| INC026577 | 2025-12-08 06:31 | Store Wi-Fi | Store #44 | 2h 33m |
| INC026706 | 2025-12-29 21:00 | Store network links | Store #30 | 13 min |
| INC026708 | 2025-12-29 21:10 | Store Wi-Fi | Store #30 | 17 min |
What it means. Tickets for "Store Wi-Fi" in the first minutes after a failure of "Store network links" are more likely a consequence than a separate problem. There are three usual reasons: "Store Wi-Fi" works through "Store network links"; both depend on a third thing - power, a shared node, a site; or two teams registered one outage. The object column (store, node, site) tells them apart: if a pair of tickets has the same object, it is one outage.
What to do. Link tickets for "Store Wi-Fi" opened in the first 15 minutes after a failure of "Store network links" to its ticket as child tickets. Then failures stop being counted twice, and people fix one cause. If "Store Wi-Fi" works through "Store network links", "Store network links" needs the backup.
Check it yourself.* Open the last three outages of "Store network links" and the tickets for "Store Wi-Fi" in the next 15 minutes: is it the same object, and did they close together with the outage?
What we looked at. How incidents spread across hours and days of the week, and how the rate changed over time.
What we found. By time of the week: weekdays 9 to 18 hold 46% of incidents (27% of the hours of a week), weekday evenings and nights 28%, weekends 27%. Thu 02:00-03:00: 34 incidents, where your usual rhythm would give about 4. Failures in this hour came back in 31 of 52 weeks. Most of them are Warehouse system (28 of 34). ERP: failures on the last three days of the month - in 11 of 12 months. Failure days on these dates: 22, where the usual rhythm would give about 7. Last time: 30.12.2025. If nothing changes, the next one by the cycle is 29.01.2026-31.01.2026. The incident rate went up: from 3.2 a day at the start of the period to 5.1 at the end. These data cannot tell a step from a slow rise. If it was a step, it came between 13.08.2025 and 24.09.2025, most likely around 01.09.2025: it was 3.2 a day, it became 4.9. If it was a slow rise, there is no single date. The new level has held for 122 days now.
The darker the cell, the more incidents. Weekdays 9 to 18 (dotted frame) hold 46%, evenings, nights and weekends 54%. The red frame is an hour that breaks your usual rhythm; the number is its incidents.
Each circle is a day of the cycle: red when the system failed that day, hollow when it passed quietly. Ticks are failures on other days. The hollow red circle after the end of the export is the next day by the cycle, if nothing changes.
Log rows: incidents in the marked hour
| Ticket | Start | Service | Object | Length |
|---|---|---|---|---|
| INC026389 | 2025-11-06 02:21 | Warehouse system | WMS-01 | 1h 59m |
| INC026467 | 2025-11-20 02:15 | Warehouse system | WMS-01 | 1h 03m |
| INC026505 | 2025-11-27 02:23 | Warehouse system | WMS-01 | 2h 04m |
| INC026633 | 2025-12-18 02:04 | Warehouse system | WMS-01 | 1h 34m |
| INC026634 | 2025-12-18 02:25 | Payment gateway | PAY-01 | 1h 19m |
Log rows: failures of “ERP” on the cycle days
| Ticket | Start | Service | Object | Length |
|---|---|---|---|---|
| INC026174 | 2025-09-29 16:54 | ERP | ERP-01 | 1h 54m |
| INC026338 | 2025-10-29 19:44 | ERP | ERP-03 | 21 min |
| INC026516 | 2025-11-28 12:40 | ERP | ERP-01 | 3h 34m |
| INC026711 | 2025-12-30 09:47 | ERP | ERP-01 | 4h 28m |
| INC026716 | 2025-12-30 15:04 | ERP | ERP-01 | 4h 14m |
What it means. A failure tied to an hour is almost always tied to a schedule: a night job (exchange, export, backup), a release window, planned works of a supplier, or a check that writes what it found at its own time. Random failures do not gather like this. A failure that comes back at equal intervals is almost always tied to a schedule or to something that fills up: a job every few days, month-end closing and reports, disk space or memory that runs out by the same time, a certificate or password that expires. Random failures do not repeat like this. The journal does not say what happened: it has no cause field. Options: more users or sites, a release or a migration, a new contractor, new logging rules - for example, monitoring started to open tickets by itself. The change log for those weeks and a "registered by" column (person or monitoring) will tell them apart.
What to do. Find what runs at this time: the job schedule, the release window, supplier works. If it is a job, move it or add a check before it starts; if it is releases, add a check right after the rollout. Compare these dates with your calendar of month-end closing, reports and payments: what loads the system on these days. Give it resources in advance and put someone on duty who knows the system. Find what changed in those weeks. If it is business growth, recount the on-call shift for the new level. If it is a release or a contractor, return the question to whoever made the change.
Check it yourself.* Open the tickets of Thu 02:00-03:00: if they share one system and similar text, it is one repeating cause. The time is taken as it is in the export. If your system writes time in UTC, shift the hours to your time zone. Open the tickets of ERP on the cycle dates (29.09.2025, 29.10.2025, 28.11.2025, 30.12.2025): if the text is the same, it is one cause. If the system was restarted between failures, the cycle may be the time it takes a resource to run out. Look at what you did from 13.08.2025 to 24.09.2025: releases, switch-overs, new contracts, monitoring set-up.
What we looked at. The distribution of the longest outages and the current rate at which incidents appear.
What we found. 2.7% of incidents last longer than 6 h - about one in 37. At your rate that is about 3 times a month. Outages longer than 6 h among critical and high incidents (priority 1-2): 3.4% of them, about once every 2 months. Tickets still open in the export: 14, of them open longer than 6 h: 13. Their end is not known yet, so they are not in the count of long incidents. Priority 1-2 among them: 2, the oldest open since 22.12.2025. Check whether an outage is going on now or the ticket was not closed. If the current rate holds, in the next 7 days expect around 36 incidents (likely range 22-53), over 30 days - around 156 (128-185). This estimate is built from your own incident rate over the observed period, using the last 60 days: the rate changed over the period.
The rate went from 3.2 to 4.9 a day; a step or a slow rise - the data cannot tell, the shaded band is where a step could be. Next 30 days: about 156 incidents (128-185).
What it means. Such a case happens about once every 2 months. It is your own result, not an industry norm: it is counted from your own incidents.
What to do. Write and rehearse once the order of actions for an outage longer than 6 h: who decides to switch to the reserve, who talks to customers, who talks to the regulator.
Check it yourself.* Ask the engineer on duty what they will do if a service does not come back within 3 h. If the answer starts with "I will call my manager", there is no plan.
What we looked at. The completeness and consistency of the journal itself.
What we found. Data quality grade - B (out of A, B, C, D). Field completeness:
What it means. The fields time registered, time closed, service or node, priority, description or cause are filled well - so everything we said based on them is reliable. Records with no closing time were excluded from the duration calculation, not given an average value.
What to do. Make the cause field mandatory at closing, but with a short list of values instead of free text. Free text is filled badly, a list of seven items is filled well. In a quarter you will get the share of incidents with a known root cause - a measure there is nothing to compute today.
Check it yourself.* Open your ticket system and look at which fields are mandatory at closing. If the cause is not mandatory, that is why it is not in this report.
| Field | Filled |
|---|---|
| Time registered | 100% |
| Time closed | 99% |
| Service or node | 100% |
| Priority | 100% |
| Description or cause | 100% |
We left method names out of the main text on purpose, to keep it light.
| What in the report | How it was computed |
|---|---|
| Total downtime | Union of overlapping intervals before summing - without it parallel incidents are counted twice |
| Time to restore a service | Median, mean and 95th percentile of durations. ITIL 4 calls this measure MTRS (mean time to restore service); in the reliability vocabulary IEC 60050-192 the same number is called MTTR |
| Time between failures | For the whole company - the period divided by the number of incidents: how often something breaks. For a service - the period divided by its incidents: this is the time between failures (ITIL 4: MTBF). We count service failures, not hardware reliability |
| Downtime concentration | Summing downtime by object. Separately, the Gini coefficient: how differently the incidents themselves cost. Yours is 0.53: 0 means every incident costs the same, above 0.6 means a few incidents give most of the downtime |
| Returns to the same object | The same object of the same system within 7 days after the fix; tickets of one outage are one failure; compared with chance where the objects of a system fail alike |
| Share of service requests | By the ticket type column where it exists and is filled; otherwise by normal service phrases in the description (grant access, reset a password, create an account) with no words about a breakdown next to them. The records found are taken out of every calculation; what is found by words is an estimate from below |
| Links between systems | A ticket of another system within 15 minutes after a failure, against that system's own rhythm (level over the weeks around, weekday, hour); corrected for the number of pairs tested (247); a link must hold without its two busiest days; days of common storms left out |
| Rate change on 01.09.2025 | Weekly means before and after every possible date, corrected for the week-to-week spread; signal strength 5.4 against a threshold of 3.3 (ordinary years cross it in about 3-5% of cases) |
| Probability of a long outage | Generalized Pareto distribution over threshold exceedances (Peaks-Over-Threshold). Shape parameter ξ = +0.12, 95% confidence interval [-0.07, +0.27]; 137 exceedances of the 3.4 h threshold |
| Rate forecast | Mean rate of the current regime, no trend carried forward; the prediction interval follows the spread of your weeks and months (Poisson or, when failures come in waves, negative binomial) |
| Data quality grade | Completeness of mandatory fields, consistency of times, share of excluded records |
The numbers in this report were checked by a second independent pass of a different model; no divergence was found.
The curve pulls clearly away from the straight line - confirmation that incidents cost very differently: 41.5% of them hold 80% of all downtime (Gini=0.534). A straight line would mean every incident costs the same.
The wording follows the official ITIL 4 glossary.
| Term | Meaning |
|---|---|
| Event | Any change of state that has significance for the management of a service |
| Incident | An unplanned interruption to a service or reduction in the quality of a service |
| Problem | A cause, or potential cause, of one or more incidents |
| Known error | A problem that has been analysed but has not been resolved |
| Workaround | A solution that reduces or removes the impact of an incident while a full resolution is not yet available |
| Change | The addition, modification, or removal of anything that could affect services |
| Service request | A request from a user for an action agreed as a normal part of service delivery. Not an incident |
| Time to restore a service | How quickly a service is restored after a failure |
* These tasks can be done more precisely and to a higher standard with our help.