Which sites break again and again
In shop No. 17 the cash register again does not print the receipt. The engineer restarts it and meets the norm of 20 minutes. Three days later the situation repeats. In the Service Desk reports everything is "green" - tickets are closed on time. But the shop manager sees that the register is down all the time.
Each single repair is done by the rules, but standard logs do not show chronic failures. To move from the subjective feeling "our cash register broke again" to exact decisions, you need numbers: which exact sites keep coming back for repair, how much worse they are than their neighbours, and whether it is just chance.
The task of the method
An analysis of the Service Desk export must answer three questions:
- Which site breaks more often than its neighbours? A specific cash register, one server out of three - whatever stands out from the norm.
- Where does an outage come back right after a fix? A site falls in series 1-3 days after repair.
- Is it a pattern or background? How many repeated outages there would be if all sites in the system broke at the same random rate.
How it is calculated
- 1. Focus on the site, not the system
The analysis looks at "the cash register of shop No. 17", not "all cash registers". If you look at the system as a whole, outages will happen every week just because of their total number.
- 2. Merging duplicates
Tickets for the same site, opened during the current failure or within 1 hour after the repair, count as one outage. This cuts off monitoring "bounce".
- 3. The return metric
It records a breakdown of the same site within 7 days after the previous ticket was closed.
- 4. Comparison with chance
The method mathematically hands out outages to sites at random. Then it compares the fact with the model: "27% of outages came back; with a random spread there would be about 28%".
- 5. Two signs of an anomaly
Breaks more often than others: more outages than its peers, with a correction for the size of the fleet (so as not to call one of 47 shops the "worst" by chance).
Outages come in series: after a repair the site breaks in 1-3 days - much more often than the general rhythm of the system gives.
Important: the method points out the problem node. The exact cause (hardware or software) has to be looked up in the ticket texts.
Working with "dirty" logs in real life
The algorithm is protected from typical problems of Service Desk logs and will not find what is not there
- No column with the site (only the service)The method does not look for repeats, so as not to give false anomalies where there are simply many calls.
- The site is rarely given (less than 20% of tickets)The analysis stops and gives the reason. It is clear which field you must make people fill in.
- Monitoring noise (5 tickets in an hour)It is merged into one incident.
- Different spelling ("shop 17" and "SHOP No. 17")It is recognised as one site.
- A mass outage (everything fell)It gives no false findings: a one-off failure is not written into the "chronicle" of dozens of servers.
- Little data (less than 4 weeks or less than 20 outages)The method draws no conclusions from a sample that is too small.
The reliability was checked on 160 years of data in total. The method found all sites that break 20-25 times against a norm of 5. It catches a series of 5 outages in a row in 28 years out of 40, and a series of 4 outages in 20 out of 40. There were no falsely named sites.
What to do with the sites that were found
| The picture in the log | Possible causes | What to do next |
|---|---|---|
| A site breaks all year (26 times against a norm of 5) | Worn hardware, abnormal load, power problems at the site | Modernise or replace the node. Check the ticket texts for identical symptoms |
| Outages come in series (a return in 1-3 days) | The symptom is treated, not the cause (for example, it is fixed with a plain restart) | Open a Problem ticket, investigate the whole series together, and do not close the single tickets |
| A series, then months of quiet | A bad update that was later rolled back or fixed | Check the change log (Change Log). Fix the right configuration |
| Many returns, but within chance | The problem is in the whole system, not in one specific server | Work with the system as a whole, do not look for a "guilty" site |
Example. One shop's cash register and a backup server
Input: a retail chain, 47 shops, 25 servers, about 1 100 tickets a year.
Situation: engineers keep restarting an old cash register in one shop. The backup server fails from time to time a couple of days after it is brought up.
What the report showed:
- The overall rate of returns is 27% (with a random spread - 28%). There is no general problem with returns in the system.
- Cash registers of shop No. 17: 26 outages a year (a typical shop gives 5). They break 5 times more often than the rest.
- Server BACKUP-02: 12 of 16 outages happened within a week after the previous repair (three clear series).
The value of such a result:
- Targeted repair. Replacing the specific cash register in shop No. 17 pays off through fewer engineer visits.
- Problem Management. The series of BACKUP-02 failures goes to a deep analysis. It is clear that simple restarts only put off the next outage.
- Justifying the budget. A request to buy equipment rests on an ironclad argument: "26 repairs a year against a norm of 5". The finance director understands this without technical details.
- Managing metrics. The share of returns (fact against chance) becomes a transparent KPI. After a quarter you can export the data again and check whether the number of anomalies on these sites has fallen.
A row is a site, a mark is an outage. A red dot is an outage within 7 days after the fix of the previous one, a dark dash is one after a break. On the right - how many outages the site has and how many a typical site of the same system has, or how many outages came back.