How much of the log is requests, not outages
The head of IT brings the annual report to a meeting: almost 1 500 incidents and more than 4 000 hours of downtime. The commercial director is horrified. The engineers are puzzled: "What four thousand hours? A cash register is back in 20 minutes". Sounds familiar?
Part of these "incidents" are requests to give access, reset a password, install a program. They live in the same queue, take days to complete and push all the figures up. Before counting outages, you need to remove from the log what is not an outage.
The task of the method
Before any calculation, the system answers three questions:
- Which part of the log is not outages? Service requests, changes, planned work.
- How much do they distort the figures? The number of outages, the recovery time, the total downtime.
- How to count everything else honestly? All sections of the report are then calculated without them.
How it is calculated
- 1. The ticket type column
If the export has a column that says "Incident", "Service request", "Change", the system takes it. It recognises the column by its content, not by its name: for some it is "Type", for others "Category" or "Issue Type".
- 2. Words in the description
There is no column or the type is not filled in - the system reads the description: "reset password", "give access", "new employee", "install a program". This is an estimate from below: a request written in other words stays among the outages.
- 3. Words about a breakdown matter more
"Password reset does not work" or "access to the site does not work" is a failure, even with the words of a request. If the text has "does not work", "error", "unavailable", the record stays an outage. To throw a real failure out of the report is worse than to leave one request in it.
- 4. Remove from all calculations
The records that were found are excluded before any counting: recovery time, frequency, Pareto, forecast are calculated only on outages. The first section of the report says how many records were removed and by what sign, with examples - so that you can check.
What happens in real exports
Logs are written in different ways, and the method is ready for this:
- The type is called "Category" or "Issue Type"It is found by the values. The service column is not mixed up with the type.
- The type is not filled in everywhereWhere it is empty, the words in the description decide, and the report calls this part an estimate.
- Changes, problems, planned workAlso not outages: they are removed together with requests and named separately.
- A request without recognisable words"Moving a workplace to another office" stays among the outages. So the share by words is always "at least".
- An outage with the words of a request"VPN access does not work for new employees" stays an outage.
- Three languagesRussian, English and Serbian logs are read in the same way.
Checked on five logs of a retail chain for a year in three languages: each has about 1 450 records and 320 mixed-in requests. By the type column, all requests, changes and problems are removed, down to the last one. By words - all requests with recognisable words; about one request in seven, without such words, was left. Not one of 36 outages with the words "access", "password", "install" was removed. After cleaning, the number of outages and the recovery time match a manual count on the table without requests.
What to do with the result
| The picture in the log | Possible causes | How to check | What to do |
|---|---|---|---|
| Many requests (for example, every fifth ticket is a request for access) | One queue for everything, no separate ticket type | Mark the last 20 tickets by hand | Create the type "Service request" and count outage figures only on incidents |
| The type exists, but is empty for some tickets | The field is optional and is filled in by mood | Filter the export by an empty type | Make the field required when a ticket is closed |
| Requests hang for days | They wait for approval, a purchase, an employee starting work | Compare the times of requests and outages in the export | Different terms (SLA) for requests and outages; do not include requests in the report on failures |
| The share by words is small, but the feeling of "we are buried in requests" is there | Requests are written in own words | Compare the examples from the report with 20 of your own tickets | Create a ticket type: then the split is exact, not an estimate |
Example. "We are drowning in outages"
Input: a retail chain, 47 shops, about 1 470 records in the log for a year. The ticket type is in the "Category" column, but for some requests it is not filled in.
Situation: management thinks that IT is "drowning in outages": by the annual report, almost 1 500 incidents and more than 4 000 hours of downtime.
What the report showed:
- 320 records (22%) are service requests, 25 more are planned work. 47 requests without a type were found by words in the description.
- Without them, outages for the year are about 1 100, not almost 1 500.
- Half of the outages are fixed faster than 50 minutes, not 72.
- The total time of tickets is about 1 500 hours instead of 4 300: almost two thirds of the "downtime" came from requests that were completed over days.
The value of such a result:
- An honest report to management. A quarter fewer failures, almost three times less downtime. The talk about the budget is about real problems, not an inflated figure.
- The right goals for engineers. The on-call staff are judged on the recovery time of outages, not on the time to issue a laptop.
- Staff and on-call duty. You can see how much of the first line's work is planned. It can be moved to the day shift or automated: password reset, issuing standard access.
- Accuracy of the other sections. Pareto, forecast, outage hours are calculated on real failures.
The grey bar is the figure for all tickets in a row, the dark one is for outages only. On the right - by how many percent the figure was inflated.