What to do after the analysis: from finding to practice
Almost every report ends with the question "what do we do with this". There is no universal recipe, but each finding has a typical cause and an ITIL 4 practice that deals with it. Below is a table of all the findings our report shows: what usually stands behind them, which practice is responsible and what the first step should be. Under the table - where to start if there are several findings, and how to check that it helped. Practice names are as in the article "Practices the business can see".
| What the analysis showed | What usually stands behind it | ITIL 4 practice | First step |
|---|---|---|---|
| How much of the log is requests, not outages requests for access, passwords, software installs in the same queue as outages how we find it → | Requests and outages are logged as one ticket type, so downtime and failures look bigger than they are | Service request management; incident management | Separate ticket types in the system and recount downtime without requests. It is better to start here: all other numbers become more honest |
| Concentration of downtime a small share of incidents gives most of the downtime how we find it → | A few chronic defects that are worked around every time instead of being removed | Problem management | Open a problem record for each of the top incidents - with an owner and a deadline |
| Fails rarely but takes long to fix high mean time between failures (MTBF), long mean time to recover (MTTR) how we find it → | Time goes not to the fix but to finding the owner, access and a decision | Incident management, escalation | Break down the five longest incidents by stage: detection, assignment, fix. More - the escalation matrix |
| Fails often but is fixed fast low mean time between failures with short recovery how we find it → | The same defect that the team has learned to work around quickly | Problem management | Find the common cause of the frequent outages and get it removed, instead of speeding up the workaround |
| A site breaks again and again one till, one node, one site - more often than its neighbours how we find it → | Repairs by the standard without removing the cause; wear of specific hardware | Problem management; IT asset management | A problem record for the site and a decision: fix the cause or replace |
| Some outages pull others along after one outage, tickets for other systems follow how we find it → | One cause creates several tickets, and different teams fix "different" outages | Monitoring and event management (linking events); problem management | Merge such tickets into one incident; describe the dependencies of the key service |
| It got worse after a change outage frequency rose from a certain date, and it is not chance how we find it → | Changes go without risk assessment and without watching afterwards | Change enablement | Introduce a watch period after changes and a rollback condition written in advance |
| Outages pile up at certain hours and days peaks at the same hours, weekdays, dates of the month how we find it → | Scheduled jobs, backups, reporting periods compete for resources; work at a bad time | Capacity and performance management; change enablement | Put the schedule of background jobs and work on top of the peak map and spread them in time |
| An outage comes back in a cycle every few days, at month end - with the date of the next one how we find it → | A scheduled job, something that builds up (logs, memory, space), a reporting period | Problem management; capacity and performance management | Find what runs or builds up with this period and check it before the next expected date |
| A metric that warns of an outage a measure behaves unusually hours before an outage how we find it → | The signal was there, but there is no alert on it | Monitoring and event management | Set an alert on the metric found, with the lead time found, and a month later count how many were false |
| Forecast: more outages or a growing risk of long downtime the expected number of outages and the probability of a long failure how we find it → | Load grows, hardware ages, the old safety margin is running out | Service level management; service continuity management; capacity management | Discuss the target level with the business and put the forecast into the on-call, purchase and recovery drill plans |
Where to start if there are several findings
Almost every report has several findings, and there is no strength for everything at once. The order below is ours, from experience: each step makes the next one more precise or cheaper.
- First clean the data. If the log has many requests, all other numbers are inflated. Separate requests from outages - and recount.
- Then changes. The cheapest step with the fastest return: according to Google, about 70% of outages are caused by changes to a live system. A watch period and a rollback condition need no new systems and no people.
- Then links, and only after that concentration and repeats. The worst system by downtime may be the victim of someone else's failure. First check whether its outage is pulled along by another one, otherwise the problem record will be opened for the wrong system. More - why one method is not enough.
- Repeats and concentration - to problem management. Not all at once, but the top three. A practice started wide usually dies in the first quarter.
- Long recovery - to escalation. Who learns, how fast, who is next by the timer. A visible result in weeks.
- Hours, cycles, early warnings, forecast - as you can. This is work with resources, schedules and monitoring. It is cheaper when the first steps are done.
How to check that it helped
Take the starting level before the work begins and repeat the measurement a quarter later in the same way - best with the same report on a fresh export. Without a starting level any result stays a claim, and the first sceptical question in a meeting will wipe it out.
- Problem management: the share of repeated incidents and the share of downtime given by the top rows of the list.
- Change enablement: outage frequency after changes and the change failure rate - changes that needed a rollback or an urgent fix.
- Escalation: time to start of work, separately from fix time.
- Monitoring: the share of outages learned about before the client.
Related articles: ranking of outage types - how to label your incidents by source; what ideal operations look like - where to move in the end.
The table is not dogma: the same finding is covered by different practices in different companies. It sets the direction of search, and the exact answer comes from your data and a talk with the team.
We usually go through the report together with the client's team: choose one or two goals, build a plan for the quarter and a quarter later check the result with a repeat run. If this is your case - .