What comes next · After the report

What to do after the analysis: from finding to practice

Almost every report ends with the question "what do we do with this". There is no universal recipe, but each finding has a typical cause and an ITIL 4 practice that deals with it. Below is a table of all the findings our report shows: what usually stands behind them, which practice is responsible and what the first step should be. Under the table - where to start if there are several findings, and how to check that it helped. Practice names are as in the article "Practices the business can see".

What the analysis showedWhat usually stands behind itITIL 4 practiceFirst step
How much of the log is requests, not outages
requests for access, passwords, software installs in the same queue as outages
how we find it →
Requests and outages are logged as one ticket type, so downtime and failures look bigger than they are Service request management; incident management Separate ticket types in the system and recount downtime without requests. It is better to start here: all other numbers become more honest
Concentration of downtime
a small share of incidents gives most of the downtime
how we find it →
A few chronic defects that are worked around every time instead of being removed Problem management Open a problem record for each of the top incidents - with an owner and a deadline
Fails rarely but takes long to fix
high mean time between failures (MTBF), long mean time to recover (MTTR)
how we find it →
Time goes not to the fix but to finding the owner, access and a decision Incident management, escalation Break down the five longest incidents by stage: detection, assignment, fix. More - the escalation matrix
Fails often but is fixed fast
low mean time between failures with short recovery
how we find it →
The same defect that the team has learned to work around quickly Problem management Find the common cause of the frequent outages and get it removed, instead of speeding up the workaround
A site breaks again and again
one till, one node, one site - more often than its neighbours
how we find it →
Repairs by the standard without removing the cause; wear of specific hardware Problem management; IT asset management A problem record for the site and a decision: fix the cause or replace
Some outages pull others along
after one outage, tickets for other systems follow
how we find it →
One cause creates several tickets, and different teams fix "different" outages Monitoring and event management (linking events); problem management Merge such tickets into one incident; describe the dependencies of the key service
It got worse after a change
outage frequency rose from a certain date, and it is not chance
how we find it →
Changes go without risk assessment and without watching afterwards Change enablement Introduce a watch period after changes and a rollback condition written in advance
Outages pile up at certain hours and days
peaks at the same hours, weekdays, dates of the month
how we find it →
Scheduled jobs, backups, reporting periods compete for resources; work at a bad time Capacity and performance management; change enablement Put the schedule of background jobs and work on top of the peak map and spread them in time
An outage comes back in a cycle
every few days, at month end - with the date of the next one
how we find it →
A scheduled job, something that builds up (logs, memory, space), a reporting period Problem management; capacity and performance management Find what runs or builds up with this period and check it before the next expected date
A metric that warns of an outage
a measure behaves unusually hours before an outage
how we find it →
The signal was there, but there is no alert on it Monitoring and event management Set an alert on the metric found, with the lead time found, and a month later count how many were false
Forecast: more outages or a growing risk of long downtime
the expected number of outages and the probability of a long failure
how we find it →
Load grows, hardware ages, the old safety margin is running out Service level management; service continuity management; capacity management Discuss the target level with the business and put the forecast into the on-call, purchase and recovery drill plans

Where to start if there are several findings

Almost every report has several findings, and there is no strength for everything at once. The order below is ours, from experience: each step makes the next one more precise or cheaper.

  1. First clean the data. If the log has many requests, all other numbers are inflated. Separate requests from outages - and recount.
  2. Then changes. The cheapest step with the fastest return: according to Google, about 70% of outages are caused by changes to a live system. A watch period and a rollback condition need no new systems and no people.
  3. Then links, and only after that concentration and repeats. The worst system by downtime may be the victim of someone else's failure. First check whether its outage is pulled along by another one, otherwise the problem record will be opened for the wrong system. More - why one method is not enough.
  4. Repeats and concentration - to problem management. Not all at once, but the top three. A practice started wide usually dies in the first quarter.
  5. Long recovery - to escalation. Who learns, how fast, who is next by the timer. A visible result in weeks.
  6. Hours, cycles, early warnings, forecast - as you can. This is work with resources, schedules and monitoring. It is cheaper when the first steps are done.

How to check that it helped

Take the starting level before the work begins and repeat the measurement a quarter later in the same way - best with the same report on a fresh export. Without a starting level any result stays a claim, and the first sceptical question in a meeting will wipe it out.

  • Problem management: the share of repeated incidents and the share of downtime given by the top rows of the list.
  • Change enablement: outage frequency after changes and the change failure rate - changes that needed a rollback or an urgent fix.
  • Escalation: time to start of work, separately from fix time.
  • Monitoring: the share of outages learned about before the client.

Related articles: ranking of outage types - how to label your incidents by source; what ideal operations look like - where to move in the end.

The table is not dogma: the same finding is covered by different practices in different companies. It sets the direction of search, and the exact answer comes from your data and a talk with the team.

We usually go through the report together with the client's team: choose one or two goals, build a plan for the quarter and a quarter later check the result with a repeat run. If this is your case - .