IT operations practices the business can see: what to invest in first
Lists of "the most important ITIL practices" usually give management little: they are written in the language of processes, while the manager asks something else - what the next quarter of the team's work will give and where it will show. There are many cases when a practice that is right by the book gave the business nothing visible for years. And the other way round: a boring discipline like watching the system after updates changed the picture in a few weeks.
So the ranking below is built not from the textbook but from the goals a company keeps an operations team for. First - these four goals. Then twelve ITIL 4 practices by return on effort. Then the same list through the client's eyes. At the end - the order of introduction, the measures to track the effect, and typical traps. The order itself is ours, from operations experience: ITIL 4 describes 34 practices but does not rank them.
1. Four goals the business pays operations for
Operations rarely brings new revenue - it protects the revenue you already have. It is worth saying this directly, otherwise every talk with the business starts in defence. There are four goals, and any practice should be checked with one question: which of them does it work for.
Keep the revenue you already have
Downtime of a key service is not a "technical incident" but hours when the company does not sell, does not ship and does not serve, and staff wait. Here the contribution of operations is direct, and it is easiest to show in hours.
Look reliable to the client
The client judges not by an availability percentage but by their own experience: did they notice the outage before you, did they get a clear status, did it happen again a week later. This is where contract renewals and referrals come from.
Get access to the market
Tenders, contracts with service level guarantees, industry audits require evidence: processes are described, logs are kept, time limits are met. Sometimes - a certificate under the international IT service management standard ISO/IEC 20000. Without this, part of the market is closed, however good the technology is.
Free up the team's time
The business notices this goal last, but it decides whether the company has the strength to grow. A team that is always fighting fires does not do projects - it promises them.
A practice that works for none of the four goals may be useful inside the team, but you will not be able to defend a budget for it - and that is normal. The problem starts when such practices are the majority of the plan.
2. Ranking of practices by return
The order is by return on effort, adjusted for speed: a practice that gives the same effect in a month ranks above one that gives it in a year. The ranking is for a typical company with an established infrastructure and no separate process team. For a cloud start-up or an industrial company the order will shift, and the "What it gives the business" column helps you rebuild it for yourself. Practice names follow ITIL 4.
| No. | Practice | What it gives the business | What the client sees | Where to start |
|---|---|---|---|---|
| 1 | Incident managementIncident Management | Cuts downtime. Removes the pauses that are often longer than the fix itself: finding the owner, waiting for access, agreeing a decision | The outage ends fast; the client gets a clear status, not silence | Break down the five longest incidents by stage: detection, assignment, fix |
| 2 | Problem managementProblem Management | Cuts the number of outages. In ITIL this is the practice that finds and removes the causes of repeated outages, not single cases | "The same thing stopped breaking" - the strongest signal of trust | Open a problem record for the three most frequent repeated outages, with an owner and a deadline |
| 3 | Monitoring and event managementMonitoring and Event Management | The system learns about the outage, not the client. It moves the very point from which downtime is counted | The company reports the outage first. This is more visible than half an hour of difference in recovery | Count the share of incidents you learned about from clients, and cover the most frequent ones with monitoring |
| 4 | Change enablementChange Enablement | Closes the biggest controllable source of failures: according to Google, about 70% of outages are caused by changes to a live system | Updates stop being a lottery; work happens in the announced window | Introduce a watch period after updates and a rollback condition written in advance |
| 5 | Service level managementService Level Management | Turns the team's work into a promise that can be checked. Without it, any talk about quality stays an argument of opinions | It is clear what is promised and whether it is kept; the report can be shown to their own management | Describe 3-5 key services in business language and agree a target level for each |
| 6 | Knowledge management and the known error databaseKnowledge Management; the known error database is part of problem management | Removes dependence on "irreplaceable" people: a known outage is fixed by an instruction, not by the memory of one engineer | The speed of response does not depend on who is on shift today | Describe the ten most frequent outages: symptoms, workaround, status of the permanent fix |
| 7 | Capacity and performance managementCapacity and Performance Management | Turns some failures into planned work: space, memory and bandwidth run out predictably | The service does not "slow down at month end"; peak periods pass smoothly | Take three resources that grow towards their limit and calculate the date when it will be reached |
| 8 | Service continuity managementService Continuity Management | Deals with rare events, but they decide the worst case of the year: loss of a site, ransomware, a disaster | Visible only at the moment of a big failure - but then visible in full | Run one drill: restore the key database from a backup and record the actual time |
| 9 | Supplier managementSupplier Management | A noticeable part of downtime comes from outside: links, cloud, payment gateways. You can manage it only with the contract and a procedure | The client does not care whose failure it is. They judge how you handle it | Collect for each supplier the contact, time limits and escalation steps - one page for the on-call engineer |
| 10 | Configuration and IT asset managementService Configuration Management, IT Asset Management | The basis for linking events, impact analysis and planning. By itself almost invisible to the business | Not visible - it shows through the speed of other practices | Do not build a full inventory. Describe dependencies only for key services |
| 11 | Continual improvementContinual Improvement | Keeps what was achieved: without regular review, the results of other practices slowly fade | Stability over a long horizon, not an event | Once a quarter: one data review, three decisions, a check a quarter later |
| 12 | Service request managementService Request Management | Takes routine off the team - access, passwords, typical requests - and frees time for engineering work | Staff notice it more than external clients | Move the three most common typical requests to self-service |
A note without which the ranking is easy to misread: this is the order of return on effort, not of importance. The lower rows are not "unimportant": without configuration records you cannot link events, without continual improvement the upper rows slip back over time. It is only about where to start when resources are short. And they are almost always short.
3. Why the top three are these
The top three rows cover three parts of the same path: an outage happened → it was noticed → it was fixed → it does not repeat. That is why they do not replace each other, and it is better to introduce them together, not one after another.
- Monitoring moves the start. While you learn about outages from clients, everything else is already late. This practice does not speed up the fix - it moves the moment the fix starts.
- Incident management squeezes the middle. In reviews of long incidents, most of the time often goes not to actions but to pauses: finding the owner, waiting for access, finding out who may decide. This is fixed by the way outages are passed on and escalated, not by hiring stronger engineers.
- Problem management removes repeats. The first two practices deal with each case separately, this one with their number. Its return is the slowest, but the longest: a removed cause does not come back next quarter.
From this comes a practical conclusion: if you choose one of the three, look at your data for six months. Many short repeated outages with an acceptable fix time - invest in problem management. Few outages, but each drags on for hours - invest in incident management and escalation.
4. What the client really judges
The operations team measures itself by availability and recovery time. The client measures something else, and the gap between these lists explains the situation "the numbers are green but the client is unhappy". Below is what, in our experience, clients name themselves, and the practice responsible for it.
| What matters to the client | Why exactly this | Which practice is responsible |
|---|---|---|
| Predictability matters more than speed | The client plans their own work. A two-hour outage with a clear time is easier to bear than a forty-minute outage in total uncertainty | Service level management, progress updates |
| Hear it from you, not notice it yourself | If the client found the outage first, trust suffers more than from the duration itself | Monitoring and event management |
| That it does not repeat | One big failure is seen as bad luck. The third identical outage in a month - as a trait of the supplier | Problem management, the known error database |
| One entry point and one owner | Retelling the problem to a third person annoys most, whatever the result. In ITIL this is called a single point of contact | Incident management, the service desk |
| Work in the announced window | Unavailability without a warning is no different from a failure for the client | Change enablement, the work calendar |
| A real time instead of an optimistic one | A missed time costs more than a long one named at once. After the second time the client believes no time at all | Incident management, discipline of updates |
Note the common feature: none of the six rows is about technology. Five of the six are covered by discipline, not money - good news for a company with a limited budget. You can raise the client's opinion noticeably earlier than availability.
5. Order of introduction
Practices are linked, and some of them make no sense to build before others. The order below is made so that each step rests on the previous one, and the business sees changes before its patience runs out.
- Put records in order. One incident log, required fields: start and end time, object, outcome. Without this there is nothing to measure the effect of the other steps with, and in six months the argument about results will be an argument of opinions.
- Clean up alerts and fix escalation. Who learns, how fast, who is next by the timer. A visible result in weeks.
- Introduce a watch period after changes and a rollback condition. The cheapest step with the fastest return: it needs no new systems and no new people.
- Start problem management on the three most frequent outages. Not on all - exactly on three. A practice started wide usually dies in the first quarter.
- Describe key services and agree target levels. Now you have evidence for the promises: the log is kept, escalation works, there are fewer repeats.
- Then - by data. Capacity, continuity, suppliers, configuration records join in the order that the analysis of your own incidents suggests, not a general list.
A common mistake in the order is to start with a full configuration inventory. It looks like a foundation, but takes many months and gives the business not a single visible result in that time. In our experience, operations projects are more often closed not because they are wrong, but because they show nothing for too long.
6. How to measure the effect
The six measures below are clear without training, five of them are counted from the incident log. All are in shares and hours: such a conversation is easier to defend than one about notional money, which the other side can always dispute.
- Share of repeated incidents: what percentage of outages repeat one seen before. A direct measure of problem management.
- Change failure rate: how many updates needed urgent action - a rollback or an urgent fix. This is one of the DORA measures, a common set of software delivery metrics. A measure of change enablement.
- Share of outages found before the client: the main measure of monitoring, and at the same time what the client feels directly.
- Time to start of work versus fix time: two parts of one interval. In DevOps the first part is measured by mean time to acknowledge (MTTA), the whole by mean time to recover (MTTR). The gap shows where to invest: in the way outages are passed on, or in skills.
- Concentration of downtime: what share of systems gives most of the lost time. It shows where effort makes sense at all.
- Share of team time on emergency and routine work: counted not from the log but from time records. Google SRE's guide is to keep it below 50%. This is the measure that explains to the business why projects do not move.
The rule of measurement is simple: take the starting level before the work begins and repeat the measurement a quarter later in the same way. Without a starting level any result stays a claim, and the first sceptical question in a meeting will wipe it out.
7. Five traps
- Introducing a practice fully by the book. ITIL is a set of recommendations, not a mandatory programme. The working minimum is almost always smaller than described, and you should start with it.
- Measuring the team by the number of closed tickets. The number grows fast and means nothing: it rewards splitting tasks and closing them formally instead of working on causes.
- Making closing time the goal. Incidents start to be closed before everything is restored, and then new ones are opened for the same thing. The data gets spoiled, and with it the chance to analyse anything.
- Thinking a practice is introduced when the procedure is written. The sign of introduction is changed behaviour of the on-call engineer at three in the morning, not a signed document.
- Building everything at once. Three practices finished give more than twelve started. This is a common mistake of strong teams: they have the strength to start everything.
A short conclusion. Practices should be chosen not from a list of what is "right", but by two questions: which business goal the result works for, and when it will show. For most companies the top of the list is the same - incidents, problems, monitoring, changes - because that is where most of the controllable downtime lies. But the order inside this four is different for everyone, and it is decided not by opinion but by your own data for the last six months.
Want to understand where to start in your case - send an export of incidents: we will return an analysis: where downtime concentrates, how fast things are fixed and which outages pull others along. A quick way to estimate the level in words - the maturity test in the bot, about 5 minutes. What to do with each finding is covered separately - in "What comes next".
Discuss a plan for the quarter with an operations engineer - .