The operations ladder: five steps of maturity
The maturity of operations does not show in how many dashboards are built or which system is installed. It shows in what the team does and records for its service: does it keep an outage log, set priorities, promise times, look for causes. On the same Zabbix one team stands on step two and another on step four.
Skipping a step does not work. With no recorded history there is nothing to count, with no priorities and promised times there is nothing to compare, with no recorded changes there is nothing to link a repeat outage to. So each step below has a sign to recognise yourself by, what hurts there, one move up, and what can already be computed from such a history.
Tap a step to read what it is.
Your step by your answers
1. Reacting
How to recognise yourself: You learn about a failure from the consumer, and no outage history piles up anywhere: it is sorted out in chats and by phone.
What hurts: A talk about reliability has nothing to rest on: there are no facts, only impressions.
The move up: Record every outage in one place: start, end, what was affected.
What we compute from such a history: Nothing: there is nothing to count yet.
Your step by your answers
2. Recording
How to recognise yourself: Every outage is recorded: when it started, when it ended, what was affected.
What hurts: You see what breaks but not what matters more: all outages are equal and nobody promised a recovery time.
The move up: Name the main services and their owners, set priority by impact and urgency, promise recovery times and put someone on call.
What we compute from such a history: Recovery time and total downtime, where downtime piles up, the worst hours of the week, which outages come back on a cycle and pull others along, outage frequency for 7 and 30 days and the chance of a long outage.
What the step is made of:
- every outage is recorded Incident management
Your step by your answers
3. Managing
How to recognise yourself: Main services have owners, outages have a priority, recovery times are promised and the on-call rota works.
What hurts: Outages close fast but come back: nobody looks for the cause, and changes are not recorded, so there is nothing to link a failure to.
The move up: Investigate repeats down to the cause, record changes with a type and a result, and learn about outages before the consumer does.
What we compute from such a history: Everything from the step below, plus a breakdown of outages by priority.
What the step is made of:
- main services have owners Service level management
- priority is set by impact and urgency Incident management
- recovery times are promised Service level management
- an on-call rota and a hand-over order Service desk
Your step by your answers
4. Preventing
How to recognise yourself: Repeats are investigated down to the cause, changes are recorded, outages are known before the consumer notices.
What hurts: Improvements are made, but nobody checks whether they worked: the measure before and after is never compared.
The move up: Give every improvement an owner and a measure, review the measures together with the consumer, and write down which equipment the services depend on.
What we compute from such a history: Everything from the steps below, and the date the outage frequency changed can be checked against the recorded changes.
What the step is made of:
- repeats are investigated down to the cause Problem management
- changes are recorded with a type and a result Change enablement
- outages are known before the consumer notices Monitoring and event management
Your step by your answers
5. Improving
How to recognise yourself: Improvements have an owner and a measured effect, measures are reviewed with the consumer, and service dependencies on equipment are known.
What hurts: There is no common recipe beyond this: a setup done once goes stale together with the infrastructure.
The move up: Hold the step with repeat exports: compare the measures against the previous period and check that what was done has worked.
What we compute from such a history: Everything from the steps below, plus a check that an improvement worked: whether the outage frequency changed after it and how many outages to expect next.
What the step is made of:
- improvements have an owner and a measured effect Continual improvement
- measures are reviewed with the consumer Measurement and reporting
- service dependencies on equipment are known Service configuration management
What the ladder rests on
The ladder rests on ITIL 4 - the practice names come from there; on the approach of formal process assessment (ISO/IEC 33000, the CMMI appraisal method) - from there comes the rule of two kinds of evidence, what people say and what the records show; and on Gartner's types of analytics - what happened, why, what will happen, what to do. The survey makes a function "declared", your export makes it "confirmed", and a mismatch makes it "disputed".
Find your step
Twelve questions, a few minutes. Your answers show the step in your own words; an export of your log shows it from your data.
What next with your step
Send an export of your outage log for 90 days or more: from it the engine computes what your step opens.