Articles
What you can learn from incident logs
Short articles on real questions of IT operations. Each one is about one question: what we find, with which method, and why you can trust the numbers. Data in the examples is anonymised.
What the engine does
Priority
Which incidents cause the most downtime
On real data, 8.8% of incidents gave 80% of downtime.
Reliability
MTBF and MTTR: how often it fails and how fast it is fixed
How often each service fails and how fast it is fixed.
Patterns
At what hours services fail most often
A pattern in time that you cannot see in the moment.
Cycles
Which outages come back in a cycle
An outage every few days or at month end, and the date of the next one.
Forecast
How many more outages and when the next one
A forecast of outage frequency and of the risk of long downtime.
Root cause
Which outages pull others along
Which outage brings tickets for other systems, and how many tickets are one failure.
Repeats
Which sites break again and again
A site that breaks more often than its neighbours or comes back right after a fix.
Early warning
Which metric warns of an outage
A metric that behaves unusually hours before an outage, and how long before it.
Cleanup
How much of the log is requests, not outages
Requests for access and passwords in the same queue as outages inflate failures and downtime.
Check
Did it get worse after a change
The date from which outage frequency changed, and a check that it is not chance.
How the analysis works
Why one method is not enough
The worst system by downtime may be the victim of someone else's failure. In what order we compute, and which method covers the blind spot of another.
Trust
Why you can trust our numbers
The engine computes the numbers; AI cannot invent a figure.
Operations
Maturity
The operations ladder: five steps
From "we learn about failures from users" to improvements with a measured effect. Where your team stands on this ladder and what the next step is.
Outage types
Ranking of outage types: what harms operations most
Planned work, providers, power, how operations is organised - twelve sources of outages ranked by lost hours, not by frequency.
Alerts
Escalation matrix and alert noise
Why most signals are not worth reading, how to measure the share of noise, and whom to wake at night - by a written rule, not by feeling.
Practices
Practices the business can see
Twelve ITIL practices ranked by return: kept revenue, client trust, access to contracts, free team capacity.
Target state
What ideal operations look like
An ordinary company, not Google: one ordinary Tuesday hour by hour, eight signs of a mature team and ten questions to check yourself.
What comes next
What to do after the analysis
Which problems we find most often and what to do about them with ITIL / ITSM.
The first step - a data quality grade - is free. All data is processed in isolation and is not shared with third parties.