20+ years in IT operations. I built and led several network monitoring and operations teams, set up Problem Management in a bank, and ran independent audits followed by infrastructure optimization. Reading failure data and finding solutions is my profession - OpsLab is the tool: an exclusive statistical engine built for modern analysis and forecasting.
A NOC built from scratch, Problem Management in a bank, independent audits. 1999-2023.
See more →Eight analysis methods with real numbers: priority, root causes, cycles, before/after, forecast, event links, MTBF/MTTR, industry norms.
On the project site →What failure data shows and what to do about it: monitoring maturity, alert noise, the incidents that cost you most.
Read →The cost of downtime, how the industry fights it, what statistics can pull out of an incident log, and where AI helps - with sources.
Download PPTX →Email, LinkedIn, Telegram and the services deck as a file.
Get in touch →Cut MTTR (mean time to repair - how long it takes to fix an incident) by 20%. Built a failure-forecast model on historical data and prevented 4+ major network incidents a year.
Owned Problem Management (the root-cause practice) plus 12 more ITIL practices. Built a monitoring service from zero. Held the SLA at 99.9% ± 0.2%.
Cut the failure rate by roughly 50% across engagements. Raised availability from 95% to 99.8%.
Radio communications engineer by training. "Master of Communications" award, 2013.
Short pieces on what failure data shows and what to do about it. Written from practice, not from textbooks.
From "we learn about outages from customers" to tuning against business metrics. Where your team actually stands, and what the next step is.
Read →Which problems the analysis finds most often, and what to do about each one - in ITIL / ITSM terms.
Read →Why most alerts are not worth reading, how to measure the noise, and who should be woken at night - by a written rule, not by feel.
Read →Twelve ITIL practices ranked by what they return: kept revenue, customer trust, access to contracts, free capacity in the team.
Read →Planned work, providers, power, the way operations are run - twelve sources of failure ranked by lost hours, not by how often they happen.
Read →A normal company, not Google: one ordinary Tuesday hour by hour, eight signs of a mature team, and ten questions to check yourself.
Read →Just name the main problem you are facing, and we will take it from there.