20+ years in IT operations. I built and ran a NOC (Network Operations Center - the team that watches infrastructure health around the clock), led Problem Management (the practice of finding root causes, not just closing tickets) in a bank, and ran independent infrastructure audits. Reading incident data is my job. OpsLab is the statistical engine I built to do it faster and to prove the pattern with numbers.
A NOC built from scratch, Problem Management in a bank, independent audits. 1999-2023.
See more →The engine finds the pattern. I read what it means for your infrastructure.
See more →Three clear formats: diagnostics and retainer, a run inside your perimeter, interim role.
See more →Six core methods out of 20+: priority, root causes, cycles, before/after, tail risk, event links.
See more →A deterministic engine computes; the AI only writes. Two full sample reports to check.
See more →Email, LinkedIn, Telegram and the services deck as a file.
Get in touch →Cut MTTR (mean time to repair - how long it takes to fix an incident) by 20%. Built a failure-forecast model on historical data and prevented 4+ major network incidents a year.
Owned Problem Management (the root-cause practice) plus 12 more ITIL practices. Built a monitoring service from zero. Held the SLA at 99.9% ± 0.2%.
Cut the failure rate by roughly 50% across engagements. Raised availability from 95% to 99.8%.
Radio communications engineer by training. "Master of Communications" award, 2013.
Three formats. Each one starts small, so you can stop after the first step.
We start with a paid diagnostic of your incident history: what causes most downtime, what repeats, what to fix first. If it proves useful, we continue as a monthly retainer, or as a project - setting up Problem Management or a NOC.
If your data cannot leave the company, the analysis runs on your own machine. You keep the data; I get only the report and read it with you.
I take the reliability role for a few months: a fixed share of the week, an agreed scope, a clear exit date. Useful while you are hiring, or during a transformation.
Six core methods out of 20+. The full catalog also covers correlations, trend diagnostics, anomaly detection and a per-service breakdown.
Which 20% of incidents cause 80% of the downtime, so you fix the few that matter. The Gini coefficient (a measure of inequality: 0 - all incidents equally bad, 1 - one incident caused everything) shows how concentrated the problem is.
Groups incidents by behaviour - time of day, duration, frequency - instead of by ticket category. A long list of "different" errors often turns out to be a small number of root causes.
Finds hidden repeat cycles in uneven time series - a method from astrophysics applied to incident logs. It reports FAP (false alarm probability - the chance the cycle is random noise).
Tests whether a deploy or a config change really moved the incident rate, and by how much. Cliff's δ (effect size - how big the shift is, from 0 to 1) turns "it feels worse" into a number.
Estimates the probability of a rare but very long outage - the one that breaks an SLA. Computed for planning, not for the average case.
Finds which events appear together and in what order, so an early alert can sit on the leading signal instead of the crash itself. Lift shows how many times more often a pair occurs than by chance.
The short path: who I am → what I look with → why you can trust the figures → what the result looks like.
Every figure comes from numpy, scipy and statsmodels - the same input always gives the same output. The AI models write the text and cross-check each other; they never compute a number.
Exports from Zabbix, Nagios, Jira Service Management, ServiceNow: CSV, JSON, TXT, XLSX. No agents, no VPN, no access to your infrastructure - only the file you decided to send.
The first conversation needs no data and no files: 30 minutes about your incidents and what is worth measuring.