Infrastructure reliability consultant

Vadim Alekseev

20+ years in IT operations. I built and ran a NOC (Network Operations Center - the team that watches infrastructure health around the clock), led Problem Management (the practice of finding root causes, not just closing tickets) in a bank, and ran independent infrastructure audits. Reading incident data is my job. OpsLab is the statistical engine I built to do it faster and to prove the pattern with numbers.

NOC / Monitoring Problem Management ITIL SLA design Incident statistics

Track record

A NOC built from scratch, Problem Management in a bank, independent audits. 1999-2023.

See more →

My approach

The engine finds the pattern. I read what it means for your infrastructure.

See more →

Ways to work together

Three clear formats: diagnostics and retainer, a run inside your perimeter, interim role.

See more →

What the engine looks for

Six core methods out of 20+: priority, root causes, cycles, before/after, tail risk, event links.

See more →

Why the numbers hold

A deterministic engine computes; the AI only writes. Two full sample reports to check.

See more →

Contacts

Email, LinkedIn, Telegram and the services deck as a file.

Get in touch →

Built reliability practices, not slide decks

2005-2019 · Rostelecom

Built the North-West Russia NOC from scratch

Cut MTTR (mean time to repair - how long it takes to fix an incident) by 20%. Built a failure-forecast model on historical data and prevented 4+ major network incidents a year.

2021-2023 · Bank

Led Problem Management and monitoring in a digital transformation

Owned Problem Management (the root-cause practice) plus 12 more ITIL practices. Built a monitoring service from zero. Held the SLA at 99.9% ± 0.2%.

2019-2021 · ITSM consulting

Independent infrastructure audits

Cut the failure rate by roughly 50% across engagements. Raised availability from 95% to 99.8%.

1999-2004 · Lensvyaz

Chief engineer - telecom, networks, power

Radio communications engineer by training. "Master of Communications" award, 2013.

The engine finds the pattern. I read what it means.

A statistical report tells you what happened in your data: a repeat cycle, a cluster of related incidents, a shift after a release. It does not tell you whether that matters, or what to do about it. That judgment comes from running a NOC, owning an SLA, and sitting in the room when a bank's systems went down. OpsLab computes the numbers. I read them the way 20 years on call taught me to.

Ways to work together

Three formats. Each one starts small, so you can stop after the first step.

1

Diagnostics → retainer → project

We start with a paid diagnostic of your incident history: what causes most downtime, what repeats, what to fix first. If it proves useful, we continue as a monthly retainer, or as a project - setting up Problem Management or a NOC.

2

A run inside your perimeter

If your data cannot leave the company, the analysis runs on your own machine. You keep the data; I get only the report and read it with you.

3

Interim / part-time role

I take the reliability role for a few months: a fixed share of the week, an agreed scope, a clear exit date. Useful while you are hiring, or during a transformation.

What the statistical engine looks for

Six core methods out of 20+. The full catalog also covers correlations, trend diagnostics, anomaly detection and a per-service breakdown.

Priority (Pareto ranking)

Which 20% of incidents cause 80% of the downtime, so you fix the few that matter. The Gini coefficient (a measure of inequality: 0 - all incidents equally bad, 1 - one incident caused everything) shows how concentrated the problem is.

3 of 90 incidents = 41% of all downtime. Gini = 0.67.

Root causes (clustering)

Groups incidents by behaviour - time of day, duration, frequency - instead of by ticket category. A long list of "different" errors often turns out to be a small number of root causes.

90 incidents → 2 clusters: "night, short" and "morning, long". Two causes, two teams.

Cycles (periodicity)

Finds hidden repeat cycles in uneven time series - a method from astrophysics applied to incident logs. It reports FAP (false alarm probability - the chance the cycle is random noise).

Period = 23.1 h, FAP = 4×10⁻⁸ → almost certainly a scheduler, not chance.

Before / after a change

Tests whether a deploy or a config change really moved the incident rate, and by how much. Cliff's δ (effect size - how big the shift is, from 0 to 1) turns "it feels worse" into a number.

Cliff's δ = +0.445, p < 0.001 → the rate grew, confidence above 99.9%.

Tail risk

Estimates the probability of a rare but very long outage - the one that breaks an SLA. Computed for planning, not for the average case.

1% probability the next major incident lasts 72+ hours - a number for SLA talks.

Event links

Finds which events appear together and in what order, so an early alert can sit on the leading signal instead of the crash itself. Lift shows how many times more often a pair occurs than by chance.

database_timeout → app_crash in 89% of cases, lift = 4.2.

Why the numbers hold

The short path: who I am → what I look with → why you can trust the figures → what the result looks like.

A deterministic engine, not a guessing AI

Every figure comes from numpy, scipy and statsmodels - the same input always gives the same output. The AI models write the text and cross-check each other; they never compute a number.

Works with what you already export

Exports from Zabbix, Nagios, Jira Service Management, ServiceNow: CSV, JSON, TXT, XLSX. No agents, no VPN, no access to your infrastructure - only the file you decided to send.

Let's talk

The first conversation needs no data and no files: 30 minutes about your incidents and what is worth measuring.