Why you can trust our numbers
The main question for any "AI analytics": where does the number come from and is it invented. Here is a direct answer: how our analysis works and why AI with us does not make up numbers.
The task
We have seen more than once AI reports with convincing but invented numbers. The model confidently names a figure that is not in the data. To the question "where is this from?" it answers "sorry, I do not know", and you have to answer for the mistake. After a couple of such cases, trust in "AI analytics" is gone.
The second barrier is the data. Logs are dirty, from different systems, not ready for upload. Nobody wants to clean them by hand just to "see if it makes sense".
How it works with us
- 1. AI - seriously and at the front. Several language models, each for its own task: one writes the analysis of the findings in plain language, another checks it independently. A suitable tool is chosen for each step.
- 2. We know the limit of AI. Language models, even the strongest, tend to invent numbers with confidence where they do not know the exact answer. This is a known property of how they are built ("hallucinations"), not a rare failure of one model.
- 3. So we do not trust AI with the calculations. All numbers are calculated by a software statistical engine in Python (numpy / scipy / statsmodels - libraries for calculations and statistical tests). AI only describes the fact that is already calculated in text and does not change the result of the calculation.
- 4. Cross-check. In the paid analysis a second independent model checks the text against the calculated facts and catches mismatches before the report reaches you. You get a verified result, without the kitchen of checking.
Like an editor and a journalist: a reader values that the text went through editing, but reads the finished article, not the notes in the margin. We show that there is an editor, and we hand over the article.
How we check each method
A method gets into the report only after a check that we call a circle. The engine does not show a method without a circle. Now 12 methods have passed the circle, 13 wait for their turn and are hidden, and 4 we removed: they did not pass the check.
- 1. A log with a known answer. We build a log in which we put a picture ourselves: for example, three big failures in a year or a task that fails every Thursday at 02:00. We calculate the right answer by hand. The report must name the same systems, dates and numbers. There are 72 such logs now, and they run at every change of the engine.
- 2. Forty years for each picture. One lucky run proves nothing. So for each picture we create 40 different "years" of the log and see in how many of them the method finds it. An example: a task that fails at night in half of the weeks was found in 40 of 40 years; the same task by day, in a third of the weeks, only in 7 of 40. We write this honestly: a weak signal is not always seen by the method.
- 3. Years without a picture. The main check is the reverse one. We create years where there is no pattern and count false findings. By time of day - 0 of 40 years, by links between systems - from 0 to 1 of 40, by repeats on a site no site was named in any of the 40 years.
- 4. The threshold of chance. Each finding is compared with what chance would give at the same work rhythm. If there are many checks, the threshold becomes stricter by the same factor: with 200 pairs of systems checked, the threshold of 5% is divided by 200 (0.025% per pair). Otherwise, among 200 checks one or two "findings" would appear by chance. And one more thing: the finding must stay if you remove the two busiest days. One big failure is not counted as a pattern.
What the checks found and fixed
Everyone has errors in calculations. It is important that they are found before the client does. A few examples of what the report wrote before the check and what it writes now.
| The situation in the log | What the report wrote | What it writes now |
|---|---|---|
| Old tickets were closed all at once on one day (a queue clean-up) | "VPN - 3066 days of downtime" | Such tickets are not part of downtime. The report names the day and the number of tickets |
| Four systems break in the same way | "Downtime is concentrated in three systems" | "Concentrated" only if the share of the worst is higher than what chance gives. Otherwise: "by this data the culprit is not visible" |
| The duration is written as "2,71" or "02:42:33" | 0 incidents: of five ordinary forms of the duration, only one was read correctly | All ordinary forms are read, including the Jira and ServiceNow formats |
| A year with no pattern by hours | "The worst window is Tuesday 16:00, 2% of all outages", with advice to put a shift there | "No unusual hours" |
| The file has only the start time, no end | Downtime 0, availability 100%, the goal "met" | "No end time": recovery time and downtime are not calculated |
| Records "outage not confirmed" | Counted as failures: median recovery 14 min | Removed before the calculations: 21.5 min |
An example of the result
A bank's failure log for 2022: 802 records. In 210 of them there is the mark "outage not confirmed": monitoring fired, but there was no failure. If you count all records in a row, the median recovery time is 14 minutes. The engine removes these 210 records before all calculations and gets 21.5 minutes on 592 real failures: false alarms are short and made the repair a third faster than it is. The report writes how many records were removed, by what sign, and shows five of them for checking.
On the same 592 failures, the 52 longest (8.8%) give 80% of all downtime. This number is not "picked": it comes from a simple count - sort the failures by duration and add them up from the top until 80% is reached. Repeat the same count in your own table - you get the same number.
Even before conclusions, the engine assesses the quality of the export. On a public log of disk failures (1057 rows) it gave the grade B: 1047 rows are usable. And it noticed that all failures are recorded exactly at 02:00 - this is the time of a night check, not the time of the breakdown. So the report will not call 02:00 a "dangerous hour", but will write "time of the record".
What is needed from your data
Usually nothing has to be prepared: you already have the data. An export from Jira, ServiceNow, Zabbix or a simple list of incidents in CSV / Excel will do: the start and end time of the event (or a ready duration) and, if possible, the name of the service. Before the analysis the engine itself assesses completeness and gaps and shows a data quality grade.
Sometimes there are non-standard exports: a new format, a rare export error, mixed sources. Then we do not throw the data away and do not fit it blindly, but sort it out and clean it by hand for your situation.