At what hours services fail most often: how to see systemic outages behind the noise in the ticket log
Most teams look at outages in the moment - through monitoring charts. But monitoring shows only the "fire" itself, not its root cause. As a result, regular breakdowns are often put down to chance, and the team spends more effort putting out symptoms instead of removing the source of the problem.
The task
A deeper analysis of incident times lets you get from the data log not just statistics, but extra practical value:
- Separate conflicting processes. Tell a one-off failure from a regular outage caused by background jobs "overlapping" (backups, data exchange, release windows, heavy night calculations).
- Audit the quality of data and processes. Find systemic distortions in the records - for example, technical spam from scanners and checks that lands in the log in one second - and allow for false morning peaks, when night outages are logged only at 09:00.
- Get arguments for contractors and shift schedules. Rely on hard facts (the number of weeks, the system, the exact hour) when you revise the SLA with vendors or plan on-call duty, instead of assigning people at random.
A descriptive analytics algorithm reads the incident log automatically, filters out single spikes and shows where the error is in the schedule, and where it is the natural working rhythm.
How the algorithm looks for patterns
To see the real picture of incidents, the algorithm uses two methods of descriptive analytics.
- Shares of the week: when a team is needed
Basis: all incidents are split into three groups: standard working hours (weekdays 9:00 to 18:00), evening and night on weekdays, and weekends. Even tickets that are still open are counted, because the time they started is already recorded.
The point: weekday working hours are only 27% of the time of the whole week. If, say, 43% of incidents fall into this period, then the other 57% of outages happen when nobody is in the office. That is already a reason to think about how to plan on-call staff and work.
- Looking for an "unusual hour": where a systemic outage hides
The busiest hour is not yet an anomaly. There are always more active users and tickets during the day. So the algorithm looks not for peaks, but for deviations from the usual rhythm.
Comparison with its own norm: Thursday night is compared with the nights of other weekdays, and Saturday with Sunday. If every week at 02:00 on Thursday there are twice as many tickets as on Wednesday or Friday, this is already a schedule, not just load.
Check for repetition: one big failure that generated 30 tickets in 60 minutes is not accepted by the system as a pattern. An anomaly is recorded only if outages at this hour repeat in different weeks.
Looking for daily peaks: separately, it finds an hour that is steadily busier than its neighbours every day. This is how heavy daily jobs look, for example backups or night calculations.
Honest data: if the chart of outages is just a random spread, the system will not invent a "worst time" artificially. The report will say it directly: "no unusual hours".
Working with real data: traps in the records
Incident logs are never perfect. The algorithm's task is to recognise the usual artefacts of data collection and not to raise false alarms where a human or technical factor is at work.
- How the algorithm handles distortions in the data
One big failure: a burst of dozens of tickets in one hour is classed by the system as a single incident, not as a repeating pattern. There is no need to look for a schedule here.
Technical spam in one minute: if hundreds of records appear exactly at 02:00, the system treats this as the time a script or an automatic check wrote them, not the time of the breakdown, and does not look for a pattern in these records.
Morning peaks instead of night incidents: if a service went down at night, but the ticket was registered only when staff arrived at 9:00, a false spike appears. The report warns about this possibility, so that the team can introduce automatic night monitoring for problem services.
Format limits: if the export has only dates, the system does not invent a fake hour "00:00" and analyses strictly by day of the week. If the time is recorded in UTC, it reminds you to shift it to the local time zone.
Decision matrix: what to do when anomalies are found
| Type of anomaly | Possible cause | What the team does |
|---|---|---|
| The same day and hour every weekfor example, Thursday 02:00 | Weekly data exchange, a heavy report, a release window, planned work by a provider | Move conflicting jobs apart in time; add an automatic check after a release; agree the window with the vendor again |
| Every day at the same timefor example, 04:00 | Backups, night calculations, a restart of the service pool | Spread the scheduler jobs; give extra capacity to the specific night window |
| Peaks at the end of work shifts or weeks | "Shadow IT" at work: heavy user macros, mass data exports by business departments | Find who starts it (from database or load balancer logs), optimise the queries, move the run to non-working time |
| Regular spikes at an hour when there are no jobs | External automated activity: aggressive scraping, password guessing, unauthorised integrations through the API | Limit the request rate (Rate Limit), update the firewall rules (WAF), block anomalous IP addresses |
| An ordinary rhythm with no patterns | Ordinary load: more by day, less at night | Stop looking for non-existent problems in the system schedules. Switch to on-call schedules and automatic resource scaling (Auto Scaling) for daytime peaks |
Example 1. A hidden outage on a schedule
Situation: a retail chain with an online shop registers about 1100 incidents a year. Staff complain that the night data exchange often fails, but against the general background this is not visible.
What the analysis showed: on Thursday from 02:00 to 03:00, 37 outages are recorded against a norm of 4 for this time. The situation repeated in 32 weeks out of 52. The main source is 1C:ERP.
Result: instead of vague arguments, the team gets the exact time and the system. Reconfiguring the schedule of the background job removes more than 30 regular IT failures a year and spares the warehouse from morning downtime.
Example 2. Optimising on-call duty instead of hunting "ghosts"
Situation: the head of support wants to change the shift schedule, suspecting regular load spikes at night.
What the analysis showed: working hours (27% of the week) account for 43% of incidents. The other 57% are spread over evenings, nights and weekends. No single hour stands out from the general norm - there are no "unusual hours".
Result: there is no sense in looking for a "worst window" and breaking people's schedule because of a random spread. It is better to work on the on-call rules: define clearly which of these 57% of outages need an engineer woken up at night at once, and which can wait until morning.
How to choose a window for planned work
Planned work (updates, migrations, hardware replacement) is risky in itself. If you put it where something already fails regularly, two outages add up, and afterwards it is hard to tell what brought the system down. The report helps you choose a window by three signs:
- Not in the unusual hour that was found. If the report found "Thursday 02:00" or "every day 04:00", some job already runs at this hour. Work on top of it competes with it for resources, and it is hard to sort out later who brought down what.
- Where there are few outages and someone is there to fix them. On the heat map in the report, light cells are the hours with the fewest outages. Night and weekends are usually quieter, but there are also fewer people: the shares by period of the week show how many outages will fall into the window, and your on-call schedule shows who will be there if the work goes wrong.
- With a correction for the registration time. If the log has the time of the ticket, not the time of the breakdown, a night outage looks like a morning one. A quiet night in the report may mean that nobody was simply looking at night.
This is how the heat map looks in the report: rows are days of the week, columns are hours. The darker the cell, the more outages; light cells are candidates for planned work. The dashed frame is weekdays from 9 to 18. The red frame is the hour that stands out from the usual rhythm (here Thursday 02:00, 37 outages): do not put work there.
What the report does not know: your on-call schedule, business peaks (sales, month close) and the cost of work outside working hours. You add this data - and of the quiet hours, those remain where there are people and no sales peak.
Summary: how to use the data in your work
An analysis of outage times is a tool for managing resources, and it affects three areas directly:
- A smart schedule for peopleThe shift schedule is built on the real share of incidents in different periods of the week, not on the illusion that "all is quiet at night".
- A safe schedule for systemsReleases, backups and heavy exports are not started at the hours that the algorithm has already marked as problem ones. An anomaly that was found is the first candidate for an audit and a change of schedule.
- An objective dialogue with contractorsIf the "cursed hour" that was found steadily matches the provider's work window, you have exact facts in hand: the number of incidents, the number of weeks, the affected systems. This is an argument for moving the maintenance window or revising the SLA in the contract.