Target state · Everyday work

What ideal IT operations look like in an ordinary company: predictability instead of heroics

When people say "world-class operations", they usually show a room with a video wall, a 24/7 shift and a team of reliability engineers. This has nothing to do with an ordinary company. An ordinary company is a few dozen systems, three to fifteen people in operations, no separate process team, a budget that has to be defended, and a manager for whom operations is one of five concerns. For such a company another question makes sense: what does good look like here.

Our answer surprises many: an ideal operations team looks boring. There are no heroics, nobody saves the day at night, and there is not much to tell at a conference. Things break there as often as at the neighbours' - the difference is in what happens after. Below is this state part by part: an ordinary day, the signs of a mature team, what it does not have, what it stands on and how much of the team's time it takes. Where we rely on common practices (ITIL, Google SRE), we say so; the rest is our experience.

  1. What "ideal" means and what it is not
  2. One ordinary Tuesday
  3. Eight signs of a mature team
  4. What it does not have
  5. Four pillars
  6. Chaos, normal, ideal
  7. What it costs in people and time
  8. Ten questions to check yourself

1. What "ideal" means and what it is not

Let us start with what ideal operations are not: this is where an impossible goal is most often set, and then the work on it is dropped. It is not the absence of failures. There will be failures: hardware ages, suppliers go down, people make mistakes, and the business needs changes faster than they can be made safely. A team that set the goal "nothing should break" will always see itself as failing.

It is not availability at any price. Each next "nine" in the availability percentage costs more and more: according to Google SRE, the next step up in reliability may cost a hundred times more than the previous one. And a user whose own smartphone is 99% reliable will not notice the difference between 99.99% and 99.999%. A mature team knows what availability each service needs and does not raise everything to the level of the most critical one.

It is not a thick folder of procedures. A document written once that does not change behaviour at three in the morning is a cost, not an asset. A mature team is recognised by behaviour, not by paperwork: it often has less written down than people expect.

The working definition we use: operations are ideal when the result does not depend on who is on shift today, and any failure ends in a predictable time and leaves a record you can learn from. Everything else follows.

2. One ordinary Tuesday

Descriptions of maturity are hard to remember, so it is easier to show a day. A medium company, a few dozen systems, six people in operations. There is no 24/7 shift - there is an on-call rota with a phone.

08:40

The on-call engineer opens the night summary. It has three lines, not three hundred: events that needed no action went to the log and did not reach the feed. One line of three ended in an incident - a service restart by the instruction, in eleven minutes, without calls.

Nobody woke up at night. Not because the night was quiet, but because the wake-up rules are written down and none of them fired.

09:15

A fifteen-minute morning meeting. Three things are discussed: yesterday's incidents with higher than medium priority, today's changes and one open problem record. The problem has an owner and a deadline, and there is progress on it today.

11:00

A warning comes in: space on the file storage will run out in about two weeks. This is not a failure, and nobody is woken - it is a task for the queue within a week. The growth was visible in advance, because resources are watched by trend, not by waiting to hit the threshold.

14:30

A planned update of one of the systems. The window was agreed a week ago, users are warned, alerts for the system are switched off for the time of the work, the rollback plan is written down together with the decision time: if it does not work by 15:30 - roll back.

It worked at 15:05. For the next day the system is watched closely: any outage on it until tomorrow is by default treated as linked to the update.

16:20

A client reports that an external service is slow. The on-call engineer sees that the alert came eight minutes ago and the incident is already open. The client is told that the problem is known, the cause is at the provider, a request to them is open, the next update - in half an hour.

This is the key moment of the day. There is a failure, and it is not their own, but the client gets a status, not silence, and does not learn the news first.

17:40

An engineer adds a paragraph to the known error database about yesterday's unusual case: symptoms, the workaround, the status of the permanent fix. Five minutes of work that next time will save an hour and a half - and most likely not for him.

Note what this day did not have. Nobody looked for who is responsible for the system. Nobody called the one person "who knows how to fix it". Nobody learned about a failure from a client. Nobody decided on a rollback in panic: the decision was made in advance and calmly. And the day is completely ordinary: two incidents, one change, one external failure.

3. Eight signs of a mature team

These signs are visible from outside and are checked by questions, not by an audit. They need no specific monitoring system and no specific methodology: we have seen them in companies that never heard of ITIL.

  • Incident log: every incident has an owner and a record that can be found in a minute. Not "somewhere in the chat", but in one place: time, object, outcome. Without a record you can neither measure nor learn.
  • A fix without an expert: a known outage is closed by an on-call engineer of normal skill using an instruction. Experts are called for the new, not for the typical.
  • Night calls by rule: the criteria are written down, the on-call engineer applies them in half a minute, and is not told off for a call made by the rule.
  • Order in changes: work happens in an agreed window, every change has a rollback plan with a specific decision time.
  • A repeat is a defect, not bad luck: a repeated outage becomes a problem record with an owner - this is how ITIL separates incident management from problem management. Our working rule: the third identical outage is not fixed a third time, a problem is opened.
  • The company reports an outage first: the share of incidents learned about from clients is known and falling. Perhaps the most reliable outside sign of the eight.
  • A dozen measures: not a hundred and not three - as many as fit on one page and are discussed in half an hour once a month.
  • A report for the business: once a quarter the team itself brings a list of what it improved - in hours and percentages, with the "before" level. This tells a partner from a contractor.

4. What it does not have

Absences say more about maturity than presences, and are noticed faster. Each of them looked sensible in its time - that is the difficulty.

  • No hero. A mature team has no person without whom nothing gets restored. A hero is not an achievement but a concentrated risk: their holiday or resignation is more dangerous than a hardware failure.
  • No manual morning checks. Walking through systems "to see if everything is alive" means monitoring is not trusted. Either monitoring is fixed, or the ritual stays forever.
  • No feed watched out of habit. If the alert flow is read out of the corner of the eye, it mostly consists of things that need no action.
  • No temporary fixes older than a quarter. A workaround put in "for a week" eight months ago is not a solution but an invisible debt that will one day be called in.
  • No reports that nobody reads. A regular document with no addressee and no decisions after it is pure waste of time. Such reports are closed, not improved.
  • No hunt for the guilty in incident reviews. As soon as a review becomes disciplinary, records become neat and useless. Google SRE calls the alternative a blameless review: you review events and rules, not people.

5. Four pillars

Everything above stands on four things. None needs big purchases, and all four are needed: remove any one - and in a few months the structure starts to fall apart.

  • Data: one log where every incident goes - with time, object and outcome. Not for reporting, but to look back in six months and see patterns that are not visible in the moment.
  • Rules: who learns, who acts, in what time, who is next and when a manager is woken. A short document known by heart works better than a detailed one opened once a year.
  • Knowledge: typical outages described - symptoms, the workaround, the status of the permanent fix. This turns one person's skill into a property of the whole team.
  • Rhythm: fifteen minutes in the morning, half an hour a month on measures, half a day a quarter on data review and decisions. Without rhythm the first three pillars quietly go out of date: the infrastructure changes, while rules and knowledge stay from the previous version of the company.

They should be built in one order: data, rules, knowledge, rhythm. Starting with rules without records makes no sense - there is nothing to check that they are followed. Starting with knowledge before rules too: there is nobody to apply the described workarounds if the event does not reach the right person.

6. Chaos, normal, ideal

Between chaos and the target state there is an intermediate one, and most companies are in it. It is not bad: you can work in it for years. The difference is easiest to see in specific situations - here are eight.

SituationChaosNormalIdeal
How they learn about an outage From a client or from management From monitoring, but with a delay and not for all services From monitoring, before the client; the share of cases when the client was first is known and falling
Who takes the incident Whoever saw it first or whoever was found The on-call engineer, then by agreement The on-call engineer by rule, escalation by timer, every role has a backup
Night call By the on-call engineer's feeling; they get blamed for a mistake either way There is an agreement, but everyone reads it differently Written criteria; a call by a criterion is not discussed, a bad rule is reviewed
Changes Whenever convenient; how to roll back is decided on the spot Agreed, but nobody watches the system closely after the work A window, a warning to users, a time for the rollback decision, a day of close watching after
Repeated outage Fixed again every time Noticed, the workaround is known, the cause is not removed A problem record with an owner and a deadline; the workaround is described, removing the cause is in the plan
Incident records In chats and in memory In the system, but filled in any old way So complete that statistics are counted without manual cleaning
Talking to the business Only when something went down An availability report that few people read Once a quarter: what was improved, in hours and percentages, with the "before" level
A key engineer leaves A disaster: part of the systems becomes unclear Painful, they recover for months Noticeable, but not critical: knowledge is described, roles have backups

It helps to go through the table with the team and mark your column in each row. The picture is almost always mixed, and that is normal: the difference between the rows shows where to move more precisely than any overall score.

7. What it costs in people and time

The main objection: "we don't have people for this". The objection is fair, so here is our estimate of the regular load for a team of five to eight people. This is not the cost of the transition, but the cost of keeping the state once it is reached.

  • Daily: a 15-minute morning meeting for the whole team and a few minutes to write up each closed incident.
  • Weekly: about an hour on the problem queue and resource trends.
  • Monthly: half an hour to an hour on the page of measures and a talk about it.
  • Quarterly: half a day on the quarter's data review, revising thresholds and alert rules, and a short summary for the business.
  • Once a year: one drill to restore a key service and one check of starting systems after a full power cut.

In total this is a few percent of the team's working time. It should be compared not with zero, but with the share of time that now goes to emergency and routine work and to finding out again what was already found out. For reference: Google SRE sets the goal of keeping such work below 50% of an engineer's time. Teams that count their share for the first time usually get more than they expected - and this very difference becomes the main argument in a talk with management.

The transition is a separate story, and it is better to plan it not as a one-year project but as a series of short steps by quarter: one step - one measurable result. In our experience, initiatives in operations die not from complexity but from duration: the team cannot stand half a year of work without a visible effect, and it cannot be blamed for that.

8. Ten questions to check yourself

Answer only "yes" or "no", without "mostly" and "we try". "Yes" counts only if it is true not for one model system, but for any key service and in any shift.

  1. Do you know the share of incidents you learned about from a client, not from monitoring?
  2. Can the on-call engineer find in a minute who is responsible for any key system, and their backup?
  3. Are the criteria for waking a manager at night written down, and does the on-call engineer know them?
  4. Does every change have a rollback plan with a set decision time?
  5. Does a repeated outage become a task to remove the cause, instead of being fixed again?
  6. Are the ten most frequent outages described so that a person with no experience in that system can close them?
  7. Are there no "temporary" workarounds in place that were put in more than a quarter ago?
  8. Can you get incident statistics for six months in ten minutes without manual cleaning?
  9. Are planned windows counted in the overall unavailability balance?
  10. Did the business get a summary of improvements in hours and percentages from you in the last quarter?

How to read the result (the scale is ours). Eight "yes" or more: you are in the target state, the task now is to keep it. Four to seven: an ordinary mature company, and it is clear where to move: take the questions with "no" in order of their numbers. Fewer than four: start with the first pillar - recording incidents. Without it, the other steps can be neither planned nor checked.

A short conclusion. Ideal operations in an ordinary company are not about scale or budget, but about predictability. Things break as often as everywhere. The difference is that the system notices the outage, it reaches the right person by rule, ends in a clear time and leaves a record thanks to which everything goes faster next time. None of these properties can be bought.

You can check yourself in two ways. In words - the ten questions above or the maturity test in our bot (7 questions, about 5 minutes, at the end a level and tips). From data - send an export of incidents: we will return an analysis: where downtime concentrates, how fast things are fixed and at what hours outages pile up. A survey shows the team's ideas, an export shows its behaviour.

Related articles: five steps of operations maturity, escalation matrix and alert noise, what to do with the findings of the report. Discuss your situation with an operations engineer - .