Skip to content

info@dotslash.co.uk  ·  LinkedIn

New: The Service Operations Framework, Second Edition.  Read it →

IT Service Operations · Monitoring & Alerting

Signal, Not Noise.

Everything Service Operations does starts with a signal. Get it right at the source and every decision after it gets sharper. Get it wrong and you pay for the noise forever.

Monitoring and alerting sits at the front of the operations loop. It is where the platform senses what is happening and raises the few signals worth acting on. Most enterprises get this wrong, and they spend millions doing it. In the Service Operations Framework, this is The Senses: the layer everything else stands on.

Hero Visual Options-selectionMA

The Problem

The SLUDGE Tax.

A large estate generates between five and twenty million raw events a day. Perhaps a tenth of one percent are worth acting on. The rest is noise, and you pay for it more than once: to generate it, to ship it across the network, to ingest it, to process it, and to store it.

The bill is real. A team of eighty engineers each losing four hours a week to noisy alerts and mismatched logs burns roughly £800,000 a year in wasted talent alone. Five terabytes a day of verbose log data at commodity rates adds another £3.5 million. Call it a £4.5 million SLUDGE tax, before you count a single delayed release or a single exhausted on-call engineer.

The worst part is what the spend does not buy you. Systems still fail in production with no warning. When a revenue-bearing service goes down, research puts the cost of major downtime at around £25,000 a minute. You are paying a fortune to be blindsided.

Enterprises are spending millions to watch their systems badly. Most of that spend buys noise.

(Engineers × hours wasted × £/hour) + (daily ingest × £/GB)
= annual drain.

5-20m

raw events a day

0.1%

worth acting on

£4.5m

annual SLUDGE tax

£25k

a minute of major downtime

What you need

What You Actually Need.

Most monitoring estates were never designed. They accumulated. Every new tool, every new team, every new agent bolted on without anyone asking what the business actually needed to see. The single pane of glass was sold as the dream and arrived as the problem.

Start the other way round. Define what each service needs before you choose anything to deliver it. What has to be detected? How fast? What does an alert have to prove before it earns the right to wake someone at three in the morning?

The secret to controlling this is standardisation. When you standardise how data is collected, you decouple your telemetry from expensive proprietary agents, which allows you to match the spend to the risk. A commodity service does not need a premium platform watching it. Where the data is common and the stakes are low, open standards do the work: OpenTelemetry and open collectors capture, standardise and route telemetry at no licence cost. Save the premium tooling for the services where minutes cost millions.

 

Our capability matrix gets you there in weeks. It maps the capabilities you need against the services you run and the risk each one carries, so you buy exactly what the business needs and nothing it does not.

You don't always need premium licences. You need the right capabilities, defined with purpose.

ServiceRiskTooling Revenue pathHighPremium Regulated pathHighPremium Back Office SystemsMediumMixed Commodity estateLowOpen standards

Digital First

Spend Where It Counts.

You might not call yourself a technology company. Your customers do not care. The checkout. The booking screen. The payment path. The login. That is where they touch you, and increasingly it is the only place they touch you. When those flows stop, you are not degraded. You are closed.

So one question should drive every monitoring decision: which paths earn the revenue, and which carry the regulatory weight? Watch those as if survival depends on them, because it does.
 

 

Every company is a digital business now. Digital is where you meet your customers, and if it is down, you are shut.

The history is not subtle. A data-centre failure grounded an airline and cost it around £58 million. A banking outage locked customers out at month end and broke trust at the worst possible moment. A retailer lost the better part of two days of online orders. The pattern holds across transport, banking, telecoms and retail: the financial hit scales with how many customers are online, how fast the money moves, and how long it takes to respond.

This is where the premium spend belongs. Instrument the revenue-bearing and regulated paths to the highest standard you can, watch them from the customer's side as well as the system's, and let the rest of the estate run lean on open tooling. That is not a cost compromise. That is putting the money where the business actually lives.

The Three Things

Strategy, Standards, Governance.

Three things keep monitoring under control. Without them, SLUDGE comes back, costs climb, and production fails in silence.

Strategy

Decide what good looks like before you tool. What are you detecting, for whom, and to what standard? Strategy turns a thousand local decisions into one deliberate design.

Standards

One schema, enforced everywhere. Every log, event and metric carries the same core fields: service, environment, host, severity, owner, business unit. Without that, you are pattern-matching on inconsistent data and debugging in three query languages at once. With it, correlation becomes deterministic and the CMDB has something real to bind to.

Governance

Every alert proves its right to exist or it dies. Run the audit: when did it last fire, who acted, what did they do, and what would have happened if it had not? If you cannot answer all four, switch it off. Then tag every telemetry stream with a cost and show it to the team that owns it. Engineers are natural problem solvers; give them visibility into the cost of their noise, and they will tune it out themselves.

Governance should feel like rails guiding a train, not a wall across the track. Wrap your standards in red tape and your best engineers will route around you by Friday.

The Suppression Delusion

The Suppression Delusion.

It is the proudest number in a lot of operations centres: look how much we suppress. But stop and ask what it means. Every one of those events was triggered by something, and you decided in advance it was worthless. So why generate it at all?

You did not save that work. You paid for it. You paid to ingest the event, paid the compute to inspect it, paid the licence to process it, then paid again to throw it away. That is a noise tax, billed every month, for the privilege of saying "ignore this."

Suppressing fifty thousand events a month is not a win. It is a confession.

Suppression is not even safe. Rules rot. Infrastructure changes. The event that was noise six months ago might be the only warning you get today. Build the wall of suppression rules high enough and you can no longer see the fire behind it.

The fix is not a bigger filter. It is a better signal. Cut the noise at the source so the event is never emitted in the first place. Where low-level signals genuinely matter, correlate them into incidents rather than hiding them, so the one that counts is still there, in context, instead of gone. Then stop measuring "events suppressed" and start measuring the only number that matters: what share of everything you generate actually needs a human or a machine to act.

The Shift to Monitoring as a Service

Built By Owners.

Central monitoring teams cannot keep pace with a thousand services they did not write. So they over-instrument, log everything just in case, and the SLUDGE builds straight back up. The model is broken.

Turn monitoring into a service the business consumes. Build standard, pre-configured patterns for the technologies teams actually run: Linux and Windows, the language runtimes from Java to JavaScript, Ruby, PHP, the databases, the container platforms and so on. Each pattern carries the right metrics, the right thresholds and the right standards baked in.

Then hand them over. Teams deploy the patterns themselves, as part of the release cycle, with the standards enforced as code so nothing ships without the monitoring it needs. Doing the right thing becomes the easy thing.

This is the shift that matters. It brings development and operations together, moves observability to where the knowledge lives, and drives better behaviour by default rather than by mandate. Clean, standardised, owned telemetry is the foundation everything else stands on: correlation that works, automation you can trust & services that begin to heal themselves. 

This is the gold under the SLUDGE. It always was.

In the Service Operations Framework, this is The Senses: the layer everything else stands on.

How We Engage

Where We Start.

We make your existing stack work harder before we ever suggest buying more.

A dotslash engagement starts by facing the estate honestly. We quantify your SLUDGE tax in real numbers, define the capabilities each service needs against the risk it carries, set the standards, and stand up the patterns and governance that keep the noise out for good.

Sense, Then Act.

Cut the noise at the source and everything downstream sharpens: cost down, signal up, fewer surprises in production. And the payoff compounds, because clean signal is what Observability and AIOps stand on. First you sense. Then you act.

92%

less alert noise

38%

faster restoration of service

£3.1m

annual run cost removed