Logo
Logo

Innocent Until Proven Green

Vriti Magee | Oct 1st 2026

IMG_2026-10-01-111351.jpeg

"Three alibis. All green. One brick declining to comment." Illustrated by DALL·E

Tens of thousands of alerts. None of them knows the others.

You are on call. An application has stopped responding, and the alerts arrive by the thousand.

"They come without any shared context."

Nobody can say what is wrong, or whose problem it is.

The Bridge

A call begins. It takes about 20 minutes to get the right people onto it. The cloud engineer is asked to join and bring evidence.

Everyone has a dashboard. A large enterprise may run around 100 of them, across network, infrastructure, application and cloud. Each shows its own slice and nothing else.

On every one of them, the logic is the same:

"If it's green, it's good. If it's red, it's bad."

Very simple colouring. And yet:

"why is the application unreachable when the VM looks healthy?"

Hours will pass before anyone reaches the first fix. Not the time spent fixing, but the time spent finding out what there is to fix.

"Majority of cost of MTTR is a coordination, not diagnostic."

"That's a coordination that you can't see in any dashboard."

The Map Is a Week Old

Someone opens the topology. The application's path runs through data centres, clouds and on-premises networks, and it keeps changing.

"The topology that they get is an old topology, it's a week ago, and it's not related to the incident that happened right now."

The Space Between

The suspects are many: accounts, service accounts, regions, VPCs, transit. The hardest place to look is between them, at the transit gateways and the interconnects. You ask who owns them.

"nobody owns any of these sections that we can fix that issue at all"

One Incident, Many Dialects

Kubernetes alone goes by GKE, AKS, EKS, OpenShift and Nutanix. Azure organises around resource groups, AWS around accounts, Google Cloud around projects. Logs arrive as syslog, as messages, as query results, in different versions and formats. The enterprise side speaks SNMP.

You are expected to speak all of it.

The Bill

Every tool added to help brought a console, a skill set and an integration that someone built and maintains by hand.

"That's a big cost. And nobody talks about it"

Now Imagine It Differently

The same night. The alerts do not arrive as thousands. The related signals are brought together and printed as one incident, with its dependencies mapped.

"these have never ever been independent issues"

Each record is already carrying context when it reaches you: topology, ownership and maintenance. It was attached as the data came in, not reconstructed on the call.

"Context attached on the move when the data moves in, not reconstructed when the data moves out."

A cloud resource has one structured name: cloud, account, region, type, then the provider's own ID. So a question about a region is simple.

"just a prefix match on this normalized ID."

The alarm says more than that something is wrong with an instance.

"It comes with where it is, where it fits"

It names the construct to start looking at and the team to engage. A failure that raises a trap on-premises and an alarm in the cloud is recognised as one event, not two unrelated ones.

The call still happens, but it starts with an owner and somewhere to look. What operators say they want is modest:

"I just want to see an alert that tells me what's going on."

None of This Happened

All of it was described.

At Cloud Field Day 26, Selector framed hybrid observability as an operational problem more than a technological one. Most tools are vertical pipelines: metrics, logs, events, configurations and CMDB data each have their own collector, store and view, so someone rebuilds the context during the incident.

The platform organises by stage instead: collection, normalisation, correlation and root cause analysis, then output an operator can follow. Naming, grouping and ownership are settled when data arrives, not when the alerts pop up. A resource has one structured name. Resources roll up into constructs that say which team to engage. Alarms in three provider formats are normalised, so related signals form one event.

The decision on what is good or bad stays with machine learning and rules, and the language model builds queries and explains.

Data first, model second.

Green is simply good. The method was to clear what is good until what is bad is left. The first incident was summed up in one line:

"Everything on the path is healthy except one firewall rule"

A question worth the time: what would it take to attach this context at ingest in your own estate?

The Selector sessions are a good place to start.

🔍 Links for Further Reference

Watch the full Cloud Field Day 26 sessions:

Recent Articles