Logiciel Solutions Contact Us
Success Stories Tech News Contact Us
framework

17% Of The Compute Bill Buys Telemetry. Most Of It Answers Nothing.

Seven dimensions, each scored on five levels, weighted into a single maturity index between 1.00 and 5.00. You tick only the level you can prove from your last real incident, and the band you land in names the one move worth funding this quarter rather than the four your vendor would like to sell you.

In depth

Ingest Volume Rises Every Year. Incidents Get No Shorter.

01

The trap most platform teams walked into: buy the platform, turn on every agent, forward all telemetry and sort it out later, then measure success in dashboards built and integrations enabled while attribute names stay whatever each SDK emitted, so nothing joins across services and the same three people answer every page.

02

What the teams who debug fast do instead: start from the ten questions they actually asked in the last five incidents, instrument to answer those, enforce OpenTelemetry semantic conventions at the collector rather than requesting them in a wiki, and settle sampling and retention as policy before the volume exists.

The detail

What Separates An Assessment From A Vendor Checklist.

Zone · 01

Proof

Not Licensed Capability

Score the level you have demonstrated, not the level your platform supports. The proof is a link: a trace that crossed the async boundary, an alert that routed correctly at 3am, a deploy marker sitting next to the latency spike. If you cannot produce that link from the last real incident, drop a level and carry on.

Zone · 02

Weights That Match Blast Radius

Instrumentation coverage carries a weight of 20 and cost control carries 5, because a missing span costs more at 3am than an untidy invoice does. The weights sum to 100, so the total runs from 100 to 500 and divides cleanly into an index. Two estates with the same flat average can sit two bands apart once weighting lands.

Zone · 03

Coverage You Have Not Counted

Coverage you have not counted is coverage you do not have. Most teams put instrumentation coverage somewhere above eighty per cent and hold no inventory to check it against. The dimension asks something narrower: is coverage counted, published and chased, and does a service with no telemetry get stopped at the deploy gate.

By the numbers

The figures that make it a board-level conversation.

17%
of total compute infrastructure spend goes on observability at the average organisation, with the median at 10%
52%
of practitioners report high volumes of false alerts, and 59% name disparate tools as a core problem
76%
of organisations now use OpenTelemetry in some form, and 47% raised their investment in it year on year
2.6
median weighted maturity index across the estates we assess, which reads as instrumented but not correlated
Inside the report

What you'll take away.

01

pull the evidence before anyone scores

Your last three postmortems, one month of alert history, the ingest bill broken down as far as it goes, and one trace from a request that crossed a queue. Two hours of collection. Scores argued without that pack are opinions about tooling, and they come out a level high every time.

02

score seven dimensions against a live incident

Instrumentation coverage, semantic conventions, trace completeness, correlation, error budgets, incident workflow, and cost and cardinality. Pick the highest level the evidence pack supports, not the highest your licence permits. A feature you have bought and never used in anger is a level you have not reached.

03

weight the scores and read the band

Level times weight, totalled, divided by 100. The index runs from 1.00 to 5.00 across five bands, and each band names one move. Most estates land between 2.60 and 3.39, which reads as instrumented but not correlated: the data arrives, joining it is still manual.

04

fix the two lowest weighted scores and nothing else

Take the two dimensions costing you most weighted points, then write the smallest change that moves each up one level, with a named owner and a date inside the quarter. Rerun the assessment after the next real incident, using that incident as evidence. A level you cannot re-prove is a level you did not hold.

Questions

Frequently asked.

Do we need OpenTelemetry in place before we can run this?

No, but it will show. The assessment scores outcomes rather than vendors, and a proprietary agent estate reaches level 4 on several dimensions without trouble. What it will not reach is level 5 on attribute hygiene, because enforcement at the collector is where naming stops drifting and queries start joining.

Is this not a scorecard designed to sell us a migration?

No, and the asset argues against one. Open Telemetry is oversold as an exit strategy, because it does not port your dashboards, alert definitions, saved queries or retention model. Fan out a copy of production telemetry to a second backend for two weeks first, then price the lock-in you find into the contract term.

How long does the assessment actually take?

About half a day. Two hours to assemble the evidence pack, ninety minutes to score the seven dimensions with the on-call engineers in the room, and the rest on the weighted total and the two fixes. The argument about which level you can prove is the part that earns its time.

Who should be in the room when we score it?

The on-call rota, mostly. A Head of SRE chairing, two engineers who carried the pager last month, and whoever owns the telemetry bill. Product joins for the error budget dimension, because a budget policy nobody in product agreed to is a document rather than a gate, and it will not hold under pressure.

Our index came out low. Where do we start?

With the band, not the list. Below 1.80 you instrument the three services on the revenue path and get trace ids into their logs. Between 1.80 and 2.59 you enforce service.name, deployment.environment and owner at the collector. One move per band, funded properly, beats seven started andabandoned.

Who is this assessment for?

Heads of SRE and platform leads who have to defend an observability budget. It assumes you already have tooling, probably from several vendors, and that the uncomfortable question this year is not what else to buy but why incident duration has not moved since the last two purchases.

Get the framework

Have it emailed to you.

Drop your details and we'll send 17% Of The Compute Bill Buys Telemetry. Most Of It Answers Nothing. straight to your inbox - no spam, unsubscribe anytime.

Download framework
Next step

Put this into practice.

Book a 30-minute observability review (logiciel.io)

Download the assessment