Logiciel Solutions Contact Us
Success Stories Tech News Investors Contact Us
whitepaper

MLOps and Model Serving: What Buyers Should Ask.

A quoted p50 is one point on a curve, and the seller chose which point. Continuous batching makes latency and throughput a single dial, so a figure published without its concurrency, batch size, token shape and accelerator model cannot be placed on that curve at all.

You are buying capacity for teams who are not in the room. This playbook sets sixteen questions across nine areas beside the weak answer and the strong one, so the operating point, the idle accelerator and the checkpoint calendar are settled on paper rather than discovered in production.

In depth

Three Faults That Make A Serving Evaluation Pleasant And Useless.

01

Latency and throughput are one dial, and sellers set it for the slide.

Hold concurrency near zero and a single request owns the accelerator, so p50 looks superb while the hardware sits mostly idle; raise the batch size and throughput climbs with queue wait and tail latency behind it. Both measurements are honest, and only one of them describes a cluster carrying your traffic. Ask at what concurrency the number was taken, and how much of it is queue wait rather than compute.

In shorthow much of it is queue wait rather than compute
02

The price list describes flat traffic and nobody has flat traffic.

Stanford's AI Index records inference cost falling roughly 280 times between November 2022 and October 2024, which makes whatever rate you negotiate this quarter a temporary fact while the shape of the contract outlives it. Reserved accelerator capacity is paid for whether or not it computes anything, so somebody funds the idle hours between your peaks. Establish early whether that somebody is you, and what stops being billed when traffic falls to a tenth overnight.

In shortEstablish early whether that somebody is you, and wh…
03

A platform multiplies whatever the provider does underneath it.

A serving tier carrying thirty teams turns one supplier decision into thirty simultaneous events: a checkpoint withdrawal, a runtime upgrade, a change in default decoding behaviour, none of which generates a release in your own pipeline for anyone to review. On a sample of nearly 5,000 practitioners, DORA found AI adoption lifting organisational performance strongly where the platform underneath it is good, and barely at all where it is not (DORA, 2025). Version pinning and the notice period attached to it belong in the contract rather than in the release notes.

In shortVersion pinning and the notice period attached to it…
The detail

Three Answers Worth More Than The Whole Technical Deep Dive.

Nine areas carry the sixteen questions and three of them decide the rest. Send these before the deep dive and keep what comes back, then ask for the artefact behind any strong answer rather than the sentence that delivered it.

Zone · 01

The operating point

Ask for a curve rather than a figure: p50, p95 and p99 across stated concurrency levels and batch sizes, on a named accelerator and a pinned inference server. A seller who separates queue wait from compute has run a loaded test. A seller who offers to divide the figure by replica count has run a spreadsheet.

Zone · 02

The idle accelerator

Either dedicated and yours to fill, or pooled with a named isolation guarantee, a per-tenant floor and stated quota behaviour when a neighbour surges. A provider pooling spare capacity across tenants is making a reasonable offer and should describe what makes it safe. One charging for peak-shaped provisioning while pooling it anyway is selling the same hour twice.

Zone · 03

Pinning and withdrawal

Ask whether you can pin an immutable checkpoint identifier, how long you may stay on it, and what arrives in your inbox before it is withdrawn. A workable answer contains a notice window counted in months, both versions addressable at once so migration is a routing change, and an evaluation delta they publish rather than you discover.

By the numbers

The figures that make it a board-level conversation.

16 questions
grouped into nine areas, sent before the technical deep dive rather than after it
280 times
fall in AI inference cost between November 2022 and October 2024, so buy structure rather than a rate
5,000
technology professionals surveyed, and where platform quality is low the effect of AI adoption is negligible
Inside the report

What you'll take away.

01

Step 1 - Replay a recorded peak hour at its real shape

Take the worst hour your gateway logs show that month and replay it with bursts and gaps intact. A constant request rate measures a constant request rate and nothing you will ever serve.

02

Step 2 - Fix the operating point before anybody tunes anything

State concurrency, the prompt and output length distribution and the p99 target in the scope note. Then make each vendor publish every parameter they changed, including the ones that made the result worse.

03

Step 3 - Let the cluster go quiet, then time the next request

Idle it for an hour, send one request and record image pull, weights load and cache warm separately. Repeat at three in the morning, when the scale-to-zero policy is actually exercised.

04

Step 4 - Export the artefact mid-trial and serve it elsewhere

Take out weights, tokeniser, configuration and registry records while the relationship is cordial, then load them on a different runtime. An export nobody has executed is a clause, not a capability.

Questions

Frequently asked.

Why does one quoted latency figure tell us so little?

It was measured at an operating point the seller picked, and the operating point never reaches the quote. Without the concurrency, the batch size, the prompt and output length distribution and the accelerator part number, the figure cannot be placed on the curve, so it is true somewhere and useless everywhere else.

How do we compare pricing when rates keep moving?

Compare structure, not rate. Inference cost fell roughly 280 times in two years, so the number you negotiate this quarter dates quickly while the contract shape outlives it. Ask what stops being billed when traffic drops to a tenth, and what is admitted when it multiplies for ninety seconds.

What exactly should we demand on checkpoint pinning?

An immutable version identifier you may serve for a contractual window, a deprecation calendar published a release ahead of withdrawal, both versions addressable at once so migration is a routing change, and an evaluation delta they publish. An alias resolving to the current best version is not pinning.

Is a bake-off worth the fortnight it costs?

It settles in two weeks what procurement argues about for a quarter, but only where the operating point is fixed in writing before either vendor touches a dial. Replay a real peak hour with its arrival pattern intact, and run the export yourself rather than reading the clause that describes it.

How does this differ from your scorecard on the same topic?

This playbook is for a platform you have not bought yet, and it works through what a seller will and will not put in writing. Benchmarking MLOps and Model Serving is the instrument for scoring a platform you already run, domain by domain, against the artefact each row needs.

Who is this playbook written for?

VPs of platform and heads of ML infrastructure buying serving capacity that other teams' features will sit on by the end of the year. It assumes you can read a data sheet and now need the questions that decide whether the data sheet means anything.

Get the whitepaper

Have it emailed to you.

Drop your details and we'll send MLOps and Model Serving: What Buyers Should Ask straight to your inbox - no spam, unsubscribe anytime.

Download whitepaper
Next step

Find out what the quoted latency becomes at your concurrency.

Bring the serving contract you are about to sign and the worst hour your gateway recorded. A working session with our engineering leads, who will replay one against the other rather than present to you. SECTION 7 - FAQ - 5 to 8 questions

Talk to our engineers