← Back

Capabilities

Cut your model spend without gambling on quality

The reason nobody moves to a cheaper model is not that cheaper models are bad. It is that no engineer will stake their name on one without evidence from your own traffic.

Your AI bill grows because switching to something cheaper requires somebody to promise the quality will hold. Nobody will promise that on the strength of a public leaderboard, so the expensive model stays, quarter after quarter.

What it changes for the business

  • The saving becomes defensible. You can show what a cheaper model would have produced for your own requests, so the decision is evidence rather than nerve.
  • Nobody has to volunteer to be wrong. The engineer who would have carried the risk personally is handed a comparison instead.
  • Customers are never the experiment. The evaluation runs alongside live traffic, and every customer keeps getting the answer they were always getting.
  • The review is repeatable. When a new model launches, the same comparison runs again, so this stops being a project and becomes a habit.

How it works, briefly

A small share of an application's real requests is quietly run against one or more alternative models in the background. The live answer is still produced by the model you are using today, so nobody waits and nothing changes for the customer. Each alternative is then compared on cost, accuracy and latency.

Shadow run · support-assistant

7 days · 4,180 sampled requests · no customer saw an alternate answer.

Observe only
ModelAccuracyCost / 1kp95
claude-sonnet-4.5-$12.401,180 msIn use today
gpt-4.1-mini96.2%$3.60640 msCandidate
llama-3.3-70b88.1%$1.10910 msNeeds review
Fig 07ACost, accuracy and latency, on your own traffic.
How this is handled elsewhere How Squidder does it
Published benchmarks Your own traffic, your own prompts
A handful of prompts someone tried by hand A representative sample over a real period
Split live traffic and watch complaints Nothing customer-facing is exposed
Switch and hope, then roll back Decide before anyone is affected
Cost compared separately from quality Cost, accuracy and latency in one comparison

Graded not equivalent

159 of 4,180. The article tells the reader to read these rather than the score.

159 cases
RequestHow the answers differedRead as
rq_4b17ee02Same conclusion, different order of reasonsAcceptable
rq_9c02af41Omitted the excess clause from the summaryMaterial
rq_1f88d730Shorter, and dropped a caveat that was not load-bearingAcceptable
rq_77a3b195Quoted a refund window of 14 days instead of 30Material
Fig 07BThe cases the grader marked as different.

Where it fits, and where it does not

This covers model conversations. Agent tool calls and ordinary web traffic have no equivalent answer to compare, so they are not part of it. Comparing two answers for equivalence is a judgement rather than a proof, so read the cases marked as different, not only the headline score.

Where to start

Choose one application with steady, unremarkable traffic. Run its current model against a cheaper one for a week and read the disagreements. If they are all acceptable, you have the evidence, and it came from your users rather than a vendor's chart.