Novycom
Book a call

Practice

Shadow mode: the cheapest month you will ever spend

The agent does the work. Nobody uses the output. It is the highest-value month in any deployment and the first thing clients try to skip.

The agent does the work. Nobody uses the output. It is the highest-value month in any deployment and the first thing clients try to skip.

Every client asks whether the shadow period can be shortened. It is a fair question — you are paying for a system that is deliberately not being used. Here is what that month actually buys, and why we have never regretted insisting on it.

It produces the only accuracy number that means anything

Lab evaluation tells you how a system performs on the cases you thought to collect. Shadow mode tells you how it performs on Tuesday. Those diverge more than anyone expects: in one deployment the held-out test score was 91% and the first shadow fortnight came in at 78%, entirely because the live queue contained a supplier format that had not appeared in three months of historical data.

Nothing was broken. The test set was just narrower than reality, which every test set is. Shadow mode is the only mechanism we know that reliably finds out how much narrower.

It finds the failure shapes, not just the failure rate

A rate tells you how often. Shapes tell you why, and shapes are what you can fix. Running side by side for four weeks typically surfaces three or four distinct clusters — a date format, a category of customer phrasing, a system that returns stale data under load — and each is a specific, closeable gap rather than a vague call for "better prompting".

Roughly two-thirds of what we find in shadow mode gets fixed in the retrieval or integration layer, not in the model. That ratio is the whole argument for the exercise.

It builds calibrated trust in the team

This is the part that does not appear in any proposal and matters most. Teams meeting an automation arrive either credulous or hostile, and neither is useful. Four weeks of watching proposals next to their own decisions produces something better: staff who can tell you which categories the agent is reliable on and which they still want to see.

That knowledge is what makes the go-live uneventful. When the team has already agreed which tasks it trusts, switching on autonomy is an administrative act rather than a leap of faith. We have watched the same deployment be a crisis at one client and a non-event at another, and the difference was three weeks of shadow running.

It costs almost nothing

The engineering is already built — you are paying for compute and a small amount of review time. Against that, the failure mode it prevents is a public mistake in week one that gets the project cancelled, which costs the entire build.

Concretely: on a mid-sized deployment the shadow month typically costs a few hundred dollars of inference and about two hours a week of one person's attention. We have never had a client complete one and say it was not worth it. We have had two who skipped it, and both came back to run one afterwards, from a worse position.

How we run it

  • Everything is logged in pairs. Agent proposal and human decision, on the same record, with the retrieved context attached.
  • Reviewers do not see the proposal first. Otherwise you measure anchoring, not accuracy. The agent's output is revealed after the human commits.
  • Disagreements get a reason code. Thirty seconds from the reviewer, and it is what turns a rate into a shape.
  • Weekly readout, not a final report. Fixes ship during the month. By week four the numbers should be visibly moving, and if they are not, that itself is the finding.
  • Every disagreement becomes a test case. The evaluation set that comes out of shadow mode is worth more than the one that went in.

When it can be shorter

Two weeks is defensible when the task is low-stakes, high-volume and the queue is homogeneous — stock-availability replies, for instance, where a wrong answer is embarrassing rather than expensive and you accumulate a thousand samples in days.

It should be longer than four weeks when volume is low, when the process is seasonal, or when a single error is genuinely costly — anything touching refunds, pricing, customs paperwork or a regulated claim. For one client handling export documentation we ran shadow mode for eleven weeks. They were right to insist, and we said so.

Keep reading

Recognise the process?

Tell us about it in thirty minutes.

Bring one process that is eating your team's week. We will tell you on the call whether it is a good candidate — and we say no to roughly a third of what we are asked about.