Practice
The agent does the work. Nobody uses the output. It is the highest-value month in any deployment and the first thing clients try to skip.
Every client asks whether the shadow period can be shortened. It is a fair question — you are paying for a system that is deliberately not being used. Here is what that month actually buys, and why we have never regretted insisting on it.
Lab evaluation tells you how a system performs on the cases you thought to collect. Shadow mode tells you how it performs on Tuesday. Those diverge more than anyone expects: in one deployment the held-out test score was 91% and the first shadow fortnight came in at 78%, entirely because the live queue contained a supplier format that had not appeared in three months of historical data.
Nothing was broken. The test set was just narrower than reality, which every test set is. Shadow mode is the only mechanism we know that reliably finds out how much narrower.
A rate tells you how often. Shapes tell you why, and shapes are what you can fix. Running side by side for four weeks typically surfaces three or four distinct clusters — a date format, a category of customer phrasing, a system that returns stale data under load — and each is a specific, closeable gap rather than a vague call for "better prompting".
Roughly two-thirds of what we find in shadow mode gets fixed in the retrieval or integration layer, not in the model. That ratio is the whole argument for the exercise.
This is the part that does not appear in any proposal and matters most. Teams meeting an automation arrive either credulous or hostile, and neither is useful. Four weeks of watching proposals next to their own decisions produces something better: staff who can tell you which categories the agent is reliable on and which they still want to see.
That knowledge is what makes the go-live uneventful. When the team has already agreed which tasks it trusts, switching on autonomy is an administrative act rather than a leap of faith. We have watched the same deployment be a crisis at one client and a non-event at another, and the difference was three weeks of shadow running.
The engineering is already built — you are paying for compute and a small amount of review time. Against that, the failure mode it prevents is a public mistake in week one that gets the project cancelled, which costs the entire build.
Concretely: on a mid-sized deployment the shadow month typically costs a few hundred dollars of inference and about two hours a week of one person's attention. We have never had a client complete one and say it was not worth it. We have had two who skipped it, and both came back to run one afterwards, from a worse position.
Two weeks is defensible when the task is low-stakes, high-volume and the queue is homogeneous — stock-availability replies, for instance, where a wrong answer is embarrassing rather than expensive and you accumulate a thousand samples in days.
It should be longer than four weeks when volume is low, when the process is seasonal, or when a single error is genuinely costly — anything touching refunds, pricing, customs paperwork or a regulated claim. For one client handling export documentation we ran shadow mode for eleven weeks. They were right to insist, and we said so.
Recognise the process?
Bring one process that is eating your team's week. We will tell you on the call whether it is a good candidate — and we say no to roughly a third of what we are asked about.