Case 01
Order exception desk, six markets
Before. Three people worked a shared inbox from 9am, and the queue was never empty by 6pm. Every case meant opening the marketplace back office, the WMS and the courier portal in separate tabs and reconciling three versions of the truth by eye.
What we built. An agent that classifies each incoming case into one of nine shapes, pulls order, inventory and shipment state through existing APIs, and either executes the resolution or routes it with all three states already assembled on one screen.
What nearly broke it. The WMS returned stale stock for up to ninety seconds after a write. In shadow mode the agent confidently told twelve customers an item was available that had just been picked. We added a read-after-write delay and a stock re-check immediately before any promise to a customer — a fix in the integration layer, not the model.
Would do differently. We scoped six markets at once. Three would have been better: the two smallest markets contributed 4% of volume and about a third of the edge cases.
11 weeks · 4,100 cases/mo · 78% unattended · 3 people redeployed, none let go
Case 02
Catalogue enrichment, four languages
Before. 18,000 SKUs, attributes entered by six people over four years, and descriptions that were often one line copied from a supplier PDF. Three market launches were blocked on listing quality, and the manual estimate to fix it was fourteen months.
What we built. Attribute normalisation against the client's own taxonomy — not a generic one — plus per-market copy written to each marketplace's rules, and a review queue where a merchandiser approves in batches of fifty.
What nearly broke it. The first pass wrote fluent English copy and translated it. The Indonesian team rejected 40% of it as technically correct and commercially dead. We rebuilt the pipeline to generate per market from the attributes directly, with a native reviewer in the loop for the first thousand SKUs of each language.
Would do differently. Involve the market teams in week one rather than week five. The rejection was avoidable and it cost us a fortnight we did not bill for.
9 weeks · 18k SKUs · 3 markets live · rejection rate 40% → 6%
Case 03
Marketplace payout reconciliation
Before. Six platforms, six settlement formats, two spreadsheets and three days of one finance person's month. Discrepancies were found late or not at all, and one platform's fee change went unnoticed for two billing cycles.
What we built. A nightly matching run across orders, fees, refunds and payouts, with unmatched residue isolated into a queue rather than absorbed into a total. Fee-structure changes now surface as an alert on the first settlement that does not reconcile.
What nearly broke it. Very little, and that is the point — most of this is deterministic matching, not AI. The model does one narrow job: reading the free-text adjustment lines that four of the six platforms use. We said so in the scoping week and the build was priced accordingly.
Would do differently. Nothing significant. This is the shape of project that works: bounded, verifiable, and honest about which part actually needs a model.
5 weeks · 3 days/month → 2 hours · one fee change caught in month two
Case 04
The one we stopped
The ask. A distributor wanted an agent to decide which backorders to prioritise when stock arrived short — a judgement made daily by two senior people balancing customer relationship, margin and contractual penalty.
What we found. In the scoping week we gave the same fifty historical cases to both decision-makers independently. They agreed 61% of the time. There was no ground truth to build an evaluation set from, because the organisation did not have a shared policy — it had two people with different instincts.
What we did. Stopped, and said why. We spent the remaining two days of the scoping week facilitating a session that produced a one-page prioritisation policy. No build followed, and the fixed scoping fee was the entire engagement.
Why it is on this page. They came back eleven months later with a different process, and that one shipped. We would rather show this than pretend our hit rate is 100%.
1 week · no build · policy document delivered · client returned