Source code, in your repository
Pushed to your git remote from the first week, not delivered as a zip at the end. Full IP transfer, no licence-back, no dependency on a platform we own.
Capabilities
Most vendors sell one of these and hope the rest is your problem. Running an AI system in production needs all six, so we do all six — and tell you which ones your project actually requires rather than billing for the full set.
01 · Agentic automation
A chatbot answers. An agent reads the order, checks stock in your WMS, drafts the customer reply, updates the marketplace, and logs what it did. The design question is never "can it do this" — it's "what happens when it's unsure", and that's where most of our work goes.
WHAT YOU GET
One agent per job, not a general assistant. Narrow scope is what makes accuracy measurable and failure predictable.
HOW IT RUNS
The agent calls your APIs the way a trained operator would — read inventory, create a note, issue a label — with permissions scoped per tool.
WHERE IT STOPS
Below the agreed threshold it stops and routes to a person with the context attached, rather than guessing and hoping.
02 · Retrieval & knowledge
Your operating knowledge lives in signed PDFs, a wiki nobody updated, and three people's heads. We turn the first two into something queryable and citable — so an answer can be checked, not just trusted.
INGESTION
Scanned contracts, spreadsheet exports, email threads, product sheets in four languages. Parsing the mess is most of the job.
RETRIEVAL
Keyword and semantic together, because your SKU codes and part numbers don't have useful embeddings.
TRUST
If the system can't point to a source passage, it says it doesn't know. That rule alone removes most of the risk.
03 · Integration engineering
An agent that can't reach your order table is a demo. We treat integration as the primary engineering effort: authentication, rate limits, retries, idempotency, and what happens when a marketplace changes its API without telling anyone.
Shopify, Lazada, Shopee, Amazon, TikTok Shop, WooCommerce; 3PL and courier APIs; customs and freight documents.
SAP, NetSuite, Odoo, custom ERPs, warehouse systems, Zendesk and Freshdesk, WhatsApp Business, accounting exports.
Queues and retries so a platform outage delays work instead of losing it. Idempotent writes so a retry never double-ships.
Contract tests against every external API. When a partner changes a field, we find out from a failing test, not from your customers.
04 · Evaluation & guardrails
Every system we ship has a written accuracy bar agreed before development starts, a test set built from your real historical cases, and a defined behaviour for every way it can be wrong. This is the part vendors skip and clients discover later.
BEFORE BUILD
"95% of address corrections match what a senior clerk would do" is a bar. "It works well" is not.
EVERY CHANGE
Prompts and models are code. Nothing reaches production without the full eval suite passing.
IN PRODUCTION
Live sampling against the eval set. When accuracy slides, you hear it from a dashboard rather than a complaint.
05 · Data engineering
The same product has three SKUs across two systems and a different name in the supplier's file. Until that's resolved, no model will reason correctly about your inventory. We do this work first and say so in the scope.
Matching products, suppliers and customers across systems that were never designed to agree with each other.
Incremental syncs with change data capture, so the agent reasons over this morning's stock rather than last week's export.
Assertions on the data itself. A silent schema change upstream shouldn't become a wrong answer downstream.
Sometimes cleaning the data delivers most of the value and the model is unnecessary. We'll tell you when that's the case.
How it fits together
Half the build is the integration layer, and that is where deployments actually fail. This is the shape of every system we ship — drawn so your engineers can argue with it before we start.
06 · Run & improve
Handing over a system and disappearing is how AI deployments quietly stop working. We stay on monthly terms you can cancel — which keeps the incentive on making it worth keeping.
STAGE 01
Process observation, feasibility, effort and expected savings — written down and yours to keep whether or not you continue.
One week · fixed fee
STAGE 02
Shadow mode on real data against the agreed bar. Miss the bar and you stop here — we'd rather lose the project than erode your team's trust.
Two to three weeks
STAGE 03
Escalation paths, logging, an off switch you control, and your team trained during the build rather than handed a manual after.
Two to four weeks
STAGE 04
Monitoring, retraining on drift, cost tuning, monthly reporting — or a clean handover to your team when you're ready.
Monthly · cancel anytime
Technology
We pick per problem rather than per fashion, and we tell you what runs where. Nothing here is proprietary to us — you can hire someone else to maintain it.
GPT · Claude · Gemini
Open-weight where data must stay in-house
Per-task routing to control cost
pgvector · Qdrant
Hybrid keyword + semantic
Reranking, citation enforcement
PostgreSQL · BigQuery
dbt · Airflow
Change data capture
Python · TypeScript
Docker · Kubernetes
AWS · GCP · Alibaba Cloud
Shopify · Lazada · Shopee
SAP · NetSuite · Odoo
Zendesk · WhatsApp
Eval suites per task
Langfuse · OpenTelemetry
Cost and latency budgets
Rate limits, stale reads, undocumented APIs and a schema that changed without warning. That is where the weeks go, and it is why the integration layer is priced as its own capability.
Deliverables
Not "a solution". These are the artefacts that end up in your repository, your drive and your team's hands, and they are the same list on every engagement.
Pushed to your git remote from the first week, not delivered as a zip at the end. Full IP transfer, no licence-back, no dependency on a platform we own.
The test set, the scoring harness and the calibration curves. Clients tell us this is the piece they would never have built themselves and the one that keeps the system honest after we leave.
How to deploy, how to roll back, what each alert means, what to do when the model provider has an outage, and who to call. Written to be used by someone who has never met us.
Not just what it does — why it was built this way, what we rejected, and which decisions we would revisit if volume tripled. The reasoning is what makes the code changeable later.
Accuracy by error type, escalation rate, cost per task, latency and volume. Running in your infrastructure, visible to your team, with alert thresholds you can edit.
One technical, one operational, recorded. Plus thirty days of questions answered at no charge after the final invoice, because a handover that ends at the invoice is not a handover.
Effort and price
These are not quotes. They are the spread of what comparable work has actually cost us to deliver, published so you can sanity-check a budget before spending an hour on a call. The binding number is the one in your scope document at the end of week one, and it has landed outside these bands in both directions.
| Typical build | Indicative band | Monthly run | Main risk | |
|---|---|---|---|---|
| Agentic automation | 5–8 weeks | USD 28k–60k | USD 3k–8k | Upstream systems that lie about state |
| Retrieval & knowledge | 3–6 weeks | USD 18k–40k | USD 2k–5k | Source documents nobody owns or updates |
| Integration engineering | 2–6 weeks | USD 12k–35k | USD 1k–3k | Undocumented APIs and silent schema changes |
| Evaluation & guardrails | 2–3 weeks | USD 9k–20k | Included in run | No agreed definition of a correct answer |
| Data engineering | 3–8 weeks | USD 15k–45k | USD 1k–4k | Scope creeping into a full data programme |
| Run & improve | Ongoing | No build fee | USD 2k–8k | Paying for monitoring nobody reads |
Prices are quoted per system, not per seat or per user, and they do not scale with your headcount. Inference and infrastructure are billed at cost with the provider invoice attached — we do not mark up compute.


First thirty days
So you can tell early whether it is going well. If any of these has not happened by the stated week, something is wrong and you should say so.
W1
Days 1–5
Two days sitting with the team doing the work — the real process, not the documented one. A random sample of 500 real items pulled. By Friday you have a written scope naming what is worth automating, what is not, and the accuracy bar we will be held to.
W2
Days 6–10
Credentials, API access, a running connection to every system involved, and the baseline timed. Nothing intelligent yet. This is the week that determines whether the rest is easy, and it is where most surprises surface.
W3
Days 11–15
A complete but unimpressive version handling the most common case shape, scored against the random sample. You see a real number in week three. It is usually disappointing, and that is exactly what it is for.
W4
Days 16–20
The agent runs on live volume and nobody uses its output. Proposals and human decisions logged in pairs, with reason codes on every disagreement. Weekly readouts start, and fixes ship inside the month rather than after it.
W5+
Day 21 onward
Task by task, as each clears its bar on a random sample. The easiest categories usually go live around week five; the ones touching money often stay in draft-only mode for months, and that is the design working rather than a delay.
Out of scope
Every one of these has been asked for. Saying no here saves a conversation later.
We do not sell an AI readiness assessment or a transformation roadmap. The scoping week produces a scope for a specific system; if the answer is that no system is worth building, you get that in writing and we stop.
We will not help build a redundancy business case. Partly principle, mostly practice: the people who know the process are the ones who make the system work, and they can tell what they are being asked to do.
We are not a body shop and do not place engineers into your team by the day rate. Engagements are scoped to a system and an outcome, because that is the only arrangement where our incentives match yours.
If you already own a tool that does 80% of this, the honest advice is usually to configure it properly rather than replace it. We have talked two clients out of a build in favour of software they were already paying for.
Agents propose refunds, customs declarations and contractual responses; they never submit them. This holds regardless of measured accuracy, and it is a policy rather than a threshold that can be tuned upward.
If two experienced people disagree on the right answer more than about 15% of the time, there is no ground truth, and a system nobody can score is a system nobody can safely improve. We say so and stop.
Source in your repository from week one, a runbook tested by somebody who did not write it, and thirty days of questions answered after the final invoice. A handover that ends at the invoice is not a handover.
Questions
Usually from the person who will have to maintain this after we leave. They are the right questions.
Python services in containers, a Postgres database, a queue, and whichever model API wins the benchmark. Deployed into your cloud account on your existing orchestration where you have one — we have shipped onto Kubernetes, ECS, Cloud Run and a single well-configured VM, and the last of those is more common than people expect.
No. Provider calls sit behind one interface with the prompt, schema and retry policy defined per task, so swapping means changing a config value and re-running the evaluation suite. Three of our deployments have changed provider since launch without touching surrounding code.
Prompts live in the repository, not in a vendor console, and every change goes through review and a regression run. Each logged decision records the prompt hash, model version and retrieval snapshot, so a result from four months ago can be reproduced exactly.
Circuit-break, queue the work, and fall back to the manual path with a visible banner rather than failing silently. Where a task justifies it we configure a secondary provider, but the default is honest degradation: an operator would rather see "automation paused" than a queue that has quietly stopped moving.
Incremental re-indexing triggered by change events or a scheduled diff, never a nightly full rebuild. Every source document is content-hashed so a silent in-place edit raises an alert instead of quietly serving a stale clause — we learned that one in production and it is written up in our field notes.
Yes, and it makes the handover far better. We work in your repository from week one, follow your review conventions, and pair with whoever will own it. Several clients have taken over the second system entirely, with us only reviewing the design.
Conventional unit and integration tests for the deterministic code, plus the evaluation suite for anything a model touches. The two are different things and both ship. The evaluation suite runs on every pull request and blocks merge on regression.
Your secret manager, your keys, issued to us scoped and revocable. We do not hold standing production access — access is granted per change and logged. On exit, revoke the credentials and nothing of ours remains connected to anything of yours.
Next step
The scoping week exists so you can find out what this costs and whether it works before committing to anything larger.