Case Study: Warehouse Ops Agent

Why I built this
I spent seven years in e-commerce, three of them building internal tools, and Iâm fascinated by where AI can actually fit into that work. Not replacing the people on the floor, but taking on the digging and cross-checking that nobody ever has time for. The Cycle Count Investigator specifically is the kind of tool that wouldâve saved me hundreds of hours at my last job, no exaggeration.
Itâs also a learning project, if Iâm being honest. I built it to get my hands dirty with the parts of AI engineering that actually matter once somethingâs in production: tool use, human-in-the-loop approval, MCP, and evals that measure whether a model actually gets things right instead of just sounding like it does. Still, I tried to design it around real warehouse problems and real constraints (source-location rules, quantity limits, duplicate-task checks) so a version of this could plausibly help run a warehouse someday.
Demo
A 90-second walkthrough: a supervisor asks about short picks and approves the replenishment tasks the agent proposes, a simulated cycle count kicks off an AI investigation, and the agent log shows every call it made.
The problem, in two parts
Short picks. A picker goes to a pick face and finds fewer units than the task expects. Usually the stock is just sitting on a pallet in reserve and nobody refilled the face in time, but the supervisor finds out late and has to piece it together from a few different screens.
Inventory drift. The systemâs numbers and whatâs actually on the floor slowly drift apart. A manual adjustment gets entered with no reason. A pallet gets put away one slot over. Someone counts a pallet in cases instead of units. A clerk cycle-counting a location logs the number and moves on, and nobody has time to chase down every mismatch, so the same errors keep coming back month after month.
What it does
Warehouse Ops Agent answers both kinds of questions from live data, using four tools against the warehouse database, and itâll tell you straight up when the data just isnât there. It can look up short picks, find stock by SKU, list what needs replenishing, and propose a replenishment task.
That last part is where I spent most of my design energy, actually. It can never approve its own proposal. Approving or rejecting a task only exists in the UI/API, never as a tool the model can call, so thereâs no path where the agent both suggests and executes a warehouse change on its own. The real business rules (valid source locations, quantity limits, no duplicate open tasks) live in a tested service layer instead of the prompt. The agent can misread a request. It canât talk its way around a rule.
The floor map shows the same data spatially. Every pick face is a cell shaded by stock level, outlined when it has an open task or when the agent is working there, and marked when it has an open count discrepancy. Selecting a bay shows its on-hand quantity, min/max, and the reserve pallets above it.

Thereâs also a second, cheaper model running entirely separately as the Cycle Count Investigator. Whenever a simulated cycle count opens a discrepancy, this background agent reads the inventory ledger, nearby slots, unconfirmed picks, and other open discrepancies, then comes back with a ranked list of likely causes and the evidence behind them, something like â+30 adjustment by Michael Mcguire, no reason given.â A supervisor still decides whether to accept the count or ask for a recount. Thatâs a click, not a tool call either.

Everything the agents do is logged, too. The agent log page shows every Claude request and tool call with its arguments, results, latency, and cost, so when an investigation says something surprising I can open it up and see exactly which lookups it based that on.

Architecture
Business logic lives in one place (services/), and the MCP tools, the agent, and the API are all thin wrappers around it, so each rule only gets written and tested once instead of three separate times.
I hand-wrote the agent loop instead of using the Anthropic SDKâs built-in tool runner for one reason: the human-approval pause can span multiple HTTP requests. A read-only question resolves in a single round trip, but a write call has to stop the loop mid-conversation and wait, sometimes minutes, for a person to click approve or reject. The same tool layer also runs as an MCP server, so the exact same logic works from Claude Desktop.
Evals: does the agent actually get it right
Unit tests cover the service layer fine, but they canât tell you whether Claude picks the right tool, passes the right arguments, and answers correctly when itâs the one actually deciding what to do. Thatâs what the evals are for. Two YAML-defined suites run the live agent against the Claude API over the same deterministic seeded warehouse, graded by code rather than an LLM judge, against ground-truth values pulled from the seed data at run time so a case doesnât go stale as the seed changes.
Chat agent, 15 cases
Each case checks up to four things: did it call the right tool and skip the wrong ones, did it pass the right arguments, does the answer have the right facts in it, and did the right database write actually happen. A handful of cases are there specifically to check what the agent shouldnât do: refusing to approve its own task (cannot-self-approve), refusing a ridiculous 100,000-unit move (create-task-over-max-refused), and making zero tool calls for an off-topic question about the weather (out-of-scope).
Results from one run per model (2026-10-02):
| Model | Passed | Tool selection | Arguments | Answers | Outcome | Cost (15 cases) | Mean latency |
|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | 15 / 15 | 100% | 100% | 100% | 100% | $0.09 | 3.4 s |
| Claude Sonnet 5.5 | 15 / 15 | 100% | 100% | 100% | 100% | $0.22 | 4.7 s |
| Claude Opus 5.5 | 14 / 15 | 93% | 100% | 100% | 100% | $0.52 | 8.4 s |
Cost and speed scale with model size, but accuracy really didnât here. Haiku matched Sonnet at about 40% of the cost and ran more than twice as fast as Opus. And Opusâs one miss wasnât even a safety problem â asked to âapprove task #1,â it correctly refused and said no such task exists, but it made a read call first just to check. The case technically wants zero tool calls, which is a stricter bar than the rule that actually matters: the agent never approves or writes on its own.
The evals caught a real bug before I ever wouldâve noticed it by hand, too. On an early Haiku run, the agent proposed a duplicate replenishment task because find_stock wasnât surfacing that one was already open. The service layer wouldâve refused the write anyway, but the proposal itself was wasted effort. I added open_task_id to what find_stock returns and that bumped tool-selection accuracy from 93% to 100%.
Cycle Count Investigator, 8 cases
This one checks something different: does the investigator actually find the true cause of each of the eight planted discrepancies, not just whether it calls a plausible-looking tool. Each case scores the top-ranked cause, whether the right cause lands in the top 3, the evidence it cites, the tools it used, and the next steps it suggests. One case has no real explanation in the data at all, so the right move there is admitting the units are just missing instead of making something up.
Results from one run per model (2026-10-04):
| Investigator model | Correct root cause | Evidence cited | Right tools | Cost (8 investigations) | Mean latency |
|---|---|---|---|---|---|
| Claude Haiku 4.5 | 7 / 8 | 100% | 100% | $0.16 | 13 s |
| Claude Opus 5.5 | 8 / 8 | 100% | 100% | $0.94 | 14 s |
Haikuâs the value pick here, roughly 2 cents per investigation against Opusâs 12 cents. Its one miss was a pallet counted in cases instead of units â Haiku guessed stock had been removed and got the math wrong, while Opus actually caught the miscount and suggested checking the same clerkâs other counts.
The evals caught mistakes in themselves too, which I didnât expect going in. The first run surfaced an accidental red herring Iâd planted in the seed data (an unconfirmed pick of 11 sitting right next to a shortage of 12), so I fixed that. A later run showed the keyword-based answer checks failing correct answers just because they said ârefillâ instead of âreplenishment,â so I loosened the matching and added a --rescore flag that regrades a saved report against updated cases for free, no API calls needed.
Caveats, for real
These are single runs per model, not averages, so thereâs run-to-run noise in here I havenât actually measured. Answer scoring is keyword-based instead of an LLM judge, which is cheap and deterministic but will occasionally dock a correct answer for unexpected wording. An LLM-as-judge pass and a bigger case set (Iâm aiming for 30-50) are next on the list for the eval harness itself.
Stack
Python 3.12, FastAPI, SQLModel. MCP Python SDK. Anthropic Python SDK. React, Vite, TypeScript, Tailwind, shadcn/ui. uv, ruff, mypy, pytest, Vitest.
Full source, including the eval cases and scoring code above, on GitHub.
Button not opening your mail app? Write to scottpeters2281@gmail.com.