Skip to content
🟱 Looking for work! Open to Software Engineering and AI Engineering roles, so let's talk.
Scott Peters

Case Study: Warehouse Ops Agent

By Scott Peters on Oct 3, 2026
The Warehouse Ops Agent finding 14 short picks in zone A, then proposing two replenishment tasks that wait for a supervisor to approve or reject them.

Why I built this

I spent seven years in e-commerce, three of them building internal tools, and I’m fascinated by where AI can actually fit into that work. Not replacing the people on the floor, but taking on the digging and cross-checking that nobody ever has time for. The Cycle Count Investigator specifically is the kind of tool that would’ve saved me hundreds of hours at my last job, no exaggeration.

It’s also a learning project, if I’m being honest. I built it to get my hands dirty with the parts of AI engineering that actually matter once something’s in production: tool use, human-in-the-loop approval, MCP, and evals that measure whether a model actually gets things right instead of just sounding like it does. Still, I tried to design it around real warehouse problems and real constraints (source-location rules, quantity limits, duplicate-task checks) so a version of this could plausibly help run a warehouse someday.

Demo

A 90-second walkthrough: a supervisor asks about short picks and approves the replenishment tasks the agent proposes, a simulated cycle count kicks off an AI investigation, and the agent log shows every call it made.

The problem, in two parts

Short picks. A picker goes to a pick face and finds fewer units than the task expects. Usually the stock is just sitting on a pallet in reserve and nobody refilled the face in time, but the supervisor finds out late and has to piece it together from a few different screens.

Inventory drift. The system’s numbers and what’s actually on the floor slowly drift apart. A manual adjustment gets entered with no reason. A pallet gets put away one slot over. Someone counts a pallet in cases instead of units. A clerk cycle-counting a location logs the number and moves on, and nobody has time to chase down every mismatch, so the same errors keep coming back month after month.

What it does

Warehouse Ops Agent answers both kinds of questions from live data, using four tools against the warehouse database, and it’ll tell you straight up when the data just isn’t there. It can look up short picks, find stock by SKU, list what needs replenishing, and propose a replenishment task.

That last part is where I spent most of my design energy, actually. It can never approve its own proposal. Approving or rejecting a task only exists in the UI/API, never as a tool the model can call, so there’s no path where the agent both suggests and executes a warehouse change on its own. The real business rules (valid source locations, quantity limits, no duplicate open tasks) live in a tested service layer instead of the prompt. The agent can misread a request. It can’t talk its way around a rule.

The floor map shows the same data spatially. Every pick face is a cell shaded by stock level, outlined when it has an open task or when the agent is working there, and marked when it has an open count discrepancy. Selecting a bay shows its on-hand quantity, min/max, and the reserve pallets above it.

The Floor map page: zones A, B and C laid out aisle by aisle, each pick face shaded from empty to full, with bay B-01-27-1 selected showing 9 units on hand (min 3, max 12) and two reserve pallets above it.

There’s also a second, cheaper model running entirely separately as the Cycle Count Investigator. Whenever a simulated cycle count opens a discrepancy, this background agent reads the inventory ledger, nearby slots, unconfirmed picks, and other open discrepancies, then comes back with a ranked list of likely causes and the evidence behind them, something like “+30 adjustment by Michael Mcguire, no reason given.” A supervisor still decides whether to accept the count or ask for a recount. That’s a click, not a tool call either.

The Cycle counts page: an open discrepancy at A-03-03-2 (system 348, counted 29) with an AI investigation concluding the counter most likely recorded 29 cases of 12 instead of 348 units, ranked Likely and Unlikely causes with evidence from the ledger, next steps, and Accept count and Recount buttons. A second discrepancy below traces 1,104 units to a pallet put away one bay over.

Everything the agents do is logged, too. The agent log page shows every Claude request and tool call with its arguments, results, latency, and cost, so when an investigation says something surprising I can open it up and see exactly which lookups it based that on.

The Agent log page: total cost of $0.3859 across 25 Claude requests and 39 tool calls, with the chat session from the top of this post expanded to show list_short_picks, list_replenishment_needs, and two create_replenishment_task calls marked Approved, including the arguments and result of the first.

Architecture

Warehouse Ops Agent architecture: a React chat and floor map call FastAPI, which runs the agent loop; the agent loop and an MCP server both talk to a shared services layer; a background Investigator reads from that same layer after a cycle count flags a discrepancy; everything ends up in SQLModel/SQLite.

Business logic lives in one place (services/), and the MCP tools, the agent, and the API are all thin wrappers around it, so each rule only gets written and tested once instead of three separate times.

I hand-wrote the agent loop instead of using the Anthropic SDK’s built-in tool runner for one reason: the human-approval pause can span multiple HTTP requests. A read-only question resolves in a single round trip, but a write call has to stop the loop mid-conversation and wait, sometimes minutes, for a person to click approve or reject. The same tool layer also runs as an MCP server, so the exact same logic works from Claude Desktop.

Evals: does the agent actually get it right

Unit tests cover the service layer fine, but they can’t tell you whether Claude picks the right tool, passes the right arguments, and answers correctly when it’s the one actually deciding what to do. That’s what the evals are for. Two YAML-defined suites run the live agent against the Claude API over the same deterministic seeded warehouse, graded by code rather than an LLM judge, against ground-truth values pulled from the seed data at run time so a case doesn’t go stale as the seed changes.

Chat agent, 15 cases

Each case checks up to four things: did it call the right tool and skip the wrong ones, did it pass the right arguments, does the answer have the right facts in it, and did the right database write actually happen. A handful of cases are there specifically to check what the agent shouldn’t do: refusing to approve its own task (cannot-self-approve), refusing a ridiculous 100,000-unit move (create-task-over-max-refused), and making zero tool calls for an off-topic question about the weather (out-of-scope).

Results from one run per model (2026-10-02):

ModelPassedTool selectionArgumentsAnswersOutcomeCost (15 cases)Mean latency
Claude Haiku 4.515 / 15100%100%100%100%$0.093.4 s
Claude Sonnet 5.515 / 15100%100%100%100%$0.224.7 s
Claude Opus 5.514 / 1593%100%100%100%$0.528.4 s

Cost and speed scale with model size, but accuracy really didn’t here. Haiku matched Sonnet at about 40% of the cost and ran more than twice as fast as Opus. And Opus’s one miss wasn’t even a safety problem — asked to “approve task #1,” it correctly refused and said no such task exists, but it made a read call first just to check. The case technically wants zero tool calls, which is a stricter bar than the rule that actually matters: the agent never approves or writes on its own.

The evals caught a real bug before I ever would’ve noticed it by hand, too. On an early Haiku run, the agent proposed a duplicate replenishment task because find_stock wasn’t surfacing that one was already open. The service layer would’ve refused the write anyway, but the proposal itself was wasted effort. I added open_task_id to what find_stock returns and that bumped tool-selection accuracy from 93% to 100%.

Cycle Count Investigator, 8 cases

This one checks something different: does the investigator actually find the true cause of each of the eight planted discrepancies, not just whether it calls a plausible-looking tool. Each case scores the top-ranked cause, whether the right cause lands in the top 3, the evidence it cites, the tools it used, and the next steps it suggests. One case has no real explanation in the data at all, so the right move there is admitting the units are just missing instead of making something up.

Results from one run per model (2026-10-04):

Investigator modelCorrect root causeEvidence citedRight toolsCost (8 investigations)Mean latency
Claude Haiku 4.57 / 8100%100%$0.1613 s
Claude Opus 5.58 / 8100%100%$0.9414 s

Haiku’s the value pick here, roughly 2 cents per investigation against Opus’s 12 cents. Its one miss was a pallet counted in cases instead of units — Haiku guessed stock had been removed and got the math wrong, while Opus actually caught the miscount and suggested checking the same clerk’s other counts.

The evals caught mistakes in themselves too, which I didn’t expect going in. The first run surfaced an accidental red herring I’d planted in the seed data (an unconfirmed pick of 11 sitting right next to a shortage of 12), so I fixed that. A later run showed the keyword-based answer checks failing correct answers just because they said “refill” instead of “replenishment,” so I loosened the matching and added a --rescore flag that regrades a saved report against updated cases for free, no API calls needed.

Caveats, for real

These are single runs per model, not averages, so there’s run-to-run noise in here I haven’t actually measured. Answer scoring is keyword-based instead of an LLM judge, which is cheap and deterministic but will occasionally dock a correct answer for unexpected wording. An LLM-as-judge pass and a bigger case set (I’m aiming for 30-50) are next on the list for the eval harness itself.

Stack

Python 3.12, FastAPI, SQLModel. MCP Python SDK. Anthropic Python SDK. React, Vite, TypeScript, Tailwind, shadcn/ui. uv, ruff, mypy, pytest, Vitest.

Full source, including the eval cases and scoring code above, on GitHub.

Let's work together
I'm looking for Business Systems Analyst, Programmer Analyst, and similar roles. Send me an email.

Button not opening your mail app? Write to scottpeters2281@gmail.com.

© Copyright 2026 by Scott Peters. Theme by CreativeDesignsGuru.