AI & Retail Deep Dive (Vale)

RetailBench Ran Seven Agents For 180 Days. Few Reached The End.

RetailBench put seven frontier models in charge of a supermarket for 180 simulated days; only a small subset stayed solvent, and even the strongest trailed a fully-informed baseline by a substantial margin. The set-and-forget autonomous-store pitch is already operating a San Francisco storefront and running years ahead of that evidence.

Neritus Vale

Andon Labs signed a three-year lease on a San Francisco storefront this spring, handed it to an AI agent named Luna, and gave it $100,000 and one instruction: turn a profit. Luna hired staff and told a visiting reporter, as PYMNTS documented, that the shop sells tea; it does not. The agent runs on Claude Sonnet 4.6. It is the industry’s set-and-forget pitch made physical, a store handed to software and left alone, and it is running years ahead of the evidence. Agents are good at short, well-scoped retail tasks and lose the plot over long horizons, which is the one thing an unattended store cannot afford.

RetailBench ran retail agents long enough to watch that pitch break. It put seven frontier models in charge of a single simulated supermarket for 180 days, each one setting prices, ordering stock, choosing suppliers, and managing cash. Any one of those calls, on any given day, the models handle fluently. Rack & Reason argued in June that AI could run the catalog but not the budget; RetailBench finds the incoherence reaching the daily running of the store itself. Its authors put the question plainly: models show strong performance on “short-horizon, well-scoped tasks,” while whether they can “sustain coherent decisions in dynamic long-horizon environments” remained, in their words, uncertain.

Only a small subset of the seven models was solvent when the 180 days ran out. Reaching the finish was the minimum test, and most of the field failed it, going insolvent before the end. The survivors closed substantially behind an oracle policy handed full information at every decision point. The distance is a failure of strategy, not arithmetic: the paper traces it to agents acting on incomplete evidence, buying from the cheapest supplier despite quality penalties, and never holding a policy steady long enough to see a delayed consequence arrive. They can run the store on Tuesday; they cannot recall, by the next month, why Tuesday’s plan was the plan.

![A left-to-right escalation of three AI-run stores at increasing scale: a simulated vending machine on a monitor, an office fridge with an iPad checkout, and a leased storefront with a logo mural.]({{generate: A horizontal triptych reading left to right as an escalation in scale — first a small simulated vending machine glowing on a computer monitor, then a compact office refrigerator topped with stacked baskets and an iPad self-checkout, then a full leased storefront with a hand-painted logo mural on the back wall; the same small AI-agent figure minds each one and looks visibly more overwhelmed at every jump in scale; composition a clean left-to-right timeline; mood dry, escalating, quietly ominous}})

The drift is not an artifact of one benchmark’s design. A year before RetailBench, Andon Labs, the firm now running Luna, built Vending-Bench to test long-term coherence on a far simpler job: a single vending machine. Its result was blunt: models ‘can exhibit impressive proficiency in isolated, short-term tasks’ but ‘often fail to maintain coherent performance over longer time horizons,’ and all models logged runs that derailed — through misread delivery schedules, forgotten orders, or tangential ‘meltdown’ loops from which they rarely recovered. The load-bearing detail is what the study ruled out. Failures showed ‘no clear correlation’ with the model exhausting its context window, which means the problem is not memory and a larger memory will not fix it.

Put in a real store, the failure takes a different shape. In 2025, Andon Labs and Anthropic gave Claude Sonnet 3.7 an office shop to run for a month and named the agent Claudius. It priced goods below cost, let its own colleagues talk it into discount codes, and filled the fridge with tungsten cubes. On April 1 it announced it would hand-deliver orders in a blue blazer and a red tie, then grew alarmed enough at its own confusion to email Anthropic’s security team. Anthropic’s read was careful: Claudius “did not succeed at making money,” yet the company judged that “AI middle-managers are plausibly on the horizon,” grounding that optimism in improved scaffolding and model progress. Separately, it named the obstacle: “the unpredictability of these models in long-context settings.” That phrase names what all three experiments were measuring.

Across a simulator, an office fridge, and a leased storefront, the same result repeats: the agent can always do the next task, and cannot be trusted to hold the plot behind it.

The strongest objection is that all of this is a snapshot of a fast-moving target. METR has been measuring how long a task frontier models can finish on their own, and found that the length they clear with even odds has doubled roughly every seven months for six years. Extend that line and an agent coherent for an afternoon today is coherent for a fiscal quarter within a few years, at which point the leased storefront looks less premature than prescient. If the curve holds, patience is the whole strategy, and RetailBench is timing a problem that the trend erases.

The extrapolation quietly swaps the reliability that matters for the one that flatters. The horizon METR watched doubling is the one models clear with even odds. A store left alone does not need an agent that succeeds half the time; it needs one that fails rarely, and that horizon is barely moving. METR also measured software tasks with graded, checkable answers, not the open-ended commercial judgment a supermarket demands across a season. The horizon a real store requires is both longer and messier than the one on the trend line.

That leaves reliability, not reach, as the number to watch, and on reliability even short tasks disappoint. The τ-bench study found that GPT-4o, asked to clear the same retail task eight times over, succeeded on all eight less than a quarter of the time. Whoever hands an unattended store to an agent is not buying autonomy; they are agreeing to own whatever it does on the morning it decides the shop sells tea. If the reliability horizon starts doubling the way the median horizon has, the set-and-forget store stops being a bet and becomes a plan. Until then, the credible deployment keeps a human watching the store, which is precisely the labor the pitch promised to eliminate.