Logistics & Supply Chain

Why Warehouse AI Pilots Stall at the Throughput Test: Six Recurring Patterns From Production Failures

Most warehouse AI pilots look great in testing then fall apart once they go live. This post breaks down the six reasons that keeps happening: the mix of products changes (from simple, predictable items to messy, varied ones), busy-season volume spikes the system was never tested against (peak-load gap), promised speeds that don’t match real-world speeds, robots and software that don’t talk to each other properly, storage layouts that go stale as demand shifts, and nobody watching for problems after launch. If a pilot shows two or more of these, it’s not ready to scale.

Why Warehouse AI Pilots Stall at the Throughput Test: Six Recurring Patterns From Production Failures

The contrarian observation that sets up this whole post

Most warehouse AI pilot post-mortems chase the wrong failure mode. The vendor blames change management. The integrator blames the model. The operations team blames the data. After auditing throughput failures across multiple [Autonomous Mobile Robot] AMR, robotic-picking, and goods-to-person deployments and reading the public receipts on Symbotic’s fiscal 2025 deployment turbulence, Berkshire Grey’s de-SPAC at $1.40 per share, and Covariant’s effective pivot into Amazon the same six patterns appear over and over. The model usually works. The pilot UPH [Units per hour] usually validates. What fails is the transition from a tightly controlled pilot lane to the real distribution center: full SKU mix, peak volumes, multi-vendor fleet,
daily velocity drift.

This post is the taxonomy. The 6-Pattern Throughput Failure Taxonomy names the recurring failure modes, attaches named-adopter receipts to each, and gives you the diagnostic question to run against your own pilot before you sign the multi-site rollout SOW. If your pilot is showing two or more of these patterns today, the multi-site rollout will not survive Q4 peak.

What “stall at the throughput test” actually means?

A warehouse automation pilot at the throughput test when the system runs successfully in a controlled scope one zone, one shift, a curated SKU subset, vendor engineers physically present but the projected production throughput collapses the moment any of those constraints is removed. The vendor-promised 1,200 units per hour (UPH) at the pre-sales demo becomes a measured 380 UPH
after the third week of production volume. The 99.4% pick accuracy on Phase-1 SKUs drops to 91% when Phase-2 SKUs enter the mix. The four-week ROI payback projection extends to eighteen months because the operator is still running manual workarounds for 35% of orders.

The public evidence on this is dense in 2025–2026. Locus Robotics surpassed 6 billion picks in October 2025 the most recent billion in just 24 weeks, the fastest pace in its history, with 17,000 AMRs deployed across more than 150 customers at 350+ facilities in 20 countries. That is the tier of execution discipline that ships. At the other end, SoftBank acquired Berkshire Grey at $1.40 per share down from a ~$10 reverse-merger debut in mid-2021 after the company’s order backlog “continued to climb, reaching $265m this month, but it could not get its systems out fast enough”. Backlog growth alone is not the metric; throughput conversion of that backlog is.

For every 33 AI pilots a company launches, only 4 make it to production an 88% failure rate at the scaling step. That is IDC’s number across enterprise AI generally. Warehouse automation projects sit inside that distribution, and the throughput test is where most of the failures concentrate.

The 6-Pattern Throughput Failure Taxonomy

Every pattern below is observable. Everyone carries a public receipt or an audit-able operational signal. Use them in order: a pilot that exhibits any single pattern can usually be remediated. A pilot that exhibits two or more is already failing multi-site rollout will magnify the failure across the network.

Pattern 1 — SKU-mix distribution shift

The pilot trained on Phase-1 [Stock Keeping Unit] SKUs: predictable cuboid cartons, consistent weights, standard packaging. Phase-2 rollout introduces the long tail polybags, blister packs, irregular geometry, dunnage variance, items with reflective surfaces or see-through packaging that confuses the robot’s cameras, items that get tangled together, or blocked from view. The pick model that scored 99.4% on the training distribution drops to 91% on the full catalog because the pilot was effectively a different problem.

The Amazon Sparrow trajectory illustrates the discipline required to manage this pattern at scale. As of late 2023, Sparrow could handle about 65% of the more than 100 million items in Amazon’s inventory; the robot reduced its defect rate by 65% during the San Marcos pilot as part of operational learning. By 2025, Sparrow handled over 200 million unique products but that expansion was earned through sustained data collection, not announced at pre-sales. Most enterprise pilots have neither the data pipeline nor the model-retraining cadence to close the same gap.

Diagnostic question: What percentage of your full production SKU catalog was present in the pilot training distribution, weighted by velocity?
Fix: Continuous fine-tuning pipeline with labeled production failures fed back into training; SKU-cohort accuracy reporting in addition to overall accuracy; vendor contract clauses tied to accuracy on the production distribution, not the pilot distribution.

Pattern 2 — Peak-vs-non-peak training data gap

The pilot ran in Q2 and Q3 at normal volumes. Q4 peak hits, volumes 4–5× the pilot baseline, order batch composition changes (smaller order sizes, higher SKU diversity per order, more rush-shipping tags), and the system collapses. Throughput drops not because the robots are slower but because it was only trained for ‘normal’ days, not chaotic ones.

Locus Robotics built peak-season capacity as a first-class product feature exactly because of this pattern: deployments of incremental, peak-season bots surged nearly 50 percent year-over-year, and Locus is on track to hit 60 million picks per week during Q4 peak season up from 45 million per week throughout 2025. The variance from baseline to peak inside a single network is not a tail event; it is
structural.

Diagnostic question: Was the pilot UPH measured under Q4-equivalent peak load with peak equivalent batch composition, or was it measured under steady-state conditions?
Fix: Synthetic peak simulation during pilot (deliberately running peak-equivalent volumes through the pilot zone for at least one full week); peak playbook codified before pilot exit; contractual UPH guarantees that specify the load conditions under which they apply.

Pattern 3 — Vendor-claimed UPH vs site-realized UPH gap

This is the pattern most operators feel before they can articulate. The pre-sales deck quotes 1,200 UPH on the picking arm; the laboratory video is real; the third-party benchmark cited is real. The deployed system measures 380 UPH after Week 4 and 410 UPH after Month 6. Nobody lied the lab conditions assumed picking from a single-SKU bin with optimized presentation, near-perfect lighting, items pre singulated. The production conditions are mixed-SKU bins, item entanglement, occlusion, intermittent lighting glare, and operator-driven exception handling that breaks rhythm. We cover this in depth in vendor-claimed UPH lies.

This is also where the Covariant trajectory is instructive. On March 11, 2024, Covariant announced RFM-1 (Robotics Foundation Model 1) a robotics foundation model giving robots human-like reasoning, with the Covariant Brain trained on tens of millions of episodes. By August 30, 2024, Amazon announced a non-exclusive license to Covariant’s technology, hiring 25% of the workforce and
the three founders. The deal was disclosed in a 2025 whistleblower complaint at $380 million plus a $20 million final licensing payment one year after close a “reverse acquihire” structured to avoid antitrust scrutiny. Even a foundation-model-grade picking technology backed by tens of millions of episodes ended up effectively absorbed into Amazon rather than scaling independently. The gap
between the laboratory demonstration and the multi-customer production economics turned out to be wider than the venture-backed timeline could survive.

Diagnostic question: What is the documented site-realized UPH on a deployed system at a customer with your SKU mix, your bin design, and your facility conditions and what was the variance versus the pre-sales projection?

Fix: Reference-site visits to a customer with comparable SKU and facility profile, not a vendor’s flagship demonstration site; UPH guarantees written into the contract with measurable SLAs and remediation tied to the realized number; structured pre-sales-to-post-sale variance tracking carried through the
procurement file.

Pattern 4 — AMR-WCS-WMS integration debt

The robot kinematics work in isolation. The fleet manager works in isolation. The Warehouse Management System (WMS) works in isolation. The Warehouse Control System (WCS) works in isolation. Production requires all four plus the upstream Enterprise Resource Planning (ERP) order release, the downstream sortation conveyor, and the labor-management interface to behave as a single coordinated system in real time. A Warehouse Execution System (WES) sits between WMS and WCS, taking high-level plans from the WMS and translating them into real-time, optimized instructions across all automation systems and human workers simultaneously. Most pilots skip the WES layer. They appear to work because the pilot scope is narrow enough that the missing orchestration is
invisible.

“Robots working in silos. The ACR fleet handles storage and retrieval. The AMRs handle transport. Without a WES, each of these systems would operate on its own schedule” and the throughput collapse is concentrated at the handoff points: dock-to-storage, storage-to-pick, pick-to-pack. The Symbotic disclosures from 2025 surface the cost of integration friction even at the vendor’s own deployment cadence. Symbotic CEO Rick Cohen cited construction delays and an expensive sensor upgrade as drivers of the surprise quarterly loss; the company brought engineering planning in-house after outsourced suppliers underperformed, and the stock dropped 24% in a single trading session. When the leading retail automation vendor itself struggles with deployment integration, the integration
debt is real for every operator in the market.

Diagnostic question: Is there a documented WES layer in the production architecture, with named owners for each integration boundary (WMS↔WES, WES↔WCS, WCS↔fleet manager, WMS↔ERP)?
Fix: WES-first architecture, not WES-as-afterthought; integration-test plan documented and signed off before pilot scope is locked; multi-vendor fleet management capability tested with at least two robot types under realistic load.

Pattern 5 — Slotting refresh lagging velocity drift

SKU velocity changes daily. New products launch, seasonal items rotate, promotional spikes redirect demand, supplier substitutions flow downstream. The pilot’s slotting plan, what SKUs sit in which locations to minimize pick travel was optimized once at pilot start and never recalibrated. Throughput degrades 15–30% within months as the slotting decays. Operators feel it as “the robots are slower lately” without realizing the model is fine; the layout it was optimizing against is stale.

This is one of the patterns where the technology genuinely is mature and the deployment discipline is the failure mode. LocusONE continuously optimizes throughput, fleet productivity, and network efficiency by analyzing billions of data points across robots and tasks — but the orchestration layer can only optimize against the layout it can see. If slotting refresh runs quarterly while velocity changes weekly, the optimization is permanently a quarter behind reality.

Diagnostic question: What is the cadence of slotting refresh, and what is the documented decision rule that triggers re-slotting (volume threshold, velocity-rank change, seasonal calendar)?
Fix: Automated slotting recommendations on at least a weekly cadence; trigger-based re-slotting when velocity rank shifts cross a documented threshold; KPI dashboards that surface slotting decay as a leading indicator of throughput risk.

Pattern 6 — Inadequate post-deploy monitoring

The pilot launched. The vendor engineers left. Six weeks later, throughput is down 20% and nobody on the operations team can explain why. There are no dashboards on per-robot UPH, per-SKU pick rate, exception frequency, model confidence distribution, or queue depth at handoff points. The first time leadership hears about the throughput collapse is when the warehouse general manager calls the VP of Operations and uses the word “broken.”

This is the cheapest pattern to fix and the most commonly skipped, because post-deploy monitoring is treated as an optional Phase-2 deliverable rather than a Day-0 requirement. “AI models in production need data quality signals measured in hours. That mismatch is where most data quality AI problems originate” — and the same applies to throughput signals. A weekly review cadence is not monitoring; it
is post-mortem theatre. Real monitoring catches a 4% degradation on Tuesday, before it becomes a 22% degradation on Friday.

Diagnostic question: Are there real-time dashboards on per-robot UPH, per-SKU accuracy, exception frequency, and handoff queue depth — and is there a documented SLA with named owner for each metric?

Fix: Monitoring stack defined and operational before pilot launch, not after; SLA-tied alerting on the leading indicators; quarterly model performance reviews with a documented retraining trigger.

Build vs Buy vs Partner — for the integration layer

The patterns above are not equally addressable by every procurement path. The table below maps the decision space.

Criterion Pure SaaS (Locus / Symbiotic / Berkshire Grey) Big consulting integrator (Accenture / Deloitte) Generic AI agency Brainy Neurals
Time to first pilot UPH 4–8 weeks 12–20 weeks 8–14 weeks 6–10 weeks
Multi-vendor fleet orchestration Vendor-locked to their hardware Strong, slow Variable Built vendor-neutral from spec
WES layer ownership Vendor’s Client’s, designed slowly Often missing Client-owned, architected in

The right answer depends on which of the 6 patterns is most likely to bite your operation. If SKU-mix distribution shift and slotting drift are the binding constraints, a SaaS-native vendor with strong orchestration (Locus, AutoStore-class) is usually right. If integration debt across multi-vendor fleets is the binding constraint, a vendor-neutral integration partner whether a big consulting firm or a specialist like Brainy Neurals wins. The default failure mode is choosing the procurement path on brand familiarity rather than on which pattern is the active risk.

What Mitesh has seen across 40+ enterprise AI deployments

“On the warehouse AI work we’ve shipped, the throughput test never fails for the reason the steering committee assumes. The model isn’t broken the pilot was just measuring a different operational reality. We’ve audited deployments where vendor-quoted UPH was 1,200 and site realized UPH stabilized at 380. The gap wasn’t model accuracy. It was that the pilot ran with single-SKU bins under optimal lighting, and production runs with mixed-SKU bins under whatever
lighting the building actually has. Get the pilot conditions to match the production conditions before you sign the multi-site rollout, or get used to the 88% pilot-to-production failure rate that IDC has been publishing since 2023.”

— Mitesh Patel, NVIDIA Certified AI Architect, Founder & Director, Brainy Neurals

A confession about how warehouse AI specifications actually fail

We have rewritten 40+ AI specifications mid-project at this point. The pattern is always identical: the spec was written by someone who couldn’t tell it was wrong, because they had never shipped a model into the failure modes the production environment would hit. In warehouse AI specifically, the spec usually under-constrains three things the SKU distribution to be supported, the peak load conditions to be tested against, and the multi-vendor integration scope. The fix is rarely a better model. It is rewriting the spec so the pilot exit criteria look exactly like the production conditions on day one of multi-site rollout. That is the unglamorous work that decides whether the pilot ships.

When this taxonomy doesn’t apply

The 6-pattern taxonomy assumes a warehouse above ~250,000 square feet, ≥10,000 active SKUs, an established WMS layer (Manhattan, Blue Yonder, SAP EWM, Körber, Oracle), and a procurement appetite for at least one AMR or robotic-arm vendor. The framework breaks in three conditions worth naming.

Single-product fulfillment at scale. Operations like Amazon Sortable or a single-SKU FBA depot collapse most of the SKU-mix variability that drives Patterns 1, 2, and 5. The integration patterns still apply; the SKU patterns largely don’t.

Greenfield builds with no legacy WMS. A greenfield site building the WMS, WES, and WCS in the same procurement window does not face integration debt the same way a brownfield retrofit does. Patterns 4 and 6 attenuate; Patterns 1, 2, 3, and 5 remain.

Sub-$50M annual revenue 3PL operations. The SaaS economics that work at enterprise scale eat margin at this tier. “The top barriers restricting automation adoption are lack of budget, clear business case, and understanding of technology, according to the MHI 2025 Annual Industry Report. When asked to name biggest obstacles to future automation plans, budget (41%) and cost/ROI (40%) topped
the list”. The taxonomy applies but the remediation budget often does not.

What the numbers actually look like

  • 88% of AI pilots fail to reach production — IDC research; 4 out of every 33 pilots launched graduate
  • 80.3% AI project failure rate — RAND 2024, confirmed by Gartner April 2026; twice the rate of conventional IT projects
  • 6 billion picks in 24 weeks — Locus Robotics, October 2025; the fastest pace in company history,demonstrating the production-tier compounding curve.
  • 65% defect-rate reduction — Amazon Sparrow during San Marcos pilot; the inflection point that enabled the Houston expansion.
  • $14M surprise quarterly loss + 24% single-day stock drop — Symbotic, July 2025; driven by construction delays and a sensor upgrade absorbed rather than passed to customers
  • $1.40 per share take-private valuation — Berkshire Grey, March 2023; down from a $10 SPAC debut on inability to convert order backlog to deployed throughput fast enough
  • $380M + $20M reverse-acquihire — Covariant by Amazon, August 2024; disclosed via a 2025 whistleblower complaint to the FTC, SEC, and DOJ

Frequently asked questions

What is the success rate of warehouse automation projects in 2026?

Industry-wide AI project success at the pilot-to-production transition runs 12–48% depending on the source. IDC reports that 4 of every 33 enterprise AI pilots reach production an 88% failure rate at the scaling step. Gartner reports 48% reach production on average across enterprise IT. RAND’s 2024 meta-analysis put the figure at 80.3% failure twice the rate of conventional IT projects. Warehouse AI specifically sits inside that distribution; the throughput test is where most of the failures concentrate, with the 6-pattern taxonomy explaining the recurring causes.

Why do AMR pilots fail so often?

AMR pilots fail when the pilot conditions diverge materially from production conditions on SKU mix, peak-load behaviour, vendor-claimed throughput rates, multi-system integration, slotting cadence, or monitoring discipline. The 6-Pattern Throughput Failure Taxonomy organizes those six recurring causes. A pilot exhibiting any single pattern is usually remediable; a pilot exhibiting two or more rarely survives the multi-site rollout. The fix is to test the pilot against production-realistic conditions synthetic peak load, full SKU distribution, multi-vendor integration scope before exit, not after.

How long does it actually take to scale warehouse automation from pilot to production?

For SaaS-led platforms (Locus, Symbotic, AutoStore, Geek+), 6–12 months from initial pilot signing to multi-zone production rollout in a single facility; 18–36 months for multi-facility network rollout. For custom or partner-led builds, 9–18 months to first production site; 24–48 months for network rollout. Wharton research notes that organizations typically underestimate production deployment complexity by 300–500% projects scoped for three-month implementations actually require 12–18 months. The variance is dominated by integration complexity, not by model development.

What is the difference between vendor-claimed UPH and site-realized UPH?

Vendor-claimed UPH (units per hour) is the throughput rate quoted in pre-sales materials, typically measured under optimized conditions: single-SKU bins, controlled lighting, pre-singulated items, no exception handling. Site-realized UPH is the throughput rate measured under actual production conditions: mixed-SKU bins, intermittent lighting glare, item entanglement, operator-driven exception handling. The gap is often 2–4× — vendor-claimed 1,200 UPH becomes site-realized 380 UPH on common picking workloads. The fix is to insist on UPH guarantees written into the contract under specified site conditions, and to validate against a reference-site visit at a customer with a comparable
operational profile.

Why did Symbotic miss Q3 2025 deployment targets despite a $22 billion backlog?

Symbotic disclosed in its July 2025 earnings call that construction delays and a sensor upgrade absorbed by Symbotic rather than passed to Walmart drove a surprise $14 million quarterly loss; CEO Rick Cohen attributed underperformance to outsourced suppliers, and the company brought engineering planning in-house. By fiscal year 2025 close, Walmart accounted for 84% of revenue and the company sat on a $22.5 billion backlog dominated by Walmart and the SoftBank-Symbotic GreenBox joint venture. The pattern is integration debt at the vendor’s own deployment cadence backlog growth is real, but throughput conversion of that backlog is the constrained resource.

What does it mean that Covariant was “absorbed” by Amazon?

In August 2024, Amazon entered into a non-exclusive license for Covariant’s Robotic Foundation Model 1 technology and hired roughly 25% of Covariant’s workforce including the three founders into Amazon’s Fulfillment Technologies & Robotics organization. A 2025 whistleblower complaint to the FTC, SEC, and DOJ disclosed the financial structure as $380 million plus a $20 million licensing payment one year after close a so-called “reverse acquihire” structured to avoid antitrust scrutiny. Covariant continues to operate under new leadership focused on apparel, health and beauty, grocery, and pharmaceuticals. The structural lesson is that even a foundation-model-grade picking technology with tens of millions of training episodes could not sustain the multi-customer production economics that the venture-backed timeline required.

Should we buy from a SaaS vendor or build with a partner?

For the well-defined production patterns where SaaS vendors have public deployment receipts (AMR based person-to-goods picking, established goods-to-person grids, established carton-handling sortation), buying from the SaaS vendor is the right starting move the integration paths are cleared, the reference customers exist, and the cost is bounded. Build with a partner only when (a) you have proprietary process IP that differentiates your operation competitively, (b) the multi-vendor orchestration layer is the binding constraint, or (c) you require full IP ownership of the model weights, annotations, and integration code. The hybrid pattern that wins most often: SaaS for the well-defined patterns, vendor-neutral integration partner for the WES layer and post-deploy monitoring.

How do we know if our pilot is showing two or more failure patterns before we sign the multi-site rollout?

Run a structured pilot-exit audit. For each of the 6 patterns, ask the diagnostic question included in this post (SKU coverage, peak conditions, UPH variance, WES ownership, slotting cadence, monitoring SLAs). A pilot exhibiting zero or one pattern is ready for multi-site rollout. A pilot exhibiting two or more is not the next 12 months of the rollout will magnify each pattern across the network. The audit is cheaper than the rollout failure; we cover the structured audit approach in validation playbook. Skipping the audit and assuming the steering-committee narrative is the most common predictor of a stalled rollout in 2026.

What this means for your next 90 days?

If you run a warehouse network of 4 or more facilities and you have AI automation budget allocated for 2026, the 6-pattern taxonomy gives you three concrete moves before the next steering committee.

Audit any active pilot against all 6 patterns this quarter. A pilot exhibiting two or more patterns is not ready for multi-site rollout pause, remediate the patterns, then resume.

Insist on UPH guarantees written into every new SOW, specified against your SKU mix and facility conditions, not the vendor’s reference site. This single contractual change is the highest leverage 2026 move see the Symbotic-Walmart lessons post for the public receipts on what happens when this isn’t done.

Stand up the post-deploy monitoring stack on Day 0, not Phase 2. Per-robot UPH, per-SKU accuracy, exception frequency, and handoff queue depth dashboarded, SLA-tied, alerting. If you cannot see throughput degradation in real time, you will see it in the quarterly review meeting after it has already cost a peak season.

The patterns above are not new in 2026. They were present in the Berkshire Grey de-SPAC in 2023, in the Covariant absorption in 2024, in the Symbotic deployment turbulence in 2025. The operators who have moved fastest Locus Robotics’ customers compounding from 4 billion picks to 7 billion picks in roughly eighteen months did not chase more sophisticated models. They wrote tighter pilot exit criteria, stood up monitoring on Day 0, and rebalanced their vendor selection toward the procurement path that addressed their binding pattern. That is the architectural decision waiting on your desk this quarter.

Take the next step

Before you sign the next multi-site warehouse AI rollout SOW, run your pilot through the 6-pattern audit.
Get the AI Project Scoping Template — the structured pilot-exit criteria document we use with warehouse AI clients, with the 6 diagnostic questions baked in. Download free →
Or subscribe to AI in Production, the monthly enterprise AI breakdown from Mitesh Patel — what shipped, what stalled, and what to fund next quarter. Subscribe →

For an enterprise warehouse AI scoping conversation: Book a 30-minute call with our NVIDIA
Certified AI Architect →

About the author

Mitesh
Director · NVIDIA Certified AI Architect

Mitesh

Mitesh Patel is Director of Brainy Neurals and a NVIDIA Certified AI Architect — a credential held by fewer than 3,000 people worldwide. He has spent nine years building production AI, beginning in C/C++ firmware and edge inference (Jetson, NVIDIA Triton, TensorRT, Qualcomm SNPE) and now leading enterprise RAG, computer vision, and generative AI delivery. He holds a B.Tech in Electronics & Communication and an M.Tech in Embedded Systems, and is an Upwork Top Rated Plus practitioner (top 3%). His team of 20 specialist engineers has shipped 70+ enterprise AI projects across manufacturing, BFSI, healthcare, logistics, and construction. Read more at the founder page or connect on LinkedIn.