Part of our work on retail
Retail Industry
Why Buy-Online-Pickup-in-Store Breaks at the Same Handoff in Every Rollout
Big Sky Consulting Group · September 28, 2026 · 7 min read

The pilot looked fine
The pilot ran in three stores. The managers were briefed, the regional lead checked in daily, and orders trickled in at a pace the floor could absorb. Pickup times were good. The board saw a chart that went up and to the right, and the rollout was approved.
Six months later, forty stores are live and the complaints have a shape. Customers arrive to find the order is not ready. Some were told it was. Cancellations climb at the busiest locations. Head office asks the vendor what is wrong, and the vendor says what every vendor says: inventory accuracy.
That answer is not wrong. It is the second failure. The first one happens earlier, at the same step every time, because it is built into how pickup works.
The store associate is the integration layer
Trace a buy-online-pickup-in-store order from the moment the customer taps confirm.
The ecommerce platform takes the payment and hands the order to the store. From there, no system does the work. A person does. Someone has to notice the order on a screen, walk the floor, find the item, bag it, label it, stage it in the right place, and mark it ready. Every one of those steps is unmanaged labour, and every one competes with the queue at the register, the delivery at the back door, and the customer asking where the paint samples are.
OneView Commerce, which sells store fulfillment software, puts the consequence bluntly: when associates are the integration layer, "the operation is one personnel change away from failure." We would put it more narrowly. The associate is the only component in the chain that has no service level, no timestamp at each step, and no one watching it from head office.
That is why the break is always in the same place. We call it the ready-time gap: the interval between "your order is confirmed" and "your order is picked." It is invisible to head office, because no system measures it. It is also the only interval the customer actually experiences.
What the public data says, and how little of it there is
The best independent measurement of this interval we can find is a secret-shopper study that IHL Group ran with OrderDynamics in late 2018. They sent 300 shoppers through the pickup process at ten large US retailers. Across the whole sample, 69 percent of orders were ready within two hours and 51 percent within one.
The averages hid the spread. The fastest retailer averaged 2.1 hours from order to ready notification. The slowest averaged 16 hours, and more than half its orders took over four. The researchers' explanation for the laggards was not a systems failure. It was that "the local stores were not prioritizing the verification of the orders." In other words, the handoff to people was where time disappeared.
They also found a strong correlation between notification time and whether the shopper used the service again, recommended it, and bought more on the visit. A later Rakuten Ready study, built on secret shoppers in eleven metro areas plus its own transaction data, found that customers who waited under two minutes at pickup were four times more likely to buy again.
Notice the age of those numbers. The most useful public data on the single interval that decides whether pickup works is years old and came from secret shoppers, not from retailers' own systems. That is not an accident of research budgets. It is the finding. Retailers do not publish this number because most of them do not have it.
Why the pilot cannot see it
A pilot is almost designed to hide the ready-time gap.
Volume is low, so the associate who picks is rarely interrupted. Attention is high, so orders get noticed quickly because someone senior is watching. The stores chosen are usually the best-run ones. Every condition that makes the gap widen at scale is absent at three stores.
Then the rollout adds volume, removes the attention, and reaches stores whose managers were not in the kickoff meeting. The work does not change. The conditions around it do. The gap opens, and because nobody timed it in the pilot, there is no baseline to compare against. Head office sees cancellations and complaints, which are lagging signals, and goes looking for a technical cause.
There is a volume dimension too. LinQu, a smart locker vendor, claims the bagged-orders-on-a-shelf model breaks at roughly 50 pickup orders per store per day, and puts each manual handoff at two to four minutes of associate time. They cite no source, and they sell the thing that replaces the shelf, so treat both numbers as a vendor's framing rather than a benchmark. The mechanism behind the claim is sound, though. Picking and handing off are fixed minutes per order. At some volume, those minutes stop fitting inside the gaps in a shift, and your store will find its own threshold whether or not anyone is measuring.
This is the general shape of the problem. Which parts apply to your process depends on answers only your systems can give.
Put us on it, from $5,000Where inventory accuracy actually fits
None of this makes inventory accuracy irrelevant. It makes it second.
When the associate walks to the shelf and the item is not there, the ready-time gap turns into a cancellation. Much of that is sync timing rather than bad counts. Mortar, which builds inventory sync for multi-location retailers, describes the common batch pattern: an integration pushes store stock to the website every five, ten, or fifteen minutes. For a moderately popular item, they say, the site sells stock that has already gone out the front door multiple times a day. Separately, if the count itself is off on your fastest movers, the problem compounds. We covered where that error tends to hide in what 92 percent inventory accuracy actually costs.
But fix inventory first and you have still left a person, with no queue and no clock, as the bridge between the website and the shelf. Accurate stock that nobody picks for four hours is still a failed order. The order of operations matters because the fixes cost very different amounts, and the inventory fix is usually the more expensive one.
The question that predicts the rollout
Ask a retailer for its median confirm-to-ready time, by store.
Not the average across the chain. Not the percentage of orders marked ready. The median time from the customer's confirmation to the moment the order was actually staged, broken out per location.
If someone can produce it, the retailer is in a good position. They can see which stores are drifting, which shifts are the problem, and whether volume or staffing is driving it. They can set a service level and hold a store to it.
If nobody can produce it, the pilot will look fine and the rollout will not. That is the most reliable prediction we know of in store fulfillment, and it holds regardless of which platform is being installed. The software can be excellent. It is still handing work to a step that nobody measures.
There is a second, quieter test. Ask when an order is marked ready. In many stores, "ready" is a button an associate taps, sometimes before the bag is staged, sometimes in a batch at the end of a rush. If the ready flag does not mean the order is physically waiting where the customer will look for it, then the one timestamp you do have is measuring the wrong thing. The customer who drives over on that notification finds out first.
What this means before you buy anything
The instinct, once the complaints start, is to buy. A store fulfillment app, a task manager, lockers, a better order management system. Some of those will help. All of them will be sold with a demo that shows an order moving cleanly from website to shelf, because in a demo there is no register queue.
The sequence that holds up is less exciting. Measure the gap first, with whatever crude instrument you have, even a timestamp an associate enters when the bag hits the shelf. Find out whether it is wide everywhere or wide in specific stores and shifts. Those are different problems with different fixes, and only one of them is a software problem.
This is the same pattern we described in what a multi-location retailer owes its own operations team: when systems do not connect, a person connects them, and that person's time never appears on a dashboard. Pickup is the most customer-visible version of it. It is also the version where the cost is easiest to see once you look, because every hour of gap shows up in cancellations and in customers who do not come back.
When the business case for a fix does get written, make sure it counts the associate minutes it claims to save against the associate minutes the new process adds. Vendor models tend to skip the second column, and we walked through why in the automation ROI math vendors show you, recalculated.
The honest edge
An article can tell you where pickup breaks and which number predicts it. It cannot tell you what your number is, which of your stores are carrying the gap, or whether your next dollar belongs in software, staffing, or process.
If your pickup rollout is under way and the complaints do not match what the pilot showed, talk to us. We will find your confirm-to-ready time by store, show you where it widens, and tell you whether the fix is something you need to buy at all.
