Part of our work on industrial goods
Industrial Goods
Exception Handling Is Manufacturing's Biggest Automation Blind Spot
Big Sky Consulting Group · September 2, 2026 · 7 min read

The queue nobody owns
Somewhere in your plant there is a person whose actual job is the exception queue. On the org chart they are a planner, a buyer, a customer service rep or an AP clerk. In practice they spend most of their week on the orders that did not go through: the short ship, the price mismatch, the reschedule message, the part that shipped against a routing that was changed in March and never updated in the system.
Nobody hired for that role. It accumulated. And because it accumulated rather than being designed, it does not appear as a line in any budget, which is why the first serious conversation about automating it usually happens five years too late.
The vendor pitch writes itself from there. Redwood Software's Manufacturing AI and Automation Outlook 2026, a survey of 300 manufacturing professionals, found that only about 40 percent of manufacturers have automated exception handling, while 22 percent name it a top operational bottleneck. Read that as vendor research, because it is, and it is still directionally right. Four in five plants are working exceptions by hand and one in five has noticed it hurts.
We would read those two numbers differently than the deck that contains them. The gap between 40 percent automated and 22 percent bottlenecked is not a technology gap waiting to be closed. It is a distribution. Most plants have exceptions and are fine. Some have exceptions that are eating a department. The difference between the two groups is not automation maturity. It is how many distinct causes are generating the volume.
An exception is a process telling you its rules are wrong
This is the part that gets skipped, and it is the whole thing.
An exception is not a random event. It is an instance where the transaction the system expected and the transaction reality produced did not match. Every one of them has a cause, and the causes are almost never evenly spread. In most operations we have looked at, a handful of part numbers, a handful of customers and a handful of suppliers generate the clear majority of the queue. Not because those parties are difficult, but because something in the setup is wrong and has been wrong for years.
The lead time in the item master is the contractual one rather than the one the supplier actually achieves, so MRP plans a fiction and raises a reschedule message when reality arrives. The customer has a shipping term that was negotiated by someone who left in 2021 and never made it into the order template. The part has a unit of measure conversion that rounds the wrong way on odd quantities. Each of those produces exceptions on a schedule, forever, and each of them is a thirty minute fix by someone who knows where to look.
Automate the handling of those and you have done something specific: you have made the workaround permanent, and cheap enough that nobody will ever fix the cause. The queue stops hurting, which means it stops getting attention, which means the bad parameter stays in the system and quietly distorts every plan that depends on it. You have not removed the cost. You have moved it somewhere it cannot be seen and cannot be measured.
The finance side of the house has already learned this in public. AP automation content is unusually honest about it, probably because the numbers are hard to hide: automation creates real efficiency for clean invoices while institutionalizing the manual cost of recurring exceptions. Reported exception rates cluster around 14 percent of invoices with top performers near 9 percent, and the exceptions that remain are mostly PO mismatches, missing goods receipts, and vendor master problems. Note what those are. Not scanning failures. Master data failures. The technology worked perfectly and the data underneath it was wrong.
The number worth knowing is not the one on the dashboard
Every plant that raises this with us reports the same metric: exceptions per week, or percentage of orders touched. It is the wrong number. It measures volume, and volume tells you how much pain you are in, not what to do about it.
The number that decides the investment is how few distinct causes produce that volume. Take last quarter's exceptions, group them by root cause rather than by type, and sort. If twelve causes produce 70 percent of your queue, you do not have an automation project. You have twelve tickets and a Tuesday. If the tail is genuinely long, hundreds of causes each firing once, then the variability is real, the judgment is real, and automation is the right conversation to have.
Most plants have never run this cut, because the exception log records what happened rather than why. That is not an accident either. The person working the queue is measured on clearing it, so the record they keep is optimized for closing tickets rather than for explaining them. You are not going to find the answer in a report. You are going to find it by sitting with the two people who work the queue and asking them to name the repeat offenders, which they can do from memory, in about ten minutes, and have been able to do for years.
The general rule the RPA world arrived at the hard way is that a process is a good automation candidate when its exception rate is under roughly 20 percent. Above that, the exception handling logic becomes the automation, and you have built a system whose main job is coping with a process you declined to fix. The maintenance burden does not disappear. It transfers to whoever now owns the bot. We have written before about when AI is the wrong answer to an operations problem, and this is the same shape at a smaller scale: the technology works, and it is still the wrong purchase.
This is the general shape of the problem. Which parts apply to your process depends on answers only your systems can give.
Put us on it, from $5,000What the vendor page will not ask you
Page one for this topic is visibility platforms, AI inspection cameras and AP automation blogs. All of it is competent and none of it asks the two questions that decide the outcome.
The first is whether the exception you are about to automate is a decision or a defect. A defect has a cause you can remove. A decision is a human weighing something nobody wrote down: whether to ship short to a customer who will tolerate it, whether to expedite a supplier who has earned the benefit of the doubt, whether this dimensional variance is inside the customer's real tolerance or just outside the drawing. Automating a defect is cleanup. Automating a decision means writing down a rule that never existed, which is sometimes exactly the right project and is never the two week one that was quoted.
The second is who loses the ability to notice. The person working the exception queue is, in most plants, the only person with a live view of where the process is drifting. They know the supplier is slipping six weeks before the scorecard does. Take the queue away and you have not just automated a task, you have removed a sensor. That cost shows up nowhere in the business case and reliably shows up on the floor about two quarters later.
Neither question is exotic. Both get skipped because the sales motion runs on volume, and volume is the number the buyer already has. Counting causes takes a week and produces a smaller, less impressive project. It is also the only version of this that pays back.
Where automation genuinely earns it
We are not arguing against automating exception handling. We are arguing about sequence, which is the argument we end up having in most of these engagements.
Fix the causes that repeat. Then look at what is left. The residual queue, after the twelve parameters are corrected, is the honest candidate: genuine variability, spread thin, no repeat offenders, real judgment applied at low stakes. That is where routing, triage and orchestration do useful work, and where workflow orchestration earns its keep by moving exceptions to the right person with the right context rather than into a shared inbox.
The order matters financially, not just philosophically. Automating first means your business case is built on a volume that was about to drop by half anyway, so the payback you calculated will not appear and nobody will be able to explain why. Fixing first means you scope the system against the volume that is actually permanent. Our view on how to calculate automation ROI applies directly here: the denominator has to be the steady state, not the mess.
There is one more reason to go in this order, and it is the one operators feel first. A plant that fixes twelve root causes gets faster and quieter. A plant that automates twelve root causes gets exactly as broken as it was, just with better reporting on it. Call it the difference between fixing the leak and buying a very sophisticated bucket.
Where this stops being an article
What we cannot do in writing is tell you which side of the line your plant is on. That depends on your cause distribution, on whether your exception log records anything useful, on how much of the judgment in that queue has ever been written down, and on who in the building is quietly holding the process together in ways the org chart does not show.
Those are answerable questions, but only from inside your data and a few conversations with the people working the queue. If your exception volume is rising and every proposal you have seen prices the handling rather than the cause, talk to us. We will tell you which of your exceptions are worth automating, and which twelve you should just go fix.
