Your on-time delivery rate is 91% this quarter. Is that good?
Depends who you ask. Industry benchmarks put 95-98% at "world-class" -- so 91% is a real problem, not a rounding error. Walmart requires 98% OTIF (on-time, in-full) from its suppliers and charges a 3% cost-of-goods-sold penalty for every miss below it -- one of the more striking numbers in supply-chain benchmarking, because it puts an actual dollar figure on what "close enough" costs a supplier at scale.
But the number that should worry you more than 91% itself is the question almost no OTD dashboard, report, or article actually answers: why is it 91%, and what is each cause costing you?
And there's a harder question underneath it. Your OTD rate only counts the orders you lost. It says nothing about what you paid to win the other 91%.
The Pain: OTD Tells You That You're Late, Never Why
Every manufacturer tracking on-time delivery already knows the formula -- on-time shipments divided by total shipments, or OTIF if you're also grading against complete quantity. The metric isn't the hard part. The hard part is what happens after the report lands on someone's desk: a red number, a general sense that things need to improve, and no clear next action.
That's because a missed delivery is never one thing. Across real manufacturing operations, late shipments trace back to a small, recurring set of root causes -- and they don't have the same fix:
- Material shortage -- a component didn't arrive in time to start the job
- Capacity constraint -- the line or cell was already committed to other work
- Machine or tooling downtime -- unplanned maintenance ate the schedule
- Scheduling decision -- a deliberate reprioritization bumped the order
- Labor shortage -- not enough hands on the shift that needed them
- Quality hold -- the part failed inspection and had to be reworked
- Supplier delay -- an upstream vendor missed their own date
A material-shortage-driven miss and a capacity-driven miss both show up identically on an OTD dashboard: one red data point. But the fix for one is a safety-stock or supplier-lead-time conversation, and the fix for the other is a scheduling or capacity-investment conversation. Treating them as the same problem because they produced the same metric is exactly how a "we need to improve on-time delivery" initiative turns into a vague mandate nobody can execute against.
The two OTD numbers in your ERP, and why they disagree
Before you attribute anything, settle which number you're attributing. Most ERPs hold two, and almost nobody reconciles them.
The first is OTD measured against the current promise date -- the date sitting in the order record today. The second is OTD measured against the original commit date -- what you told the customer when they placed the order, before anyone revised it.
Every time a job slips and someone updates the promise date, the first number heals itself. The order ships "on time" against a date that was moved specifically because it wasn't going to ship on time. The second number doesn't heal, because the original commit is what the customer actually planned their own production around -- and it's the number their scorecard is quietly keeping.
The gap between those two figures is the most useful diagnostic in the whole exercise, and it takes one query against data you already have. A wide gap means your OTD rate is measuring your rescheduling discipline, not your delivery performance. It also explains the conversation every operator in this position eventually has: the internal report says 96%, the customer says you're a problem supplier, and both of them are reading real data.
The Proof: Nobody in the Category Actually Attributes the Miss
The OTD/OTIF content space is real and reasonably crowded -- shipping platforms, freight-visibility tools, and supply-chain consultancies all publish some version of "what is on-time delivery and how do you calculate it," usually with a benchmark table and a handful of generic improvement tips (real-time tracking, more realistic delivery windows, better carrier selection).
What none of them do is take a specific missed-delivery number and break it down by cause, with a dollar figure attached to each category. The content stops at "here's the metric and here's the industry average" -- exactly the level of insight you'd get from a dashboard, not the level you'd need to decide what to fix first.
That gap matters because the industries most exposed to this are the ones where a single cause tends to dominate without anyone formally measuring it. Material shortages alone are estimated to drive roughly 20-30% of late deliveries industry-wide -- a large enough share that a manufacturer guessing at their own breakdown is more likely to be wrong than right, and more likely to fix the wrong lever first.
Our Actual Opinion: OTD Is a Working-Capital Metric Wearing an Operations Costume
Here's the part we'd argue with a consultant about.
On-time delivery gets filed under operations, assigned to a plant manager, and reviewed in an operations meeting. That filing decision is why the money stays invisible. A late order is not primarily a service failure -- it is cash you have already spent, sitting on your floor, not yet permitted to start the collection clock.
Walk the causes through the cash cycle and they stop looking like operational categories:
| Root cause | What it does to cash |
|---|---|
| Material shortage | You bought inventory that can't convert -- and you'll buy safety stock to cover it, raising days-sales-of-inventory permanently |
| Supplier delay | Your supplier's schedule now sets your inventory days; the fix is a terms-and-lead-time negotiation, not a shop-floor fix |
| Capacity constraint | Work-in-process ages while it waits for a cell -- inventory days again, plus deferred revenue recognition |
| Quality hold | The most expensive inventory in the building: fully converted, fully costed, not shippable |
| Scheduling decision | You chose which receivable to delay. Usually without pricing the choice |
That last row is the one worth sitting with. Every reprioritization is a decision about which customer's invoice gets to age. Almost nobody makes it that way, because the scheduler is optimizing a queue and the collections consequence lands in a different department's report two months later -- the same disconnect that shows up on the finance side in reducing your cash conversion cycle: a lever gets pulled in one department and the cash consequence gets read in another.
This is also why "improve OTD" initiatives so often produce a better number and a worse business. Which brings us to the money the metric is structurally incapable of showing you.
The saves cost more than the misses
The following is a composite, illustrative pattern -- not a specific engagement. It's assembled from what operators and advisors in discrete manufacturing routinely describe, and it's common enough that it's worth checking against your own numbers.
A manufacturer runs an on-time-delivery improvement push. It works, on paper: OTD climbs several points over two quarters. Leadership is pleased. Gross margin, over the same two quarters, quietly declines -- and nobody connects the two, because they're different reports with different owners.
What happened is that the organization got very good at rescuing orders. Premium freight to make a date. Saturday overtime to close a gap. Splitting a shipment so the first half lands on time and the balance follows next week. Buying a component from a distributor at a steep premium rather than waiting for the contracted supplier.
Every one of those is a root-cause failure that was successfully prevented from becoming a late delivery. Every one of them cost real money. And not one of them appears anywhere in the on-time delivery rate -- by construction, because the order shipped on time. They're booked to freight, to labor, to purchase price variance, to three different cost centers, and no single line item in the P&L is labeled "what we paid to protect the OTD number."
So the true volume of delivery-schedule failure in most operations is materially larger than the late-order count, and the true cost is split between a customer-facing number (the misses) and a margin-facing number (the saves). Reading only the first one and calling it delivery performance is the single most common measurement error we see in this category.
The practical consequence: root-cause attribution has to run on saved orders too, not just late ones. If a job required expedited freight or unplanned overtime to hit its date, it belongs in the material-shortage or capacity bucket exactly like a miss does -- with its rescue cost attached. Do that, and the ranking of your categories often changes outright, because the causes that get rescued most aggressively are the ones that were never showing up.
The Path: What Root-Cause Attribution Actually Looks Like
Here's a composite, illustrative example of what this looks like when a manufacturer's late-delivery data actually gets categorized rather than left as a single aggregate number -- the kind of pattern discrete-manufacturing operations commonly see once they break it down:
| Root cause | Share of late orders | Illustrative cost impact |
|---|---|---|
| Material shortage | ~40% | Largest single category |
| Capacity constraint | ~25% | Second largest |
| Supplier delay | ~15% | Meaningful, upstream-owned |
| Machine/tooling downtime, scheduling, labor, quality (combined) | ~20% | Long tail, individually smaller |
The exact split varies by operation -- a shop running older tooling will skew toward downtime; a shop with thin supplier redundancy will skew toward material shortage and supplier delay. The point isn't the specific percentages, it's that some category almost always dominates, and until it's isolated, a general "improve OTD" initiative is really just hoping the average gets better without knowing which lever moves it.
Note also that this taxonomy is the same one manufacturing already uses to decompose equipment losses -- availability, performance, quality. A downtime-driven late order and an availability loss on an OEE report are the same event, counted twice, in two systems, by two people who don't compare notes. If you're already measuring OEE on a constrained asset, half the attribution work is done; it just isn't connected to the delivery number yet.
Where this stops being a report and starts being a running job
This is the specific gap between a monitoring dashboard and how we think the work should be done.
A consultant can produce this breakdown. It's a legitimate engagement, it takes a few weeks, and you get a well-argued Pareto chart of your late orders. The problem is what it is the moment it's delivered: a photograph. Your material-shortage share was 40% during the sampling window. Two quarters later a supplier gets re-sourced, a new program launches, tooling ages -- and the ranking has moved, but the deck hasn't. Re-running it means re-hiring.
A BI dashboard has the opposite failure. It never goes stale, but it never concludes anything either. It renders more granular numbers and hands the categorization, the ranking, and the cost attachment back to whoever remembers to open it on Monday. The analysis labor didn't disappear; it moved from a spreadsheet to a filter panel.
What we build is neither. The agents read the order, job, and tooling records the business already keeps in its ERP, apply the taxonomy above to every late and rescued order, rank the categories by cost impact rather than by count, and flag when a specific tool, cell, or supplier crosses a threshold worth acting on. The categorization and the ranking are the agent's job, done on a standing schedule -- so when material shortage stops being your top category and supplier delay takes over, you get told. You don't rediscover it next year during the next initiative.
That distinction is most of the value. The one-time attribution is worth having. The standing attribution is what actually changes how the plant is run, because the thing being tracked -- which cause is costing you most, right now -- is exactly the thing that keeps moving.
Two honest caveats, since this piece is about not fooling yourself with numbers. First, this depends on your ERP data being good enough to attribute against: if late orders carry no reason code and promise-date revisions aren't retained in history, the first work is instrumentation, not analysis. Second, our manufacturing work is founder-led and design-partner stage today -- it isn't a product you sign up for and switch on this afternoon. Finance and healthcare are further along. We'd rather tell you that up front than let you discover it after a demo.
The Prompt
If your on-time delivery rate is a number you track but can't yet explain -- or if you suspect the real cost is hiding in the orders you rescued rather than the ones you missed -- that's worth a conversation before the next quarter's report looks the same as this one.
Two things you can check this week without us: pull your OTD against original commit date alongside OTD against current promise date and see how far apart they are, and total your premium freight and unplanned overtime for the same period. If the gap is wide and the number is large, you don't have a delivery problem you can fix with a dashboard.
Performis works with manufacturers to build exactly this kind of root-cause, dollar-quantified breakdown from data you already have -- talk to us about an Operations Intelligence assessment.
More on the cash side of this problem: see our guide on reducing your cash conversion cycle, or browse the full Business Solutions library.