Every supplier metric gets disputed once
The first time you put supplier performance metrics in front of a vendor, they will dispute the calculation. Harbor Provisions told our 62-store operator their fill rate was 96% during the same thirteen weeks our number said 82.3%. This is not bad faith. It is that most of these measures have three defensible definitions, and the supplier is using the one their own system reports. Harbor Provisions told us their fill rate was 96% during the same thirteen weeks our number said 82.3%, and both numbers were computed correctly. They were measuring cases shipped against cases confirmed. We were measuring cases received on the promised date against cases ordered.
The difference is everything. Their number excluded the lines they had already cut at confirmation, which is exactly where the failure was happening. This page is about defining each metric so the calculation survives that meeting, because a metric that collapses under one round of questioning is worse than no metric: it costs you credibility you need for the next conversation.
Fill rate is three metrics wearing one name
Write down which of these you mean before you report a fill rate.
| Measure | Numerator | Denominator | What it hides |
|---|---|---|---|
| Confirmation fill | Cases confirmed | Cases ordered | Nothing, but it measures the supplier's promise, not delivery |
| Ship fill | Cases shipped | Cases confirmed | Everything cut at confirmation |
| Receipt fill | Cases received | Cases ordered | Timing |
| OTIF | Cases received on the promise date | Cases ordered | Nothing, which is why it is the one to use |
Ship fill is the number most suppliers quote, because it is measured after the cut has already happened. If a supplier confirms 800 of a 1,000-case order and ships all 800, their ship fill is 100% and your shelf is short 200 cases. Receipt fill against the original order catches that, and OTIF catches it plus the delivery that showed up four days later.
The formula worth standardizing on:
OTIF = cases received on or before the promise date, capped at ordered
-------------------------------------------------------------
cases originally ordered
The cap matters and gets forgotten. Without it, an over-shipment of one line offsets a short on another and the aggregate looks healthy. A supplier who ships you 300 cases of the fast item and none of the slow one has not delivered 100% of anything useful.
Splitting the failure so it is actionable
An aggregate OTIF tells you a supplier is failing. It does not tell them what to fix, and a supplier who cannot see the mechanism will fix nothing. Decompose every miss into late, short, or both.
Harbor Provisions across thirteen weeks, 62 stores. The shares below decompose the 42-case weekly gap against the 94.6% shelf benchmark, not misses against a theoretical 100%:
| Failure mode | Share of the 42-case gap | Cases/week | Reads as |
|---|---|---|---|
| Short at confirmation | 61% | 26 | Allocation logic: we lose when supply is tight |
| Late, full quantity | 24% | 10 | Transport or scheduling, recoverable inside the week |
| Short at delivery | 11% | 5 | Pick accuracy at their DC |
| Both late and short | 4% | 2 | The genuinely broken orders |
Sixty-one percent concentrated at confirmation is a specific, nameable problem: when Harbor is constrained, their allocation puts our 62 stores behind larger accounts. That is a commercial conversation with a commercial answer (committed volume, earlier ordering, or a different service tier). Without the split, the same data supports a vague complaint about service that the supplier can absorb and outlast.
Velocity index, and why it is indexed
Raw velocity in dollars per store per week is not comparable across a dairy supplier and a shelf-stable supplier, so a raw target either flatters fast categories or punishes slow ones. Index each item to the median of the category it competes in, then roll up to the supplier weighted by dollar contribution.
Item velocity index = (item $ / carrying store / week)
----------------------------------- x 100
(category median $ / store / week)
An index of 100 means the item performs at the category median for the stores that carry it. The rollup weight has to be dollar contribution rather than a simple average across items, or a supplier's twelve slow tail items drown their two strong ones and the number stops describing the business.
Two traps in the denominator. First, carrying stores, not total stores: dividing by all 62 when an item is only authorized in 28 produces a velocity that says "slow" when the truth is "narrowly distributed." That confusion between rate and reach is the same error the velocity, share, and TDP decision tree exists to prevent. Second, exclude the weeks an item was out of stock at that store, or on-shelf-availability problems silently depress the velocity line and you end up penalizing a supplier twice for one failure.
On-shelf availability is partly your fault
This is the line that most needs an honesty clause. On-shelf availability measures whether the shopper found the item, and the causes split between the supplier (they did not ship it) and you (it sat in the back room, or the planogram was never reset, or the order was never placed). Charging the whole number to the supplier is unfair and, more practically, gets the metric dismissed.
Split it at the receiving dock. If the store received the cases and the shelf was still empty, that is yours. If the store never received them, that is theirs. Report both halves and only score the supplier on their half.
In the Harbor case the split ran roughly 70/30 supplier-to-store, which was the strongest evidence in the review, because it survived the obvious counterargument. When a supplier says "your stores are not putting it out," a number that already concedes the 30% and still shows 70% on their side ends that line of defense.
For the mechanics of measuring availability itself, the on-shelf availability entry covers the detection methods and their false-positive rates.
Promo ROI, computed against a defensible baseline
Promotional ROI is the most contested line on any scorecard because the entire answer depends on the baseline, and the baseline is an estimate. State the method explicitly on the card.
Promo ROI = (incremental units x unit margin) / total promotional cost
Incremental = actual promoted units - baseline units
Baseline should be the trailing non-promoted weeks for that item at that store group, seasonally adjusted, excluding any week within two weeks of a prior promotion on the same item. That exclusion window is the part people skip, and skipping it inflates every result, because the post-promotion pantry-loading dip drags the baseline down and makes the next promotion look better than it was.
A promotion that returns below 1.0x destroyed margin. That sounds obvious and is routinely ignored, because the dollar line goes up during a promotion and the dollar line is what gets reported. Summit Beverage ran the best promo ROI in the vendor base at 1.62x and still finished third overall on the supplier scorecard, which is the correct outcome when service is soft.
Data errors that quietly break your supplier performance metrics
Every metric above sits on the same handful of joins, so a defect in the underlying data corrupts several lines at once and does it silently. Four are worth checking before you publish a card for the first time.
Case pack changes with no item-file update. A supplier moves from 24-count to 30-count, the item file still says 24, and every subsequent unit conversion is wrong by 25%. This corrupts velocity, on-hand, and weeks of supply simultaneously, and it does not look like an error: the numbers stay plausible. Check the unit conversion whenever a supplier's case cost changes, because the two usually move together.
Store openings, closings, and remodels in the denominator. Velocity per carrying store breaks when a store is closed for a six-week remodel and stays in the store count. The item did not slow down; the store was shut. Exclude stores from the denominator for any week they were not trading.
Promotional periods bleeding into the baseline. Covered above for promo ROI, but it also distorts the velocity index, because a supplier with a heavy promotional calendar shows inflated velocity if promoted weeks are counted as normal. Either exclude promoted weeks from the index or compare like to like.
Substitutions recorded as fills. When a supplier ships a similar item in place of the one ordered and receiving accepts it against the original line, OTIF records a success and the shelf still has the wrong product. This one is particularly worth auditing because it inflates precisely the supplier who is managing their own shortage well.
None of these are exotic. All four were present in the first version of the card at our 62-store operator, and fixing them moved two suppliers by more than three points of weighted score, which is the difference between amber and green.
Putting a review-proof number on the table
The pattern that makes these metrics durable in a meeting is the same for all four:
State the formula on the card, not in an appendix. Publish the window and the store set. Show the split (late versus short, supplier versus store) rather than the aggregate alone. And translate the percentage into cases and stores before you walk in, because 42 cases a week concentrated in 28 stores is a fact a supplier can act on and 82.3% is a fact they can argue with.
Doing this in Scout
Each of these metrics is a join across systems that do not naturally talk: the purchase order, the receipt, the POS scan, and the promotional calendar. The reason buyers fall back on the supplier's own reported number is that assembling the honest version takes half a day and breaks whenever a case pack changes.
Scout computes these as standing definitions over your receiving and POS feeds. OTIF is defined once, with the cap and the promise-date logic baked in, so every supplier is graded the same way and the definition does not drift between analysts. The late/short decomposition and the supplier-versus-store availability split are slices of the same view rather than separate rebuilds. Because the definitions live in one place, when a supplier disputes the calculation you can show them the definition instead of reconstructing the spreadsheet.
Scout measures and grades supplier performance. It is not your purchasing or receiving system, and it does not replace the ERP that raises the order.
Publishing the definitions
Whatever you decide, write the supplier performance metrics down and publish them to the vendor base. A one-page definition sheet, covering the formula, the measurement window, the store set, and the exclusions, ends most disputes before they start, because the argument moves from "your number is wrong" to "your definition is different from mine," which is a conversation with an answer.
The sheet also protects you internally. Analysts change, and an OTIF that quietly means something different this year than last is worse than no trend at all. A published definition is the only thing that makes a two-year comparison meaningful.
Summary
- Fill rate has three defensible definitions and suppliers quote the flattering one. Standardize on OTIF against the original order, with an over-shipment cap.
- Decompose every miss into late, short, or both. Harbor's 61% concentration at confirmation named an allocation problem that an aggregate never would have.
- Index velocity to the category median over carrying stores, and split on-shelf availability at the receiving dock so the supplier is scored only on their half.
Further reading: building a retail supplier scorecard assembles these into a weighted card, and the retail purchasing process shows where in the cycle each measurement is actually captured.