Skip to content

Building a retail supplier scorecard

The supplier review that has no scorecard

A supplier review without a retail supplier scorecard is a meeting where the best storyteller wins. I have sat through a 45-minute quarterly with a snack vendor whose service had been visibly bad for a quarter, and watched them spend forty of those minutes on a new-item deck. Nobody in the room could say what their fill rate had actually been, so nobody could say the thing that needed saying. The review ended with an agreement to "keep an eye on service," which is what a room says when it has no number.

The number existed. Across a 62-store Northeast grocery operator, that supplier had run 82.3% on-time in-full over thirteen weeks while the shelf average sat at 94.6%. That is not a rounding error or a bad month. It is roughly one delivery in six arriving late, short, or both, and every one of those is a hole in a planogram that a store team filled with something else or left empty. A scorecard is the artifact that puts that 82.3% on the table in minute two instead of never.

What belongs on the retail supplier scorecard

The failure mode is not too few metrics. It is too many. A scorecard with twenty-two rows is a data export, and a data export does not change a supplier's behavior because it does not tell them what to fix first. The discipline is to pick the handful of measures that a supplier can actually move inside one review cycle, weight them, and let everything else live in the appendix.

Five lines is the working set.

Fill rate / OTIF30%Sell-through velocity25%On-shelf availability20%Promo ROI15%Admin accuracy10%
A five-line scorecard, weighted to 100. Service and velocity carry 55 points between them because they are the two the buyer can act on inside one review cycle.

Service and velocity carry 55 of the 100 points on purpose. Those two are the measures where a supplier's own decisions dominate the outcome. Fill rate is a function of their production planning and their allocation logic when supply is tight. Velocity is a function of whether the item earns its facings. The other three matter, but they are either slower to move (on-shelf availability depends partly on your store execution) or smaller in dollar terms (administrative accuracy is real money, but it is basis points, not points).

Two notes on the weighting that get argued every time.

Promo ROI at 15 is deliberately modest. Promotional performance is the loudest number in most supplier conversations and it deserves less weight than its volume of discussion suggests, because a single well-timed feature can swing it and because attribution is genuinely contested. Weighting it at 15 keeps it on the card without letting one promotion launder a bad service quarter.

Administrative accuracy earns its 10 points. Invoice mismatches, wrong case packs, and pricing that does not match the agreement cost real hours in accounts payable and real margin at the register. Ten points is enough that a supplier who is sloppy here cannot score in the top band, which is exactly the incentive you want.

Grading the suppliers

Here is the working set applied across the five largest suppliers in the center store and perimeter mix, trailing thirteen weeks, all 62 stores. Each raw measure is converted to a 0-100 sub-score before weighting, so a fill rate of 97.4% against a 98% target scores 99, not 97.4.

SupplierFill/OTIF (30)Velocity (25)OSA (20)Promo ROI (15)Admin (10)Weighted
Northwind Dairy29.222.518.611.49.591.2
Cedar Creek Bakery28.720.017.812.09.087.5
Summit Beverage28.218.816.413.28.084.6
Coastal Snacks27.521.315.29.87.581.3
Harbor Provisions24.714.511.610.56.067.3

The spread is the output. Northwind at 91.2 and Harbor at 67.3 are not two suppliers having a similar quarter, and before the scorecard existed they were discussed in the same tone of voice. Note also that Summit Beverage has the best promo ROI on the board (13.2 of 15) and still finishes third, which is precisely the behavior the weighting was built to produce. A strong promotion does not buy forgiveness for a soft velocity line.

Northwind Dairy97.4%Cedar Creek Bakery95.8%Summit Beverage94.1%Coastal Snacks91.6%Harbor Provisions82.3%
OTIF across 62 stores, trailing 13 weeks. The 15-point gap between Harbor Provisions and the shelf average is the entire conversation the scorecard exists to start.

The one metric that carries the review

If you only ever track one line, track fill rate, and track it as on-time in-full rather than as a raw fill percentage. The distinction matters more than it sounds. A supplier who ships 100% of the cases three days late scores well on fill and badly on OTIF, and the three days late is what emptied the shelf. A supplier who ships on the promised day at 88% of the ask scores well on timing and badly on completeness, and the missing 12% is what emptied the shelf. OTIF is the measure that refuses both excuses.

The 15-point gap in the chart above between Harbor Provisions at 82.3% and the shelf average is worth translating into shelf terms, because a percentage does not motivate anyone. Harbor ships roughly 340 cases a week into the chain across their 41 active items. At 82.3% OTIF against a 94.6% benchmark, the gap is about 42 cases a week that either arrive late or do not arrive. Spread across 62 stores that is not dramatic per store, but concentrated as it actually is in the 28 stores that carry the full Harbor set, it is a recurring hole in the same planogram positions every week.

That framing is what turns the scorecard from a report card into a negotiation. "Your OTIF is 82.3%" invites a discussion about measurement methodology. "Forty-two cases a week are not making it to shelf, mostly in these 28 stores, mostly in these six items" invites a discussion about their allocation logic.

Setting targets that are not just last year plus one

A scorecard needs a target band per line or the scores are decoration. The temptation is to set the target at the current chain average and call it done, which rewards the middle of the pack for being average and gives your best supplier nothing to hold.

Set three bands instead, and set them against what good actually looks like in the category rather than against your own mean.

LineTarget band (green)Watch (amber)Action (red)Why this line sits here
Fill rate / OTIF96%+92-96%Below 92%Below 92% the store team stops trusting the order and buffers
Velocity index105+90-105Below 90Indexed to category median, so it is comparable across sets
On-shelf availability97%+94-97%Below 94%Below 94% the gap is visible to the shopper on a normal trip
Promo ROI1.4x+1.0-1.4xBelow 1.0xBelow 1.0x the promotion destroyed margin
Admin accuracy99%+97-99%Below 97%Every point is invoice-matching labor in AP

The velocity line is indexed rather than absolute for a reason worth stating. Absolute velocity is not comparable across a dairy supplier and a shelf-stable supplier, so an absolute target either flatters the fast categories or punishes the slow ones. Indexing each supplier's velocity to the median of the categories they actually compete in makes the number portable across the whole vendor base, which is the only way one scorecard covers every supplier.

The same logic that governs a category scorecard applies here: at least one line should lead rather than lag. Fill rate leads. It moves before the sales line records the damage, which means a supplier sliding from 96% to 91% is a problem you can raise this month rather than a problem you explain next quarter.

Running the review off the card

The scorecard changes the shape of the meeting, and that is most of its value.

Open on the weighted score and the two lines that moved most since last quarter. Five minutes. This front-loads the conversation onto the graded reality before anyone opens a deck.

Spend the middle on the red and amber lines only. If Harbor is at 67.3, the agenda is fill rate and velocity, and the new-item presentation waits for the back half or the next meeting. This is the part that requires actual discipline, because a supplier who has prepared a launch deck will fight for the room.

Close by writing the next-quarter target for each red line and naming who owns it on both sides. A target with no name attached is a wish.

One structural warning from doing this badly: do not let the scorecard become a quarterly-only artifact. A number that appears four times a year is a number that gets contested four times a year, because each appearance is a surprise. When suppliers can see their own fill rate monthly, the quarterly review stops being a reveal and becomes a checkpoint, and the arguing collapses.

Doing this in Scout

The reason most scorecards die is not disagreement about the metrics. It is that assembling one takes a buyer four hours of exports: sell-through from the POS extract, receipts and shorts from the purchasing system, promo periods from a calendar in someone's inbox, and then a manual join in a spreadsheet that breaks when a supplier changes a case pack.

Scout builds the scorecard as a standing view over the data you already have. The supplier grid above is a saved view: sub-scores computed from your receiving and POS feeds, weights configurable per vendor class, and the whole thing refreshed as the underlying data lands rather than rebuilt by hand each quarter. The velocity index resolves against the category median automatically, so adding a new supplier does not mean rebuilding the comparison set.

Two things that matter in practice. The card is sliceable by store group, so the "mostly in these 28 stores" framing above is two clicks rather than a re-export. And because the same view backs the monthly and the quarterly read, the supplier sees the same number you do, which is the fastest way to end methodology arguments.

Scout is the analytics layer here. It grades supplier performance and shows you where the shelf is losing; it is not your purchasing system and it does not transmit orders or manage contracts.

Rolling it out without a fight

Introduce a retail supplier scorecard to the vendor base before you grade anyone on it. Send the definitions and the target bands a full cycle ahead, invite challenges to the calculation, and fix whatever is genuinely wrong. The suppliers who dispute the method in advance are doing you a favour: every objection raised before the first card is one that cannot be raised as a defence after it.

Grade the first quarter without consequences attached. The first card almost always surfaces data problems rather than supplier problems, and burning credibility on a number that turns out to be a case-pack error is expensive. From the second quarter the card is live, and by then nobody can claim surprise.

Summary

  • A retail supplier scorecard is five weighted lines, not twenty-two, and service plus velocity should carry more than half the weight because they are what a supplier can move inside one cycle.
  • Track fill rate as on-time in-full, and translate the percentage into cases and stores before the meeting, because 42 cases a week in 28 stores starts a different conversation than 82.3%.
  • Set three target bands per line against category-good rather than your own average, and index velocity so one card covers every supplier.

Further reading: supplier performance metrics goes deeper on how each line is computed, and vendor management best practices covers what to do with a supplier who stays red for two quarters running.

Want this as a Google Sheet?

Drop your email and we'll send the worked example.

Book a demo with your data