How to choose a retail supplier on the worst available information
You choose a retail supplier before you have any performance data about them, which is how a 62-store grocery operator ends up with an incumbent running 82.3% on-time in-full. That is the structural problem. Every measure that actually predicts whether this relationship works, fill rate, velocity against the set, how they behave when supply is tight, only exists after you have already committed shelf to them. Selection is the one supplier decision made entirely on proxies.
Which means the goal is not to pick perfectly. It is to pick with proxies that correlate with the outcomes you will later grade on, and to structure the first two quarters so a bad pick is cheap to reverse. When our 62-store operator replaced a center-store snack supplier, the winning bid was not the lowest quoted cost. It was second-lowest by 2.1 points of margin and won on a service commitment that turned out to be worth far more than the difference: the incumbent had been running 82.3% on-time in-full, and two points of margin does not buy back a shelf that is empty one week in six.
Score the decision, do not discuss it
The same discipline that makes a supplier scorecard work applies before the relationship starts. Weight the criteria before you see the bids, or the weights will quietly rearrange themselves to justify whoever presented best.
| Criterion | Weight | What you are actually assessing |
|---|---|---|
| Landed cost | 25 | Delivered cost per selling unit, not case cost |
| Service capability | 25 | Proven OTIF at comparable accounts, DC coverage, lead time |
| Category fit | 20 | Does the range fill the gaps in your set, or duplicate what sells |
| Commercial terms | 15 | Payment terms, promotional support, return and damage policy |
| Operational maturity | 10 | EDI capability, data quality, case pack sanity |
| Growth partnership | 5 | New-item pipeline, category insight, willingness to co-plan |
Landed cost and service tie at 25 deliberately. Cost is the criterion everyone over-weights because it is the only one that arrives as a hard number in the bid document. Service arrives as a claim, so it feels softer and gets discounted, even though it is the criterion that produces most of the regret.
Landed cost is not case cost
The single most common selection error is comparing case costs across bids whose case packs, delivery frequency, and minimum orders differ. Normalize to delivered cost per selling unit before comparing anything.
Three bids for the same shelf-stable set:
| Bid | Case cost | Units/case | Delivery | Min order | Landed $/unit |
|---|---|---|---|---|---|
| Supplier A | $28.80 | 24 | 2x/week, free | 15 cases | $1.20 |
| Supplier B | $26.40 | 24 | 1x/week, $45 | 40 cases | $1.14 |
| Supplier C | $31.20 | 30 | 2x/week, free | 12 cases | $1.04 |
Supplier B has the lowest case cost and is not the cheapest. Supplier C looks most expensive per case and lands cheapest per unit because of the deeper case pack. But the interesting line is B's 40-case minimum against weekly delivery: for a slow item in a small-format store, that minimum forces roughly six weeks of supply onto the shelf at every order, which is how a good unit price becomes markdown. The center-store departments in our chain already run 6.3 weeks of supply, and a supplier whose minimums push that higher is buying you overstock.
The rule: normalize to landed cost per selling unit, then check the minimum against your slowest store's rate of sale. A cost advantage that only exists at volumes your small stores cannot turn is not a cost advantage.
Assessing service before you have any service data
You cannot measure a prospective supplier's fill rate against you. You can do three things that correlate.
Ask for OTIF at named comparable accounts, defined your way. Not "we run 98% fill." Ask for on-time in-full against original order, with the promise-date logic spelled out, at two accounts of similar size and format. A supplier who cannot produce that number in your definition either does not measure it or does not like the answer, and both are informative. The definitional gap between ship fill and OTIF is where the flattering numbers hide.
Ask what happens when they are short. This is the highest-signal question in the whole process, because allocation behavior under constraint is exactly what you will experience and never what gets discussed. Ask them to describe the allocation rule. Suppliers who allocate by historical volume will put a 62-store account behind their national accounts every time supply tightens, which is precisely the failure mode that produced Harbor's 61% short-at-confirmation rate.
Call the references they did not give you. The provided references are selected. Ask a category peer at a non-competing chain. Two questions: did they hit their service numbers, and what happened the first time there was a real problem.
Category fit: does the range fill a gap or crowd the set
A supplier whose range duplicates what already sells adds SKU count without adding demand, and the tail you create will show up in the next assortment review as something to cut. Before committing, map their proposed range against the gaps in your current set rather than against your current best sellers.
The useful test is incremental: for each proposed item, what shopper need does this serve that the set does not already serve? If the answer is "it is a cheaper version of the item at position 4," you are trading margin for cannibalization, and the honest forecast is not incremental sales. If the answer names a segment, a price tier, or a dietary attribute the set genuinely lacks, that is real. The assortment optimization discipline applies to what you add, not only to what you cut.
Structuring the first two quarters so a bad pick is cheap
Selection error rates are not going to zero, so build the exit into the entry.
Authorize narrow and deep rather than broad and thin. Six items in all 62 stores gives you a clean read on velocity and service. Twenty items in twelve stores gives you neither, because no item accumulates enough weeks to be measurable and the service sample is too small to distinguish a bad supplier from a bad month.
Set the first scorecard date before the first delivery. A review at week 13 with agreed metrics is a different conversation than a review triggered by a complaint, and it means the supplier is optimizing for the measure from day one.
Agree the exit terms while everyone is optimistic. What happens if OTIF is under 92% for two consecutive months. Nobody negotiates that well after it has happened.
Scoring three bids without arguing about it
Weights set in advance are only useful if the scoring is done before the discussion. The mechanic that works: each evaluator scores independently against the published criteria, scores are collected, and only then does the group meet. Scoring in the room converges on whoever speaks first.
Here is the snack-set replacement scored out. Each criterion is scored 0-10, then multiplied by its weight.
| Criterion | Weight | Supplier A | Supplier B | Supplier C |
|---|---|---|---|---|
| Landed cost | 25 | 7 (17.5) | 8 (20.0) | 9 (22.5) |
| Service capability | 25 | 9 (22.5) | 5 (12.5) | 8 (20.0) |
| Category fit | 20 | 8 (16.0) | 6 (12.0) | 7 (14.0) |
| Commercial terms | 15 | 6 (9.0) | 9 (13.5) | 7 (10.5) |
| Operational maturity | 10 | 8 (8.0) | 7 (7.0) | 6 (6.0) |
| Growth partnership | 5 | 7 (3.5) | 6 (3.0) | 8 (4.0) |
| Weighted total | 76.5 | 68.0 | 77.0 |
Supplier C wins by half a point over A, and that margin is the honest result: the framework says these two are effectively tied and B is not close. When the answer lands inside the noise, the tiebreak should be an explicit judgment about which risk you would rather carry, not a fourth decimal place. Here it went to C on landed cost, with A held as the named fallback if C's service commitment slipped in the first two quarters.
Supplier B is the instructive one. Best commercial terms on the board, second-best landed cost, and it finishes eight points back because service scored a 5. B was the bid that quoted 98% fill without being able to define it, which is the exact signal the service question is designed to surface. A selection process that weighted cost and terms at 60 points instead of 40 would have chosen B, and the chain would have spent two quarters discovering what one question had already revealed.
Doing this in Scout
The part of selection Scout changes is the part that happens after the pick: how fast you find out whether it worked. Most chains discover a supplier problem somewhere between month four and month nine, because the data to see it sooner exists but is spread across receiving, POS, and a promotional calendar that nobody joins until a quarterly review forces it.
Scout stands the scorecard up before the first delivery. The new supplier lands in the same graded view as the incumbents, with OTIF, velocity index against category median, and the on-shelf-availability split computed on the same definitions as everyone else. Week 13 arrives with the number already built, and the trailing weeks show whether the service commitment in the bid survived contact with your order pattern.
That also makes the counterfactual visible. When you switch suppliers, the question that decides whether it was the right call is whether the category improved, not whether the new supplier is pleasant to work with. Holding the category view across the switch answers it.
Scout grades supplier performance and category outcomes. It is not a sourcing or procurement system: it does not run tenders, hold contracts, or issue purchase orders.
Revisiting the choice
The decision to choose a retail supplier is not permanent, and treating it as permanent is how chains end up carrying a poor performer for four years. Put a formal reconsideration on the calendar at twelve months, whether or not anything has gone wrong.
Twelve months is long enough for real performance data and short enough that switching costs have not compounded into an argument for inaction. The reconsideration does not have to produce a change. It has to produce an explicit decision, recorded, that this supplier is still the right one on the criteria you originally weighted. Most of the time the answer is yes, and the fifteen minutes are cheap. The rest of the time it surfaces a decision everyone had been avoiding.
Summary
- Weight the selection criteria before you see the bids, and let landed cost and service capability tie, because service is the one that produces the regret.
- Normalize every bid to delivered cost per selling unit, then test the minimum order against your slowest store: a price that only works at volumes you cannot turn is overstock in advance.
- Ask what the supplier does when they are short. Allocation behavior under constraint is the highest-signal answer available before you have data.
Further reading: supplier performance metrics defines the measures you will grade them on, and vendor management best practices covers running the relationship after the pick.