Skip to content

See a demo

30 minutes with Sasha Zhang · video link on confirmation

Loading scheduler…

Store clustering: how to build and test clusters

The question a chain average cannot answer

A merchandising meeting hits the same wall every year. Somebody asks which stores should get the six-foot beverage set and which keep four, and the only number on the table is the chain roll-up: packaged beverage is 25.5% of sales. That figure is true of the chain and of almost no store in it. In the 48-store convenience operator I use throughout this page, the same category runs 22% in eleven stores and 31% in nine others.

Store clustering is the method that turns that spread into a small number of groups a merchandising team can actually plan for: you score every store on a handful of features drawn from its own point-of-sale history, group the stores whose scores sit close together, and then build one plan per group instead of one plan for the chain or 48 plans for 48 stores. Done well it collapses an unmanageable number of decisions into three or four. Done badly it produces a tidy dendrogram, four names nobody can define, and a set of planograms that all look the same.

The difference is almost entirely in three choices: what you cluster on, how many clusters you cut, and whether you tested them before anyone built a plan.

What to cluster on, and what to leave out

Every feature you feed the model is a claim that stores differing on it should get different plans. That test kills most of the candidate list immediately. Four families survive, and all four come out of the retailer's own transaction data.

Category mix share. Each store's sales by category expressed as a percentage of that store's own total. This is the workhorse feature, and the share is doing the work rather than the dollars. Cluster on dollars and you rediscover store size, which you already knew from the sales report. Cluster on within-store share and you learn that a store selling half as much as its neighbour still sells a third of it as foodservice.

Rate of sale by category. Units per store per week, normalised by selling area or by linear feet where you have it. Mix share tells you what a store's customers buy relative to each other; rate of sale tells you how hard the category works per foot of space, which is the number a space decision actually needs. The two disagree more often than people expect, and the disagreement is usually the interesting part.

Price sensitivity. From the same POS: what share of a store's units move on deal, and how far the store's average selling price in a category sits from the chain. A store where 41% of packaged beverage units ring at a promoted price is a different pricing animal from one at 18%, and that is a fact about the register, not a theory about the shopper.

Basket composition. Items per basket, the share of single-item baskets, and the attachment rate between the categories you care about. A store where 62% of baskets are one item is a different merchandising problem from one at 38%. Market basket analysis is the underlying read, and it is worth doing before clustering rather than after, because a couple of strong pairs are usually more informative than ten weak features.

What to leave out is shorter but more important. Leave out anything you cannot measure in the data you hold. Square footage, format and delivery schedule are fine as descriptors and belong in the cluster profile, but if you put them in the distance calculation you get clusters that reproduce your real-estate file. And leave out anything that would require observing shoppers rather than transactions. The clusters on this page are built from what stores sell, and nothing in the method sees who walked in or where they came from.

Preparing the features so distance means something

Clustering is a distance calculation, and a distance calculation quietly weights whatever has the largest numbers. Four preparation steps decide whether the output means anything.

  1. Convert to within-store shares first. Every category feature becomes a percentage of that store's own total. Do this before anything else or the biggest store dominates every distance.
  2. Standardise each feature across stores. After the shares, z-score each feature column: subtract the mean across stores and divide by the standard deviation. Without this step a category with a wide spread swamps one with a narrow spread purely because of its variance. In this chain, beer and wine swings across a wide band store to store while tobacco and other tobacco products barely moves, so raw shares would have let beer decide almost every store's membership on its own.
  3. Cut collinear features. If two features correlate above about 0.85 they are one feature entered twice, and entering it twice doubles its weight. Keep the one a merchant can act on.
  4. Fix the input window. Use 52 consecutive weeks so seasonality averages out, and exclude stores that were not open for all of it. Also exclude the remodel windows: a store closed six weeks for a foodservice buildout has a category mix for that year that describes construction, not demand. In this chain three stores came out of the input on those grounds and were assigned to clusters afterwards by nearest centroid.

Choosing how many clusters

There are two constraints, and the second one usually binds.

The statistical constraint is the familiar one. Run the clustering at a range of k and watch within-cluster sum of squares fall. On this chain it drops 41% going from two clusters to three, 18% from three to four, and 6% from four to five. That is a clean elbow at four. The silhouette score, which rewards tight, well-separated groups, actually peaked at six.

The operational constraint is why six lost. Six clusters split the estate into groups of 3 and 4 stores at the bottom end, and every cluster is a plan somebody has to write, maintain, reset and re-audit. Two rules keep this honest:

  • A cluster has to be big enough to earn a plan. Roughly 10% of the estate is a reasonable floor. Below that, the cost of maintaining a separate planogram, order guide setting and promo variant exceeds what the distinction is worth.
  • The number of clusters is per decision, not per chain. Nothing requires the assortment partition and the price-zone partition to be the same. This chain runs four assortment clusters and two price zones, and both are correct, because the features that separate assortment decisions are not the features that separate pricing decisions.

Worked example: 48 stores, four clusters

Here is the output, expressed the way a merchant reads it: each cluster's mean category mix, with the chain column for comparison. Every column is a percentage of that group's own sales, so each column sums to 100.

CategoryBreakfast-led (11)Full-basket (17)Snack-and-drink (9)Beer-led (11)Chain
Foodservice31%14%19%12%18.4%
Packaged beverage27%24%31%22%25.5%
Salty snacks9%13%17%11%12.4%
Candy6%9%11%7%8.2%
Beer and wine5%18%6%24%14.1%
Tobacco and OTP14%13%11%12%12.6%
Packaged grocery and other8%9%5%12%8.7%
Breakfast-led (11)Full-basket (17)Snack-and-drink (9)Beer-led (11)Chain averageFoodservice31%14%19%12%18.4%Packaged beverage27%24%31%22%25.5%Salty snacks9%13%17%11%12.4%Candy6%9%11%7%8.2%Beer and wine5%18%6%24%14.1%Tobacco and OTP14%13%11%12%12.6%Packaged grocery and other8%9%5%12%8.7%
Every amber line is the chain average, and it describes no cluster: foodservice runs 31% breakfast-led against 12% beer-led, a factor of 2.6 (worked example)

The chain column is the store-count-weighted mean of the four clusters and sums to 100.0% before rounding: foodservice, for instance, is (11 x 31 + 17 x 14 + 9 x 19 + 11 x 12) / 48 = 882 / 48 = 18.4%.

Two numbers carry the whole page. Foodservice is 31% of sales in the eleven breakfast-led stores and 12% in the eleven beer-led ones, a factor of 2.6 on a chain average of 18.4% that describes neither group. Beer inverts the same stores: 24% against 5%, a factor of 4.8. This chain had been running one foodservice plan and one beer plan across all 48.

What changed as a result was capital, not reporting. A second brewer and a hot case reset went to the eleven breakfast-led stores, where foodservice already carried nearly a third of sales and the equipment was the constraint. The same money in the beer-led stores went to cooler doors. Before the clustering, both proposals had been sitting in the same queue competing on chain-average category growth, where neither could win.

Validating a store clustering before you build plans on it

This is the step that gets skipped, and it is the only one that tells you whether you have found structure or drawn boundaries through noise. Four tests, in order of how quickly they fail:

TestWhat to computeThis chainVerdict
Face validityCan a merchant name each cluster in one phrase, without the numbers?4 of 4 namedPass
SeparationWithin-cluster spread of the decision metric vs chain-wide spread2.6 pts vs 7.7 ptsPass
Temporal stabilityRebuild on the prior 52 weeks and count stores that keep membership43 of 48, or 89.6%Pass
Decision differenceDo the resulting plans actually differ, category by category?5 of 7 identical for two clustersPartial fail

Face validity is first because it is free and it fails loudly. If the merchandising lead cannot look at a cluster's profile and say "those are the breakfast stores" without being walked through the feature table, the grouping will not survive its first planning meeting whatever its silhouette score says.

Separation is the arithmetic version of the same question. Foodservice share has a standard deviation of 7.7 points across all 48 stores. Inside the four clusters the mean standard deviation is 2.6 points, a 66% reduction in the spread of the thing you are about to make decisions on. If clustering does not shrink that spread materially, it is decoration.

Temporal stability is the one that catches clusters fitted to a year rather than to a chain. Rebuild the whole thing on the previous 52 weeks and count how many stores land in the corresponding cluster. Below roughly 80% agreement the partition is describing events, not stores, and the five movers here were worth reading individually: three had genuinely changed after a foodservice buildout, and two sat almost exactly between two centroids and will flip every year no matter what you do.

Decision difference is the test that produced the useful finding on this chain. Full-basket and beer-led stores took the same planogram in five of the seven categories, differing only in beer and packaged grocery. For planogram purposes those two are one cluster, which is why the chain runs three assortment plans off a four-cluster partition. Two clusters that receive the same instruction are one cluster with extra maintenance.

How often to re-cluster

Cluster membership that changes every quarter makes cluster performance unreadable, because you can never tell whether the cluster improved or the membership did. The cadence that works:

  • Recompute the features quarterly, and monitor each store's distance to its own centroid. The stores in the top decile of that distance are a review list, not a reassignment.
  • Re-cut the clusters annually, on a fixed 52-week window, ahead of the planning cycle that will use them.
  • Move a store only after two consecutive quarters outside its cluster, or immediately on a structural event such as a remodel or a foodservice buildout, which is a known cause rather than drift.

Where store clustering goes wrong

Clustering on dollars. The most common failure and the hardest to spot, because the output looks reasonable. Feed raw category dollars into k-means and you get large, medium and small stores with category labels attached. Everything the clustering told you was already in the sales ranking.

Letting the label become evidence. A cluster name is a mnemonic for a mix pattern and nothing else. Somebody in the third meeting will start saying the beer-led stores skew younger, or that the breakfast-led ones are commuter sites, and both statements are inventions. The data behind the cluster is what the registers rang. It contains no information about who shopped or where they came from, and the moment a cluster label starts carrying demographic or catchment freight, the clusters have become a story.

Too many clusters to execute. Nine clusters produce nine planograms, of which four get maintained and five drift into whatever the store did last year. The chain then concludes clustering does not work, when what did not work was nine of them.

Building on a promo-heavy window. If the input year included a beer program that ran in 20 of the 48 stores, beer share is partly a record of where you ran the promo. Either use a window without a lopsided program in it or cluster on baseline rather than total sales.

Treating the cluster as the answer. A cluster is a prior, and a good one, but the store-level exception list still has to run. Half the value of a clean partition is that an outlier inside a tight cluster is now visible: a breakfast-led store doing 19% foodservice against its cluster's 31% is a specific, actionable question, and before the clustering it was invisible against a chain average of 18.4%.

Reusing the assortment clusters for everything. Replenishment targets cluster on volume, format and delivery schedule, which is a different partition built from different features. Retail replenishment guidance covers that one. Seasonal amplitude is a third. Pushing one partition through every decision is how a useful cluster set stops being trusted.

Where Scout fits

Every feature on this page is a POS aggregation: within-store category shares, rate of sale, promoted-unit share, basket composition. Scout computes those from the retailer's own transaction data, holds them current as new weeks land, and carries the store-to-store comparison and the distance-to-centroid drift monitor that makes the annual re-cut a review rather than a project. The clusters that come out of it feed assortment, space and seasonal planning work.

One boundary, stated plainly. This is a first-party POS read and nothing more. Scout holds no mobility, footfall, catchment or visit-panel data, so it does not observe store traffic, trade areas, or where a store's shoppers come from, and none of the clusters above are trade-area clusters. POS is a record of transactions, not of people, which is also why it supports no claim about shopper demographics or attitudes.

The short version

  • Store clustering groups stores on features from their own POS so a merchandising team writes three or four plans instead of one or forty-eight.
  • Cluster on within-store category mix share, rate of sale, promoted-unit share and basket composition. Convert to shares, then z-score, or the largest category decides everything.
  • Pick the number of clusters on the elbow and then cut it down to what your team can maintain. A cluster below about 10% of the estate rarely earns its plan, and the right number differs by decision.
  • Validate on four tests before anyone plans against it: a merchant can name each cluster, within-cluster spread drops materially against chain-wide spread, membership is stable when rebuilt on a prior year, and the resulting plans genuinely differ.
  • Re-cut annually, monitor drift quarterly, and move a store only on two consecutive quarters out of position or a structural change.

Want this as a Google Sheet?

Drop your email and we'll send the worked example.

Book a demo with your data