Overview
Both claims circulate with equal confidence, and neither is usually accompanied by evidence. This study assembles the evidence.
A supermarket shelf price is off-contract market pricing. Nobody negotiated it, no schedule fixes it, and it is set mostly against what the shop across the road is doing. That makes the comparison a competitor benchmarking problem rather than a curiosity about groceries, and it is why the measures on this page are the ones a pricing team would recognise: parity rate, mean absolute gap, promotion frequency and depth, gap persistence, and reference-price integrity.
The first version answered the question for a single day, and the answer was a dead heat. That turned out to be the least interesting thing in the data. The question worth asking is not which chain is cheaper this morning but which lines the two chains are willing to compete on at all, because those are not the same list and the difference is where the money sits.
7.1%
Parity rate in household, against 62.5% in pantry. Same two retailers, same morning.
100%
Reference-price integrity: every advertised “was” price tested against a price observed beforehand.
39 of 44
Promotions that ended returned to baseline. Median depth 41%, median run 7 days.
Building the series
A collector runs daily and writes one immutable CSV per day. That file is the asset: a price snapshot cannot be backfilled, so every morning the collector does not run is a day of history nobody can recover. The warehouse is a build artefact and is not committed to version control. The daily files are, precisely because they cannot be reconstructed.
- 01CollectPython, requests, two retailer JSON APIs, one dated CSV per day
- 02Loadpandas, CSVs landed as a tested dbt source with a freshness check
- 03Stagedbt, pack sizes canonicalised, unit prices rebased to per 100g
- 04Screendbt, coverage and per-line relevance rules applied before any total
- 05Trackdbt snapshot, SCD2 on price, was-price and the promotion flag
- 06Matchrapidfuzz, identical products paired across the two catalogues
- 07Modeldbt, six marts, incremental history, grain tested on every one
Price history is modelled as a slowly changing dimension rather than a pile of daily rows. A product whose price holds for six weeks is one record with a six week validity window, and a price change closes one record and opens the next. Historical days are replayed one at a time, so a change is dated the day the price actually moved rather than the afternoon the model was built.
That structure also gives the right sample size. The collection has produced 26,852 row-days, but a price sitting unchanged for nine days is one event rather than nine, so the effective sample is the 1,918 observed price changes underneath it, across 3,070 products. That is the number every history measure below rests on.
The matched-pair history is incremental: a day’s run touches a day’s rows. That is a claim, so a script checks it, by deleting the last two days, replaying them, and asserting the result is identical to a full rebuild. An incremental model whose answer depends on how many times it has been run is a defect that surfaces months later, in figures already published.
One rule governs everything downstream: a day nobody collected is never filled in. A product missing from a day’s search results does not close its history record, because it has not been discontinued and its price has not changed. It was simply not observed that day. The matched-pair series leaves that day null and flags it, a data test enforces that the flag and the nulls agree, and the dashboard draws those stretches dashed rather than joining them with a confident straight line. Inventing continuity is how a price series starts lying.
The matching layer
Mapping a line to the right competitor SKU is the hard part, and it is the part that transfers. Candidates are generated within the same search term, fuzzy-matched with rapidfuzz, then accepted by tier: national brands need the same brand, a pack size within 2% and a name score of 80 or better. Home brands are matched separately, as substitutes rather than as identical products, because treating a private label as the same SKU is how a benchmark quietly becomes a comparison of two different things.
Assignment is greedy and one to one, and every accepted pair carries its similarity score, so any pair can be pulled up and shown why it was accepted. That is the difference between a rule and an override: the rule decides which lines get benchmarked, and the score is the audit trail. Overrides are what happen when nobody wrote the rule down.
The same discipline applies to what enters a basket line at all. A retailer’s search endpoint is a ranking API, not an inventory feed, and it will return something that is not the product without raising an error. Lines therefore carry identity rules on name and unit-price basis rather than a plausible-price band, because a screen that filters on price can silently discard the price movement the study exists to measure.
Porting the original SQL into dbt came with one hard rule: it could change how the marts are built and not what they contain. Every mart was reconciled row for row against the pre-port output before the old models were deleted. Nothing moved.
The measures
Every figure below is computed in the warehouse from the collected days, and rebuilt from the raw CSVs on every run. The definitions matter more than the numbers, because a reader who cannot tell what a measure counts cannot tell whether it is being counted honestly, so they are stated here rather than left implied.
- 01Matched pair. The same product at both chains: equal brand, pack size within 2%, and a fuzzy name score of 80 or better. Home brands are matched only against each other, as substitutes rather than as the same product. Every accepted pair keeps its score, so any pair can be pulled up and shown why it was accepted.
- 02Parity rate. The share of matched pairs priced identically to the cent at both chains on the same day. Not within a tolerance: equal.
- 03Mean absolute gap. The mean of the absolute dollar difference across matched pairs in an aisle. It weights every pair equally, so it describes the shelf rather than the till, and a dear line nobody buys counts as much as milk.
- 04Promotion episode. A run of consecutive collected days on which the retailer’s own promotion flag is set for a product. Depth is measured against the price this collector observed before the run began, not against the advertised was-price.
- 05Reference-price integrity. The share of advertised was-prices that match a price this collector actually observed on an earlier day. Only episodes with an earlier observation of my own can be tested, so the denominator is 100, not 203.
- 06Repricing rate. Over the backfilled year, the share of pair-days on which either chain moved its price by at least half a cent. This is the effective sample size for the brand-tier comparison: a price sitting still for nine days is one event, not nine.
- 07Gap episode. A run of consecutive days on which the two prices differ by more than 5% of their mean, in the same direction. It closes when the gap narrows to 5% or less while both prices are still being observed; a run reaching the last observed day is open, and its length is a lower bound rather than a measurement.
Parity rate first. On identical national-brand products the two chains rarely allow themselves to be undercut, and 43% of pairs are priced identically to the cent.
Who wins each matched pair
count of 128 identical products, 23 AugThe parity rate is stable, holding between 39.5% and 43.0% on every August day in the series. Parity is equality, not a tolerance: these are pairs where both chains show the same price to the cent on the same morning. The count runs only over products both chains stock in the same size, which is the population the matching layer is built to identify and the only population on which the word identical means anything.
That headline rate hides the finding. Split the same pairs by aisle and they come apart.
Parity rate by aisle
share of matched pairs priced identically, 23 AugPantry is matched nine times more often than household. The trade has a name for the exposed end: known value items, the lines a shopper can price from memory. Price perception for the whole store is built on them, so they get matched to the cent. Everything else is the long tail, where shoppers hold no reference price against which a difference could register.
Mean absolute gap by aisle
AUD per matched pair, 23 AugThe two charts are the same story twice. Household sits at the bottom of one and the top of the other, with a gap six times pantry’s. This is the price-visible and price-opaque split: a small benchmarked head that sets what customers believe about you, and a long tail that nobody checks. The gap is a mean of line-level differences and is not revenue-weighted, so it describes the shelf rather than the till.
Drinks is the cell that does not fit, and it is worth keeping rather than smoothing. It has the second-lowest parity rate and the smallest gap of any aisle, which says exposure and matching are two axes rather than one: two retailers can track each other closely without ever landing on the same number. A two-bucket model cannot represent that, and the right response is to refine the model rather than defend it.
Promotion frequency and depth. The two chains run visibly different promotional machines.
Averaged across the collected days, Woolworths flags 14.3% of the catalogue on promotion at any one time, against Coles’ 1.5%. The relationship inverts on depth: where a promotion is a genuine cut, the median observed discount is 41.0% at Coles against 28.6% at Woolworths. Coles promotes rarely and deeply, Woolworths often and shallowly, which is a promotional-planning difference rather than a pricing one.
Reference-price integrity, and the question a single snapshot cannot reach.
From a single day it is possible to see that a product is flagged on promotion and what the retailer claims it previously cost. Whether the price actually fell cannot be established, because the preceding day was not observed. Across 319 promotion episodes, 186 began on a day the product had already been priced here.
What the promotion flag did to the price
186 episodes with an observable baselineTwenty-five of 186 promotions moved the price the wrong way or not at all, and all twenty-five were at Coles, whose 20 genuine cuts are outnumbered by Woolworths’ 141. The clearest case: four Schweppes mineral waters at $3.00, flagged on promotion at $3.30.
The obvious next suspicion is inflated reference pricing, a "was" price the retailer never really charged. That one does not hold. Of the 157 episodes carrying both an advertised was-price and an earlier independent observation to test it against, all 157 matched: reference-price integrity of 100%. This is not fake was-prices. It is the promotion flag and the shelf price being managed separately, which is a duller finding and a truer one.
Promotion or position. The single question the snapshot cannot answer.
What happened when a promotion ended
93 episodes with a price observed on both sidesA promotion on these shelves is deep and brief and then it is over: a median 38% off over a median run of 7 days, ending where it started. Nine in ten are a promotional cycle rather than a price change, a distinction between timing a purchase and changing where it is made.
Run the same question at the pair level and it answers differently. Of the pairs observed on at least five days, 117 opened with a gap wider than 5%, and 100 of them, 85%, were still wider than 5% on the last day observed. Over this window gaps mostly do not close, though that proves to be a fact about the window rather than about the gaps, and the three-year series below settles it. The split therefore governs whether the two chains match at all, rather than how quickly they correct once they differ.
Largest same-product gaps
identical product, AUD, 23 AugThe same OMO bottle was $30 at Woolworths and $15 at Coles on the same day. Every line here is laundry, coffee, toothpaste or dish soap, and not one is a known value item. A bare label is one product. A label reading (5 variants) is five products of the same line carrying the same gap to the cent, grouped into one bar rather than repeated five times. They are grouped because a whole range goes on promotion at once, which is the actual unit of decision: the choice times a range rather than a product.
Where the two chains compete, they compete to the cent. Where they do not, a gap can sit for a month.
Three years, and a test I could fail
Thirteen days answers which chain was cheaper. It cannot answer whether a gap is a promotion or a position, because on any single morning those look identical and only the following months separate them. An open price tracker has been scraping both chains daily since September 2023 and publishes the result, and it keys products by the retailers’ own product ids, which are the same ids this collector already stores. So the pairs matched here extend backwards on an equality join, with no re-matching at all.
Borrowed data is worth what it can be checked against. For every day this project priced a product itself, the backfill is asked what price it implies for that day and the two are compared: 22,616 overlapping observations, 99.97% agreeing to the cent, from two scrapers built independently by two people who have never spoken. That check is a gate rather than a report. It runs before anything is built, and below 99% the whole arm refuses to build.
Eligibility and window are two different things, and conflating them cost most of the study before it was caught. A product must have been observed for at least a year to be matchable at all, which is what keeps the pair set honest. The window each matched pair is then measured over is simply all the history it has. Demanding instead that every pair span the same three years left 67 pairs out of 1,523, because the upstream tracker only grew its own catalogue over time. Separated, the same 1,523 pairs carry 775,418 pair-days back to September 2023, against 557,418 under the twelve-month window they replaced.
775,418
Pair-days with a price at both chains, across 1,523 matched pairs, back to September 2023.
99.97%
Agreement between the backfill and this project’s own collected days, over 22,616 observations.
67.6%
Of every gap that closed, the share that lasted exactly seven days.
The first thing the history buys is a correction to the thirteen-day answer.
Over thirteen days, gaps looked permanent: four in five that opened wider than 5% were still open on the last day observed. Over three years they are not. Of 42,674 closed gap episodes, the median lasts seven days. The short window was not measuring how long gaps persist. It was measuring the length of its own window, which is the failure mode of every study reporting persistence over a period shorter than the thing being measured.
Promotion or position
gap between the chains as a share of the mean price, sampled every third day, Jun 2025 to Aug 2026Two real pairs, both accepted by the same matcher, plotted on the same axis. The chocolate block is a square wave: the gap slams open to about 67%, holds for a week, shuts to nothing, and does it again and again. The olive oil opens a 20% gap and simply never closes it. On any single morning both look like the same finding, which is exactly what the thirteen-day study could not tell apart. Only the following months separate a promotion from a position, and the two products here are also a name brand and a store brand, which is the same split the charts below measure in aggregate.
How long a gap lasts before it closes
42,674 closed gap episodes, Sep 2023 to Aug 2026Sixty-eight per cent of every gap that closed lasted exactly seven days, with a second bump at fourteen. This is not a median landing near a week, it is a spike on the week itself, and it holds in both brand tiers and in both pair sets across three years. The promotional week is the unit of Australian grocery pricing, and a study shorter than one cycle cannot see the shape at all: a thirteen-day window catches one visit from this distribution and reports it as a level.
Ask how far apart the two chains usually are and the answer is bimodal. There is almost no middle.
How far apart the two prices are
775,418 observed pair-days, Sep 2023 to Aug 2026Half of all pair-days are priced identically to the cent, and a further third sit more than 20% apart. Between them, the entire band from a fraction of a per cent to five per cent accounts for 2.4% of observations. Prices are not distributed around a small average difference; they are either matched exactly or they are nowhere near each other. That shape is the argument against ever quoting a mean gap on its own, because the mean lands in the valley where almost nothing actually sits, and it is why every measure on this page is a rate or a median rather than an average.
Three years is also long enough to ask whether repricing has a season.
How often prices move, by month
share of pair-days on which either chain repriced, Oct 2023 to Aug 2026The floor of every year is December: 5.4% in 2023, 6.1% in 2024, 6.0% in 2025, against 7 to 11% in most other months. Both chains stop moving prices over Christmas and start again in the new year, three years running. The early months of the series rest on far fewer pairs than the later ones, because the upstream tracker was still growing its catalogue, so the left of this chart is noisier than the right; the December floor is the one feature that survives that unevenness, and it survives it three times.
The second is a split the aisle chart could not separate: store brand against name brand.
How often a price moves
share of pair-days on which either chain repriced, Sep 2023 to Aug 2026A name brand is the same physical good on both shelves, so a gap between the chains is a pricing decision. A store brand is a substitute from a different supplier, so a gap is partly a product difference. Pooled, store brands reprice about three and a half times less often than name brands. Split by aisle that single number turns out to be an average of two regimes: nearly nine times less often on packaged staples, and only 1.6 times on produce. Crop and weather move a price whoever owns the label, and private-label pricing discipline is something you can only exercise over a manufactured good.
Writing the prediction down first is what makes it a test rather than a story.
The aisle split earlier on this page was found in the data rather than predicted before it, which makes it a hypothesis this study generated rather than one it tested. So for the longer series the item list, the bucket assignment and the expected numbers were committed to version control before a single bucket-level figure was computed, and the commit hash is published beside the results. Seven of twelve predictions landed inside the registered range. Packaged staples returned five of six and fresh produce two of six, which is the cell the registration itself had flagged as least certain.
The claim that mattered was whether the store-brand effect survived holding aisle constant, with a commitment made in advance to retract it publicly if it did not. It survived. The prediction that failed hardest is the more useful result.
Do the two chains’ prices move together?
median per-pair correlation of monthly mean price, Sep 2023 to Aug 2026Those are monthly means over three years. Day to day the two name-brand figures collapse to zero and faintly negative, against a registered prediction of 0.80 to 0.95. Two chains buying the same crop in the same weather ought to move together, and on national brands they do not, because national brands are the promotional vehicles and the two chains run their cycles out of phase. The prices take turns rather than moving together. Store brands are lightly promoted, so what remains is shared cost arriving at both chains at once, and they track each other most closely of all. The axis separating correlated from uncorrelated prices is not fresh against packaged. It is promoted against not.
A failed prediction of that kind is more informative than a successful one. It was wrong in a direction that named its own cause, and that cause is visible in the other measures on this page rather than invented to account for the miss.
The longer view
For the longer horizon, Savings.com.au has run a monthly Coles-versus-Woolworths index since late 2023, pricing a fixed basket at both chains and publishing the annual movement. It is not my data, not my basket and the methodology is theirs, so it sits here as context and nothing above is derived from it.
Australian grocery inflation, both chains averaged
annual change in basket cost, %. Source: Savings.com.au Grocery Price IndexGrocery inflation peaked above 9% in February 2025 and had cooled to 1.38% by July 2026. That frames everything above. The gaps measured here open and close inside a market whose overall level is close to flat, so they are competitive behaviour between two chains rather than a cost-of-living signal, and a gap that swings 67% on a chocolate block says nothing about what groceries cost in aggregate.
Their July 2026 basket came to $262.02 at Coles and $261.27 at Woolworths, 75 cents apart on roughly $260. A different basket and a different method, reaching the conclusion this study reaches from the other direction: at the level of the whole shop there is nothing to choose between the two chains, and the interesting variation is underneath that, in which lines each is willing to compete on. Figures read from their index on 23 August 2026, covering November 2023 to July 2026.
Conclusion
Five things the series supports, in the order I would defend them.
- 01The average hides the finding. Parity runs from 62.5% in pantry to 7.1% in household on the same morning, and the mean household gap is six times pantry’s. The two chains compete hard where a shopper can price from memory and barely at all where they cannot.
- 02A promotion flag is not a price cut. Of 121 promotion starts with an observable baseline, 17 moved the price the wrong way or not at all, including four mineral waters going from $3.00 to $3.30 while flagged. All 17 were at one chain. The advertised was-prices themselves check out, 100 of 100, so this is the badge and the shelf price being managed apart rather than fake discounting.
- 03Nine in ten promotions are a cycle, not a price change. Of 44 promotions observed on both sides, 39 returned to exactly the price they started at, median 41% off over a 7-day run. That is the difference between timing your shop and switching your shop.
- 04Store brands are priced to hold still and name brands are the promotional vehicle. Across three years a store-brand packaged staple reprices nearly nine times less often than its name-brand equivalent, and the two chains’ name-brand prices barely correlate day to day because their promotional cycles run out of phase.
What the study cannot establish is worth stating as plainly. These are shelf prices, not transacted prices: member pricing, multi-buys and loyalty offers sit underneath them. There are no quantities anywhere in public price data, so nothing here is elasticity and the word does not appear. It is online national pricing from one collection point, and the promotion and reference-price measures rest on thirteen collected days rather than on the backfilled history, because the backfill carries prices and no promotion metadata at all.
This is a study of retailer behaviour. It measures how two chains price against each other in public: where they match, by how much they differ, how often and how deeply they promote, and whether a difference closes or holds. Behaviour is the useful signal here, because what each retailer chooses to match is a direct read on which lines it believes are competitive.
The answer it produces is a shape rather than a verdict. There is a small benchmarked head where the two chains track each other to the cent and earn almost nothing, and a long tail where a gap can sit for weeks because nobody is checking. Knowing which line is which, and being able to prove it rather than assert it, is the work. The same shape holds anywhere prices are set against a competitor rather than fixed by a contract, which is what makes the method portable off the supermarket shelf.
The measure I would add next is follow latency: the number of days before one retailer answers the other’s move. It reads competitive intensity more directly than parity does, because parity tells you where two chains have landed and latency tells you how hard they are watching. On this split it should separate the head from the tail sharply, and it is the natural next thing this warehouse is already shaped to compute.