The Apparel OSby RetailNorthstar

The order you meet the number in

How a number reaches a merchandising team is an input to the decision it produces, not packaging around it. The order it is met in, the comparison base it is shown against, the level it is aggregated at, the precision it is carried to and the label on its source all change what a merchant is willing to conclude from identical evidence — and therefore what gets bought. That makes presentation an operating decision with a margin consequence rather than a reporting preference. It is also a bounded claim, and this essay spends as much effort marking the boundary as making the argument.

RetailNorthstar Editorial15 min read

Two people read the same reforecast and place different buys

Two merchandising directors at two comparable brands open the same kind of file on the same Tuesday. Same classification, same eight weeks of selling, same sell-through, same open-to-buy position, same vendor, same lead time on a chase. One places the chase that afternoon. The other holds and revisits the following week. Neither is careless, both know the category, and in a room together they would agree on essentially every fact in front of them.

The difference sits upstream of judgment and upstream of data quality. One of them had the vendor’s chase quote and a capacity window in hand before opening the sell-through file. The other read sell-through first and took the vendor call afterwards. Two pieces of information, met in different order. Nobody decided that order — it was set by whoever scheduled the vendor call, and that person was solving a calendar problem rather than a margin problem.

Which is the argument, and it is worth stating in its strong form before qualifying it. The same plan, met in a different order and framed with different confidence, produces a different buy. Not a different write-up of the same buy — a different buy, with different units against it and a different markdown exposure at the end of the season. If that holds, the presentation layer of a planning system is not the part you leave until last. It is part of the plan.

One constraint, stated at the front: this is a claim about ambiguous evidence, not about all evidence. Where the evidence is unambiguous, presentation does very little. It matters anyway, because almost every merchandising decision worth a meeting sits on the ambiguous side of that line — which describes the job rather than criticising the people doing it.

The question has two unrelated answers

This piece started from a maxim that circulates in business writing in roughly the form most half-remembered maxims do: focus on what makes your beer taste better, and stop spending on everything that never reaches the glass. It is a useful sentence. It also turns out to have two entirely separate answers, and choosing one while pretending it is canonical would lose any reader who knows the other.

The first answer is Amazon’s, and it is an argument about where a company spends its effort. The idea has been associated with Jeff Bezos and AWS for years, and the honest state of its provenance is that the idea is documented while the sentence is not. The citation that usually travels with the line, the one that comes with a venue and a year attached, does not survive checking: the retellings converge on one another rather than on any primary source, and they do not agree among themselves on the year or on the wording. So this piece attributes the idea and not the sentence, and prints no venue or date for the line itself. The formal name for it is undifferentiated heavy lifting. Amazon’s own earlier word was blunter: muck. On the AWS blog in September 2006, in a post by Jeff Barr reporting Bezos’s MIT keynote, the muck was named explicitly — server hosting, bandwidth management, contract negotiation, scaling, and the accumulated complexity of mismatched hardware. That post contains no beer and no brewery at all.

Bezos has told the brewery story himself, much later. At the New York Times DealBook Summit in 2024 he described visiting a 300-year-old brewery in Luxembourg whose museum held a hundred-year-old electric generator — built because there was no power grid to plug into. He called the trip one of the small catalysts for founding AWS: every company was running its own data centre, and that was not going to last. In that telling he makes the utility point and only the utility point.

The second answer has nothing to do with any of that: a separate body of published work, in a different discipline, about what happens to the experience of a drink when you are told something about it beforehand. The two share a noun and nothing else, so it is worth being blunt — the maxim has no behavioural-science backing, and the behavioural-science work has nothing to say about build-versus-buy. This essay uses the second answer, and disposes of the first next.

A good question resting on a weak history

Grant the first answer its real value. As a diagnostic it is sharper than the budget exercise it gets confused with: it does not ask what is expensive, it asks what would change for the customer if this disappeared overnight. A merchandising organisation that ran that test across its own calendar would find work it has done for years because it has always done it — reconciliations between files nobody reads, decks assembled for meetings whose decisions were made elsewhere.

Then three disclosures, all worth more to a sceptical operator than the parable is. The first: it is a parable, and should be used as one. Nobody has produced the cohort of breweries that lost to competitors because they kept generating their own power, and nobody has debunked it either — the claim is unevidenced in both directions. A thinking tool is not weakened by being labelled as one. It is weakened by being presented as history.

The second is about who was telling it and to whom. By early 2009 the image was doing sales work: contemporaneous trade-press coverage — Roger Smith writing in InformationWeek on 21 January 2009 — describes an AWS presentation slide showing Bezos standing in front of a vintage 1890s electric generator in a European beer factory. The sources disagree about which country, and that disagreement is itself instructive. The parable was told by the company selling the utility, to people deciding whether to buy the utility. That does not make it wrong. It does mean the honest use of it is to aim the same question at the pitch as at your own operation.

The third is that Amazon itself falsified the simple reading. Its own muck became AWS — the thing it had judged undifferentiated became a business in its own right. So the maxim is a question to re-ask every year rather than a rule to answer once, and it is a question about where effort goes. The rest of this piece is about something else: what happens to a decision when the evidence behind it is genuinely ambiguous, and someone has to present it anyway.

The other beer question is about timing

Start with the clean, independent study rather than the famous one. A team at ETH Zurich — Michael Siegrist and Marie-Eve Cousin, publishing Expectations influence sensory experience in a wine tasting in Appetite 52(3) in 2009 — gave tasters negative information about a wine either before or after they drank it. Delivered before, it lowered ratings. The identical information delivered after did not. Their own conclusion is the load-bearing part: the information affected the experience itself, not merely participants’ overall assessment of it.

The beer version is the better-known one and works here as colour rather than as the foundation. A 2006 study in Psychological Science, run in two Cambridge, Massachusetts pubs by Lee, Frederick and Ariely (DOI 10.1111/j.1467-9280.2006.01829.x), found that patrons rated a beer dosed with balsamic vinegar less favourably when they were told about the additive before tasting than when they were told afterwards. Disclosure only depressed preference when it came first.

The honest limit is worth printing at the same size as the finding. This is a two-study result, not a settled one. Siegrist and Cousin is a conceptual replication in a different lab with a different product, not a direct replication of the pub study, and no direct replication of the pub study exists. Two independent findings pointing the same way is a reasonable basis for an argument and a poor basis for a law.

The operator translation is short and uncomfortable. Information that arrives before the read changes the read; information that arrives after changes the write-up. A merchant who takes the vendor’s capacity warning before opening the file is not reading the same file as a merchant who takes it afterwards, even though the cells are identical. And the difference is invisible from the inside. Nobody experiences a framed read as framed. They experience it as the read.

This only works where the evidence is ambiguous

The most important citation in this essay is the one that limits it. Stephen Hoch and Young-Won Ha, Consumer Learning: Advertising and the Ambiguity of Product Experience, Journal of Consumer Research 13(2), 1986: where the evidence of quality is unambiguous, judgments depend only on the objective physical evidence and are unaffected by advertising. Where the evidence is ambiguous, advertising has dramatic effects — and it pushes people toward confirmatory hypothesis-testing and search. Their own published example of an ambiguous product experience is the quality of a polo shirt.

That boundary converts into a sorting rule a merchandising team can apply on Monday morning. Some of what a planning system shows is unambiguous: units received, on-order position, landed cost on an order already placed, weeks of supply against a known exit date, a vendor cutoff. Facts, and arithmetic on facts. Presentation does not move any of these, and nothing here claims it does.

The rest of the screen is ambiguous, by construction rather than by neglect. Whether six weeks of sell-through on a fashion classification is a trend or a weather artifact is a judgment about a signal that has not finished arriving. Whether a size break is a genuine curve problem or a distribution artifact depends on which doors got which units and when. Whether a soft category deserves one more full-price week is a bet on a demand curve nobody can observe. These are not lower-quality facts; they are the decisions the job consists of, every one resolved by inference from incomplete evidence.

The confirmatory-search finding is the sharp end, and it is what makes framing expensive rather than merely interesting. The merchandising version of it is inference rather than a measured result, and worth reading as such: a framed read does not only change the conclusion a merchant reaches; it changes what they go and look at next. Primed to see a category as recovering, a merchant opens the doors that are recovering; primed to see the same category as in trouble, they open the doors that are in trouble. Both find what they went looking for, both are looking at real data, and both leave the screen more confident than they arrived.

Before anyone oversells this

There is a version of this argument that collapses into perception being reality, and the disconfirming evidence against it is strong and specific. Tina Kurz, Emir Efendic and Caroline Goukens, Pricey therefore good? in Psychology & Marketing 40(6), 2023: across six studies with 2,842 participants, a higher price raised expected quality every time — and failed to consistently raise perceived quality or liking once the product was actually in people’s hands. Anyone selling you perception as reality is quoting the first half of that sentence.

The lever also has an asymmetric shape rather than a general one. In a 2021 framed field experiment with 140 tasters, published in Food Quality and Preference, a deceptively inflated price raised enjoyment of a cheap wine, while a truthful price and no price at all moved enjoyment not at all. The nuance underneath is the useful part: intensity-of-taste ratings tracked actual retail price rather than the stated price. The framing moved liking. It did not move discrimination.

Read across, the operator claim is precise. Framing changes what a merchant is willing to conclude from thin evidence. It does not change what the evidence is, and it does not survive contact with a fact. A confidently presented projection can persuade a team to hold a week longer. It cannot make the units sell.

One more finding, because it turns the argument back on the reader. In a sample of more than 6,000 blind tastings reported in the Journal of Wine Economics in 2008, the correlation between price and overall rating was small and negative — untrained drinkers enjoyed the more expensive wines slightly less. Among tasters with training, the relationship was not negative. Expertise changes the exposure. The reader of this essay is well trained on apparel and untrained on the presentation layer of their own planning system, because nobody has ever asked them to be. The screen is the one part of the job that arrived pre-decided.

A guard rail to close, since this is an essay about numbers that persuade. The most-quoted number in change management — that 70 per cent of initiatives fail — was traced back through five published instances by Mark Hughes in the Journal of Change Management in 2011 and found to rest on no valid or reliable evidence in any of them. This essay carries no number of that kind, and where it cites one it names the study, the venue and the year so the claim can be checked rather than absorbed.

Where this lands first: the order the week is read in

The weekly read has an order, and in most merchandising teams that order was never chosen. It is an artifact of the meeting calendar, the file structure and whichever report opens on the first tab. Whether the vendor’s chase quote arrives before or after the sell-through read; whether the markdown scenario grid is built before or after the weeks-of-supply position is stated out loud; whether the buy review opens on last year’s actuals or this year’s plan — each is a sequencing decision with the same evidence on both sides and a different decision at the end.

The comparison base is the same problem in a smaller frame, and it is probably the highest-leverage single default in the system. Identical actuals shown against the pre-season plan read as a miss. The same actuals shown against last week’s reforecast read as a trend. One invites a defence, the other invites a projection. Most teams have never decided which is the default for the Monday read. The file decided, and it decided years ago.

Aggregation level does the same work again. The level at which a number is first met determines what gets investigated, because nobody drills into a green cell. A category that reads fine at the top can be carrying a class that does not, and the class is found only by somebody who had a reason to open it. The default rollup on the first screen is not a display setting. It is an attention policy, applied to the whole assortment, every week, by whoever configured the view.

The quietest of these is the variance threshold. The percentage at which a cell turns from neutral to flagged was set once, quite possibly as a template default that shipped with the file, and it now allocates merchant attention across the range for the rest of the season. None of this is a bias to be trained out of anybody. It is a set of defaults chosen by accident and worth choosing on purpose — which costs a meeting, not a project.

Precision, confidence, and a plan that looks more certain than it is

A projected sell-through carried to one decimal place reads as a measurement. The same projection expressed as a range reads as an estimate. Same model, same arithmetic, same underlying uncertainty — and a materially different willingness to override it. False precision is not a cosmetic flaw in a planning output. It is a confidence claim the model never made, added by the format.

Evidence weight goes missing the same way. A chase recommendation built on three weeks of selling and one built on eight weeks are, in most systems, rendered in identical typography at identical precision on identical rows. The merchant is asked to weigh evidence whose weight was removed from view before it reached them. They can go and find it — the weeks are in the data somewhere — but going to find it is exactly the effort a well-presented number persuades you not to spend.

Provenance is the third variable and the least examined. A number a planner built themselves gets defended; the same number arriving from a system gets deferred to, or dismissed, and only rarely interrogated. Three behaviours attached to one figure, and which one you get depends on whether the merchant knows where it came from. Most screens do not say.

The size curve is where all three converge. An identical distribution presented as a template default and presented as a derived recommendation with its history attached produces different override behaviour — not because the numbers differ but because one invites inspection and the other invites acceptance. How the curve should be constructed is a separate argument belonging to the margin is made before the buy, which establishes that the curve is free to change until the purchase order is issued and frozen after it. The claim here is narrower and sits on top of that one: presentation determines whether anybody interrogates the curve inside the window where interrogating it is still free.

The prescription is craft rather than research, and worth flagging as such. Confidence belongs on the face of the number: how much data sits behind it, how wide the plausible range is, who or what produced it. A plan that hides its own uncertainty is not neutral. It is framed, in the direction of not being questioned. For the worked numeric version of a read done properly, retail-plan.com covers it in how to read a WSSI.

The label on the recommendation, including ours

The effect reproduces inside software, and the study that shows it implicates this company directly. Kosch, Welsch, Chuang and Schmidt, The Placebo Effect of Artificial Intelligence in Human-Computer Interaction, published in 2023 in ACM Transactions on Computer-Human Interaction: across two experiments with 369 and 100 participants, people told they were being supported by an adaptive AI held higher expectations of their own performance, and those expectations tracked how many puzzles they actually solved. No AI was running in either condition. The label was doing the work.

RetailNorthstar sells AI-assisted planning. So the consequence should be printed without softening: the AI-assisted label on a recommendation is itself a framing device, and it changes how that recommendation is received before the model has contributed anything at all. That is true of our product and of every other product in the category that carries the same words on the same screen.

Inside a merchandising team it cuts both ways, which makes it an operating problem rather than a marketing one. A style flagged by the system draws attention that an identical unflagged style carrying identical evidence does not — so the flag has allocated the merchant’s time before anyone evaluated whether the flag was right. And a merchant who has decided the model is unreliable discounts a good recommendation for the same reason, on the same evidence. Neither reaction is irrational. Both are responses to a label rather than to a finding.

Tie that back to the boundary. The label matters most precisely where the evidence is thinnest — and thin evidence is where recommendations get generated in the first place, because a classification with eight clean weeks of unambiguous selling does not need one. The framing effect is strongest exactly where the underlying signal is weakest. That is not an argument against AI-assisted planning. It is an argument that a system producing recommendations owes the merchant the evidence behind each one, at the same moment and in the same place.

The honest position for any vendor is that a system cannot be neutral about how it presents a number. There is no unframed way to show a projection. The only real choice is whether the framing was designed on purpose or inherited from whatever the template did.

Framing as an operating protocol, not a design preference

None of what follows requires a purchase, and that is deliberate. Four items, all adoptable by a team that changes nothing about its systems this quarter.

One: decide the order of encounter for the weekly read, and write it down. Sell-through before the vendor conversation. Inventory position before the scenario grid. The comparison base named explicitly — against plan, against last week’s reforecast, against last year — rather than defaulted by whichever tab opens first. Two: put confidence on the face of every projection, meaning the weeks of data behind it and a range instead of a decimal place. Three: make provenance visible, so a merchant knows whether they are overriding a template default, a derived recommendation or another human’s judgment — three very different acts that currently look identical on screen.

Four, and this is the one that pays for the meeting: audit the defaults nobody chose. The variance threshold. The aggregation level of the first screen. The order options appear in a markdown scenario grid, and whether holding and doing nothing appears there as a scenario at all — because the set of options presented is the set of options considered, and an option missing from the grid is not rejected, it is invisible. Most were set once, by someone who has probably left, in a file copied forward through four seasons.

One narrow product claim, then stop. When plan, buy and actuals sit on a single data model, every function meets the same number in the same context — which removes framing variance between artifacts rather than removing framing itself. Framing does not go away. It becomes a thing somebody chose.

The number was never met cold

Go back to the two merchants. One chased, one held, and by the end of the season one of them will have been right — and the post-mortem will treat that as a difference in judgment, because the post-mortem will look at the decision and not at the order the inputs arrived in. The order is the part nobody records, and it is the part that is free to change.

The uncomfortable version of the beer question for a merchandising organisation is not whether the data is good. It is which of the things the team spends on actually reaches the decision — and which of them the brand is, in effect, generating its own electricity for. Answer it in both directions and the presentation layer comes out on the wrong side of the usual assumption. It is not decoration around the plan. It is part of the plan, it changes the buy, and it is the one input that costs nothing to change and is almost never audited.

The argument holds only where the evidence is ambiguous. That happens to be everywhere the job is hard.

Frequently asked questions

What does it mean to call presentation an operating decision rather than a reporting preference?
It means the way a number reaches a merchandising team is an input to the decision that comes out of it, in the same sense that the number itself is an input. The order the number is met in, the comparison base it is shown against, the level it is aggregated at, the precision it is carried to and the label on its source all change what a merchant is willing to conclude from the same evidence. Because those choices change the buy, they carry the same margin consequence as the buy, and they deserve to be decided rather than inherited from a template.
Is this just saying that merchants are biased?
No, and the distinction matters. The published work on expectation does not show that people are careless; it shows that information arriving before an experience changes the experience itself rather than only the verdict afterwards. Applied to planning, that means a framed read is not a lapse in rigour a merchant could train out by concentrating harder. It is a property of when information arrives. The lever is the arrival order and the presentation, not the person.
Where does framing stop working?
Where the evidence is unambiguous. Hoch and Ha showed in the Journal of Consumer Research in 1986 that when the evidence of quality is unambiguous, judgments depend on the physical evidence alone and are unaffected by advertising; framing takes over only when the evidence is ambiguous. In merchandising that draws a clean line. Units received, on-order position, a landed cost and a vendor cutoff date are not framing-sensitive. Whether six weeks of sell-through on a fashion classification is a trend or a weather artifact is.
Does the AI-assisted label on a recommendation change how a planner receives it?
The evidence says a label of that kind changes expectations on its own. In two experiments published in ACM Transactions on Computer-Human Interaction in 2023, participants told an adaptive AI was supporting them held higher expectations of their own performance, and those expectations tracked how many puzzles they actually solved — with no AI running in either condition. RetailNorthstar sells AI-assisted planning, so this applies to us directly: the label moves reception before the model has contributed anything, in both directions. A flagged style draws attention an identical unflagged one does not, and a merchant who has decided the model is unreliable discounts a good recommendation for the same reason.
What is the smallest change a planning team can make on Monday?
Write down the order of the weekly read and stop letting the calendar set it. Sell-through before the vendor conversation, inventory position before the scenario grid, and the comparison base named out loud rather than defaulted by whichever file opens first. It costs nothing, needs no new system, and it removes the single largest source of variance between two merchants looking at the same classification in the same week.
Does putting plan, buy and actuals on one data model remove framing?
It removes framing variance between artifacts, which is a narrower and more honest claim. When every function meets the same number in the same context, two people stop arguing from two differently framed versions of one fact. Framing itself does not disappear — a shared record still has a default aggregation level, a default comparison base and a default precision. What changes is that those defaults become a choice somebody made rather than an accident that survived from a template.

For the adjacent argument about which decisions fix the outcome and when, read the margin is made before the buy, see how the stages actually connect in the connected apparel workflow map, or work a real read line by line with how to read a WSSI on retail-plan.com.

Book a Demo →