Home/Concepts/Abduction and inference to the best explanation in retail operations
Abduction and inference to the best explanation in retail operations
If the best explanation is only best relative to the alternatives in play, then any system whose alternatives were fixed at a moment in the past is committed to explanations that…
The category manager's problem
A category manager plans an assortment against a demand curve. The curve is a hypothesis: it says which explanation of past sales best accounts for what happened, and by extension what will happen next. The forecast is not a fact about the future. It is the winner of a ranking exercise conducted over whatever causes were on the table when the plan was built. The trouble in retail is not that the ranking is done badly. It is often done well, on the evidence supplied. The trouble is what was never supplied.
Take a mid-size grocery or apparel operation running a quarterly assortment cycle. The planning cadence is fixed: point-of-sale history feeds a demand model, the model ranks candidate explanations for observed sell-through — seasonality, promotion lift, a competitor's pricing, weather — and the best-fitting one sets next quarter's buy. That ranking is a textbook case of inference to the best explanation, in Gilbert Harman's sense: accept the hypothesis that best explains the sales pattern, judged by fit to the data and coherence with the planner's other beliefs about the category. The category manager is doing exactly what Charles Sanders Peirce described as the only inference that generates a new idea rather than merely testing one already held.
The failure mode this page is built around is precise: an assortment is planned against a demand curve that already moved. The plan is not wrong about the evidence it saw. It is silent about the evidence it never received, because the candidate that would have explained the shift was never in the slate being ranked.
What "best" quietly assumes
The word doing the work in inference to the best explanation is "best", and best is a comparative, not an absolute. A hypothesis wins only against the field actually entered. A demand model trained on twelve months of point-of-sale data will rank "holiday timing shifted" against "promotion cannibalised full-price sales" against "last year's cold snap inflated the baseline." It will not rank "a regional competitor opened forty stores last month" if that fact never entered the corpus the model was trained on. The absent hypothesis does not lose the ranking. It is not in the race. No confidence interval widens to announce its absence, because there is nothing in the system's evidence to be uncertain about. This is the structural point: an unconsidered cause produces no warning, no flag, no low score. It produces silence, and silence looks exactly like confidence.
This is where the intake axis becomes the whole story for a category manager, not an abstraction about model architecture. A system trained on a frozen slice of point-of-sale history — the retail analogue of a Large Language Model's fixed corpus — abduces beautifully over last year's causes and cannot notice that this year introduced a new one. A system that also ingests a live store's sensor feed — shelf cameras, footfall counters, a single site's real-time telemetry, the retail analogue of a Large World Model's bounded scene — widens the candidate slate while that scene is live. A stockout event registers; a queue-length anomaly registers; a competitor promotion visible from the car park might even register if someone pointed a camera at it. But the scene ends at the site boundary and at the reporting window's close. The widened slate does not travel to next quarter's plan, and it does not travel to the next region.
What retail operations actually generates, continuously, whether or not any planning system uses it that way, is a genuine analogue of the third position: point-of-sale streams updating by the minute, inventory telemetry from warehouse and shelf, supplier notices about input costs and shipment delays, and demand signals from search, returns and customer service contacts. None of these close. If a category manager's system held that as a persistent, tagged, revisable object — this candidate raised by a supplier notice dated Tuesday, that one raised by a footfall sensor an hour ago, this other one inherited from last year's seasonal file and now due for review — then hypothesis generation would not conclude at the planning meeting. It would continue past it. That is the Large Universe Model position applied to a category plan: not a better forecast, but a candidate set that stays open because the streams feeding it stay open.
The Bayesian's rebuttal, and its limit
Put a prior on a catch-all "unmodelled cause" term. When the named hypotheses fit the sales data badly, the posterior mass on the catch-all rises. That is a formal, well-understood fix. No new intake required, only disciplined probability hygiene.
This is a serious answer and category managers who work with demand-planning statisticians hear versions of it often. A residual term genuinely earns its keep: if promotion lift, seasonality and weather jointly under-explain a 30% deviation in sell-through, a well-specified catch-all mass rising to, say, 0.4 is a real signal that something structural changed. It is cheap, it is principled, and it requires no new data source at all.
But the catch-all cannot name the cause. It tells the category manager the model is wrong; it does not tell her whether the competitor opened stores, whether a supplier substituted an ingredient that changed shelf life, or whether a demographic shift in the trade area moved the customer base entirely. Ordering stock, cancelling a purchase order, renegotiating a supply contract — these require a named hypothesis, not an elevated residual. Converting "something I didn't model" into "the new discount grocer eleven minutes' drive from store 214" is itself an act of hypothesis generation, and generation needs material the model has not yet seen: a lease filing, a footfall drop correlated with a specific competitor's opening date, a supplier notice about a substituted input. The catch-all is the alarm. Live intake is what answers it.
The proliferation objection, and its limit
Widening intake multiplies candidates faster than it multiplies truth. Infinitely many hypotheses explain any sales pattern; a curated twelve-month corpus is a feature, encoding which explanations a competent planner already knows are worth considering.
This has force in retail specifically, where noisy signals are cheap and plentiful: every price change anywhere, every weather station reading, every social-media mention could in principle be fed into a demand model, and most of it would be irrelevant. A curated corpus does encode real domain judgement about which causes matter.
The answer turns on provenance, not volume. A candidate cause arriving from a supplier notice carries a date, a supplier identifier and a specific claim — input cost rising 8%, effective a named week. It can be checked against invoices and retired the moment it stops matching reality. A candidate cause silently baked into an undated training corpus, by contrast, cannot be checked or retired at all; it simply persists as an unexamined assumption inside the ranking. Wide intake with provenance on every candidate is a large space that can be audited and pruned. A narrow corpus without provenance is a small space that cannot be, because nobody can tell which of its embedded assumptions have expired. The genuine cost of proliferation is real. It is a ranking cost, and ranking is tractable in a way that silent absence is not.
What does not resolve
None of this makes the forecast true. Bas van Fraassen's argument from the bad lot applies to a category plan as much as to any scientific inference: the best-ranked explanation of a sales anomaly might still be a bad explanation, because the true cause was never in the lot, however wide the lot became. Continuous intake changes how candidates enter — from a modeller's habitual list of "things that usually move sell-through" to whatever the streams actually surface — but it does not guarantee the true cause surfaces. A competitor's plan announced only in a private board meeting will never appear in any stream a retailer can legally access. Open intake narrows one specific failure: the silent one, where an absent hypothesis leaves no trace. It leaves every other difficulty of inference to the best explanation intact — bad lots, underdetermined rankings, the trade-off between a simple story and a true one.
The category manager's real choice is not between a good model and a bad one. It is between a system whose candidate set closes at the planning meeting and one whose candidate set stays open because the point-of-sale stream, the inventory telemetry, the supplier notice and the footfall sensor never stop arriving. That openness does not produce a correct forecast. It produces a forecast that can be caught moving before the shelves are.