That gap is worth sitting with, because it is the normal outcome rather than an unlucky one. Accuracy is what forecasting teams are measured on. It is not what determines whether a forecast changes a decision.
The Override Is Where the Value Is Decided
In most planning organizations a statistical baseline is produced, and then people adjust it. That adjustment step is not a pathology — in a properly run sales and operations planning process it is the designed path from a statistical forecast to a consensus one. The question is whether the adjustment is informed.
The best-known study on this is worth knowing before forming a view. Fildes, Goodwin, Lawrence and Nikolopoulos examined more than 60,000 forecasts across four supply-chain companies and found that in three of them, judgemental adjustments increased accuracy on average. So the reflex that human intervention is simply noise to be engineered out does not survive the evidence.
What the same study found next is the useful part. Larger adjustments tended to improve accuracy, while smaller adjustments frequently damaged it. And upward adjustments were markedly less likely to help than downward ones, which points at an optimism bias that anyone who has sat in a demand review will recognise.
Read that together and a sharper picture emerges. The problem is not that planners adjust. It is the small, low-conviction, mildly optimistic adjustment — the nudge made because the number felt uncomfortable rather than because the planner knew something the model did not. That is the category worth attacking, and it is precisely the category an unexplained forecast produces.
Measuring It Is a Solved Problem That Few Organizations Have Adopted
There is an established technique for this and it deserves naming rather than reinventing. Forecast Value Added analysis measures the change in forecast performance attributable to each step and participant in the process — including the override step — by keeping the pre-adjustment baseline as its own series and comparing both against outcomes and against a naïve model.
Published FVA work is not flattering to the override step. Multiple documented analyses have found manual adjustments failing to beat a simple baseline in a substantial share of cases, and in some reported instances making the forecast worse more often than better.
So this is not an unexamined question. It is an under-adopted method, and the reason is more mundane than institutional cowardice: FVA requires retaining the statistical baseline as a separate, preserved series alongside the adjusted one, and a lot of planning systems and processes simply overwrite it. That is a data-plumbing problem, and it is fixable.
If you take one thing from this piece, it is that the number to establish is not your forecast error. It is whether your adjustments are adding value or subtracting it, and you cannot know that unless the baseline survives.
Why Planners Adjust Anyway

There is a well-replicated finding here too, and it should be named rather than dismissed. Dietvorst, Simmons and Massey documented algorithm aversion: people abandon an algorithm after seeing it err, even when they have been told it outperforms them. That is not a rational defensibility calculation. It is a reaction to visible machine error that human error of the same magnitude does not provoke.
Alongside it sits a second, more practical mechanism. A planner is usually accountable for the outcome, and “the model said so” is a weak position in a review meeting, whereas reasoning a person can articulate is defensible even when wrong. How much work that does relative to algorithm aversion is contested — at least one study testing blame-avoidance against aversion found aversion doing most of the explaining — and it also weakens in organizations where the planner is measured on forecast error directly, because then they own the number either way.
The more useful finding is what actually fixes it. The same researchers later showed that people will use imperfect algorithms if they are allowed to modify them, even slightly. Retained agency, not demonstrated accuracy, moved adoption.
That is the whole argument in one result. The route to a forecast being used is not proving it is right. It is giving the person who has to act on it something to hold.
What Explanation Actually Changes
If adjustment is the designed step, and small uninformed adjustments are the damaging ones, and retained agency is what drives adoption, then the design goal is not fewer overrides. It is better-informed ones.
- Disagreement moves to the driver. A planner who can see that the model is weighting last year’s promotion heavily can dispute that specific input, rather than substituting a number. A dispute about an input has a resolution; a substituted number does not.
- The adjustment becomes a signal. An override made because there is a competitor closure the model cannot see is not a disagreement — it is a missing input, and it belongs on a backlog rather than in a spreadsheet.
- The review answers a different question. “Is this number right?” cannot be settled in advance. “Is this reasoning right?” can be settled in the room, by people who know the market.
- Bad inputs surface. Explanations expose the model’s beliefs, and some will be artefacts — a store that closed and was never recorded, a calendar with a holiday in the wrong week. An accuracy score cannot tell you this; it can only tell you that something is wrong somewhere.
- Assumptions become ownable. Six named drivers can have six owners. One opaque number can only be owned by the team that produced it.
The Objection That Actually Lands

Most critiques of explainability are weak. This one is not, and it should change how any of this gets implemented.
Bansal and colleagues found that explanations increased the likelihood that people accepted an AI recommendation regardless of whether it was correct. Performance improved slightly when the recommendation was right and degraded when it was wrong. Explanation did not make people better discriminators. It made them more compliant.
That is a serious problem for the position in this piece, and the honest response is not to wave at it. It means explanation has to be designed against over-acceptance rather than assumed to produce scrutiny — which in practice means leading with what the model does not know and what would falsify it, not with a tidy attribution of what it does know. An explanation that only justifies is a persuasion tool. An explanation that also exposes its own limits is a decision tool.
Two smaller limits, stated plainly. For fully automated, high-volume, low-value decisions nobody reviews, none of this applies — there is no human to inform, and accuracy is the whole game. And explanation makes adjustment arguable without making it rarer: give a planner six drivers and one who wants to adjust can always find one to dispute. A falling override rate is the wrong success measure.
What a Useful Explanation Contains
It is worth being precise here, because a common shortcut does not clear the bar. Global feature-importance rankings — price matters more than seasonality across the model as a whole — tell a planner nothing actionable about this week’s number. Local, per-forecast attributions expressed in the units of the forecast are a different matter and do real work: plus 180 units attributable to the promotional calendar, minus 60 to the price move is something a planner can argue with.
Attribution gets you part of the way. The list below is what a planner can actually use, and the last two items are the ones attribution methods do not supply.
- What changed since the last cycle, and why. Planners do not re-derive a forecast each period; they want the delta explained.
- Which drivers moved, in units and names a planner recognises rather than transformed feature labels.
- What the model does not know. An explicit statement of what sits outside the inputs, which tells a planner where their knowledge is additive rather than redundant.
- The shape of the uncertainty. A point invites acceptance or rejection. A range with a stated distribution invites a decision about risk, which is the decision actually being taken.
- What would have to be true for this to be wrong. The item most often missing and the most valuable, because it turns a forecast into a falsifiable claim someone can monitor — and because it is the direct antidote to the over-acceptance problem above.
The Reporting Change
Keep the accuracy metrics. A well-explained bad forecast is not an improvement. But report them alongside:
- Override rate, by planner and by category, so the pattern is visible rather than anecdotal
- Forecast value added at the adjustment step, which requires preserving the statistical baseline and is the single highest-value plumbing change available to most planning functions
- Adjustment size distribution, because the literature suggests the small ones are where the damage concentrates and the average hides them
- Time to agreement in the planning cycle, which is where explanation tends to pay back first and where nobody thinks to look
Why This Matters Now
Two things are converging. Explanation tooling has become substantially cheaper to deploy. And there is a growing argument in the forecasting literature that accuracy metrics are weak proxies for the decisions they feed — that forecasting research has tended to treat accuracy as an end in itself while inventory and planning decisions depend on rather different properties. If that gap is real, a business can improve its accuracy number while nothing downstream changes, which is exactly the scene this piece opened with.
That does not mean accuracy is finished as a pursuit. Where explanations reveal a wrongly recorded store closure or a mis-dated holiday, the error was never irreducible — it was a data defect, and fixing it improves accuracy and explanation together. The two are less in competition than the framing of this piece implies.
What has changed is where the cheap improvement sits. For a lot of planning functions the next real gain is not a better model. It is whether the output survives contact with the people who have to act on it — which is a question about how planning meetings run, what gets preserved in the data, and who owns which assumption.
DAX Software Solutions is a Microsoft Dynamics 365 consulting, implementation and managed-services firm. We are not a demand-planning consultancy, and where work like this touches forecasting our contribution is the layer underneath it: whether the data can support what is being asked of it, whether the baseline and the adjustment are both being preserved where they need to be, and what has to be true in the operating model for an output to change a decision rather than a slide. On the reporting side that usually means Power BI on top of Dynamics 365, and on the process side it means the ownership questions that tend to sit unclaimed between an analytics team and a planning function.
If you want somewhere to start, check whether your planning system still holds last quarter’s statistical baseline separately from the adjusted forecast. If it does not, that is the first thing to fix, because every question in this article depends on it. We are happy to look at that with you.

