Accuracy is not the finish line: taking a forecasting model to production
Six lessons from moving a spare-parts forecasting model from a benchmark to a system that planners use every day.
A forecasting pilot usually ends with one number: the model is more accurate than the old method. That number is necessary, but it is not sufficient. When the model goes to production, the questions change. Planners ask what to order, not what the error is. Operations teams ask what happens when the model is wrong. Security teams ask who can see the data.
I am building Velocis Foresight, a forecasting platform for the spare parts of a leading automotive manufacturer: about 22,000 parts, with 52 months of history. These are six lessons from taking it past the benchmark.
1. Beat a baseline the business already trusts
A model that is “accurate” means nothing until you compare it with the method that the planners use today. We benchmarked against 15 statistical and deep-learning models, and we reported forecast value added: the improvement over a simple naive forecast.
The final system cut pooled forecast error (WAPE) from 17.2% to 14.9%, and raised forecast value added from 18.5% to 29.5%. The second number is the one that a planning manager understands, because it compares the model with something they already know.
2. Choose the metric that matches the money
Spare-parts demand is very uneven. A few parts carry most of the volume, and thousands of parts sell once or twice a year. A statement such as “the model is accurate for most parts” can therefore be true by demand volume and false by part count.
Both views are legitimate, but they answer different questions. Volume-weighted accuracy tells you about inventory value. Per-part accuracy tells you about the planners’ daily work. Say which one you report. In our case, the parts forecast within 20% error carry 83.6% of demand volume, and we state it in exactly those words.
3. A forecast is not a decision
Planners do not order a forecast; they order a quantity. The model therefore predicts a range of outcomes (demand quantiles), not one number. A business rule then chooses which quantile to plan to, from the risk class of the part. For example, a part whose shortage stops a repair can be planned higher than a part that is easy to reorder.
This step is where most of the business value is, and it is not machine learning. It is a clear, auditable rule that the planning team can read and change.
4. When two methods tie, pick the one you can govern
Two fine-tuning approaches can score almost the same. When they do, choose the one that is easier to govern. For Foresight, that is a small adapter over a frozen foundation model. The customer receives a file of about 10 MB. The base model is never modified, and removing the adapter restores the original model exactly.
A small accuracy difference is often within the noise of one test period. A governance property is permanent. When the numbers are close, choose the method that your security and compliance teams can approve.
5. Make the assistant unable to invent numbers
The platform includes a planning assistant, in the console and in Webex. A language model chooses a tool, the forecasting engine computes the answer, and the model writes the reply. Before the reply is shown, a grounding check compares every figure in it with the figures the tools returned. A reply with a number that no tool produced is refused.
For a planning tool, one invented number is enough to lose the users’ trust. The check makes that failure impossible, not only unlikely.
6. Plan for drift from the first day
Demand changes: new models launch, old ones retire, and seasons shift. The system logs every forecast and compares it with actual demand. When the error drifts past a limit, an alert goes to the planners and requests a model refresh. The alert does not start the retraining by itself; a person decides.
Monitoring is not a phase after launch. It is part of the design, because a forecasting model starts to age on the day it ships.
The common thread
None of these lessons is about a better model. They are about the system around the model: the baseline, the metric, the decision rule, the governance, the guardrails and the monitoring. That system is what turns an accurate pilot into a tool that people use.