veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesOperations & ERP

What language models add to demand planning, and what they don’t.

Forecasts are a job for statistical and machine learning models on sales history. Language models help around them: reading supplier and customer messages for signals, explaining forecast changes in plain words and answering planners’ questions with sources.

veridive6 min read

The forecast for a fast-moving item jumped overnight, and the planner has a review meeting in an hour. There are three places to look: the forecasting tool, which shows the new number but not the reason; the promotion calendar in a shared spreadsheet; and a sales inbox full of customer order notes. Most of the hour goes into finding out why.

Language models don’t belong inside that forecast. They belong around it. The forecast comes from statistical and machine learning models on sales history; the language model never produces the numbers. What it can do is read the text those models can’t, explain changes in plain words and answer planners’ questions with sources.

What does a forecasting model already do well?

Statistical and machine learning forecasting models learn from sales history, calendars, prices and promotions. They capture trend, seasonality and known events for thousands of items at once, the same way every time, and their accuracy can be measured by testing them on past periods they didn’t see. Those three properties, scale, repeatability and measurable error, are exactly what a forecast needs.

A language model has none of them for this job. It produces text by prediction, and asked for a forecast it will return a plausible number without the discipline that makes a number trustworthy. Our note on generative AI or machine learning draws the same line more broadly.

The forecasting model produces the numbers. The language model explains them.

Where do planners lose time today?

Rarely in the forecasting tool. The time goes into the work around it:

  • Reading. Supplier emails, customer notes, sales comments and promotion plans arrive as text, in several languages and formats.
  • Chasing reasons. The tool shows that a number changed, not why, so planners rebuild the story from several systems.
  • Re-typing. Dates and quantities from supplier confirmations are copied into the plan by hand.
  • Explaining. Overrides and exceptions have to be justified in the planning meeting, usually from memory.

Re-typing is a document-extraction job, covered on the supply chain page linked at the end. The other three are language work, and that is where a language model helps most.

Which text signals can a language model read?

Text becomes useful to planning once it is turned into a structured signal with a source attached. Three kinds matter most:

  • Supplier delay emails. “Our line is down for maintenance; your order will ship late.” The model extracts the supplier, the order, the new date and the reason, and checks them against the open order. Documents from carriers and forwarders follow the same pattern, as our note on logistics documents shows.
  • Customer order notes. “We are adding two stores in the region and expect larger orders from the autumn.” The signal is the customer, the items, the direction, the timing and the note itself.
  • Promotion calendars. Entries such as “holiday campaign on the drinks range” become events with items, dates and a mechanism, ready for the planning team to check.

None of these signals changes the forecast by itself. A planner decides whether it justifies an override. Separately, the data team can test whether a type of signal improves accuracy when fed to the forecasting model as an input, and add it only if the test says so.

How can AI explain a forecast change?

Many forecasting systems can show which drivers moved a forecast: recent sales, seasonality, a promotion flag, a price change. The language model receives the old and new forecast, the driver breakdown computed by the forecasting system, and the text signals linked to the item and period. It writes the explanation, and every statement cites a driver or a document.

In an illustrative case, a planner asks, “Why did the forecast for this item rise?” The answer names the promotion flag as the main driver and cites the promotion calendar entry behind it, then adds two customer order notes that mention larger orders for the same weeks, each linked. The planner can check all three sources in a minute and decide whether the increase is right, too high or too low.

The same pattern answers the other questions planners ask every day: which open orders a supplier delay affects, which customers ordered well above forecast, what changed since the last planning meeting. Each answer is built from figures the planning system computed, and cites them.

This fails in one predictable way: when the sources are thin, a language model will offer a plausible cause anyway. The rule that prevents it is simple. An explanation may cite only recorded drivers and linked documents, and when neither accounts for a change, it says “no recorded reason” and asks. Our note on management reports with AI applies the same rule to commentary.

Where should planners stay in charge?

  • Replenishment. The system suggests; planners approve every replenishment.
  • External commitments. Purchase orders, confirmed delivery dates and anything promised to a supplier or customer always need a person.
  • Overrides. Planners decide whether a signal changes the forecast, and each override is recorded with its source and reason.
  • Uncertain readings. Extractions the system is unsure about go to a person for checking, not into the signal list.
  • The plan itself. The language model reads the forecast and the plan; it never edits either.

Recorded overrides pay off later: they show whether planners’ adjustments actually improved accuracy, a question memory alone can’t answer.

How do you test the value?

Record a baseline first, then test each part on its own terms:

  1. Planner time: hours spent reading, chasing reasons and preparing explanations, before and after.
  2. Explanation quality: a sample of explanations reviewed by planners against the drivers and sources, using past forecast changes as the test set.
  3. Signal extraction: a labeled sample of emails and notes, checked for missed signals and wrong dates.
  4. Forecast accuracy: measured on selected items, and credited to a text signal only when a backtest shows the gain.

Expect most of the value in time and in better-documented decisions. Accuracy gains are possible, but they have to be shown, not assumed.

Start with one weekly question

Pick one planning team and one question it asks every week, such as why forecasts moved for its top items. The supply chain intelligence page shows how a planner copilot fits next to supplier confirmation work, and AI strategy and discovery is the step that tells you whether the time saved justifies the build.

Ask an assistant about this note

Operations & ERPPlanningSupply chain

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

Can a large language model forecast demand?

It isn’t the right tool for the forecast itself. Demand forecasts should come from statistical and machine learning models trained on sales history, which are repeatable and can be tested against past results. Language models add value around the forecast: reading text signals such as supplier emails and order notes, explaining forecast changes and answering planners’ questions with sources.

What is a planner copilot?

A planner copilot is an assistant that answers demand and supply planners’ questions in plain language, such as why a forecast rose or which orders a supplier delay affects, citing the figures and documents behind each answer. It reads the forecast and the plan but doesn’t change them. Planners approve replenishment, overrides and every external commitment.