Field notesStrategy & leadership
Generative AI or machine learning: which does your problem need?
Forecasting, scoring and anomaly detection on structured data are usually classic machine learning. Reading, writing and reasoning over documents and conversations is where language models help. Many good systems use both, each for the part it does well.
veridive6 min read
Two requests arrive in the same planning meeting. The sales director wants generative AI to predict which customers will leave. The service manager wants a machine learning model to answer customer emails. Each has picked the tool that suits the other’s problem.
Forecasting, scoring and anomaly detection on structured data are usually classic machine learning. Reading, writing and reasoning over documents and conversations is where language models help. Many good systems use both, each for the part it does well.
What is the difference, in plain terms?
Classic machine learning learns from labeled history, rows and columns of past cases with known outcomes, to predict a number or a category: demand for the coming weeks, the chance an invoice is paid late, whether a transaction looks unusual. It is usually trained on your data for one task.
Generative AI, in business mostly large language models, comes pre-trained on vast amounts of general text and is steered with instructions, examples and documents. It reads and produces language: it drafts, summarizes, extracts fields, classifies messages and answers questions from sources.
Both are probabilistic, both can be wrong, and both need evaluation on your own cases. The difference lies in what goes in and what comes out.
Which problems suit classic machine learning?
Problems with structured inputs, plenty of history and a numeric or categorical answer:
- Forecasting: demand, cash, staffing, call volumes.
- Scoring: churn risk, lead quality, the likelihood of late payment.
- Anomaly detection: unusual transactions, sensor readings, return patterns.
- Recommendations: the next product, the next best action.
Quality is measured against history the model hasn’t seen, for example forecast error on past months, and a trained model is cheap to run per prediction. It needs care of its own: a demand model trained before a new sales channel opened knows nothing about it until it is retrained, so retraining is part of the running cost. And some of these problems need neither kind of AI: a moving average or a rule may be enough, and is easier to explain.
Which problems suit language models?
Problems where the input is language or documents:
- Reading: supplier emails, contracts, invoices and call transcripts, with fields extracted into a structure.
- Classifying: the intent of a message or the category of a request.
- Drafting: replies, summaries, explanations and first versions of reports.
- Answering: questions from a document library, with a citation for every claim.
Language models need no training on your data to start, only instructions, examples and sources, but each call costs more than a classic prediction, and its cost grows with the length of the text. Side by side:
| Aspect | Classic machine learning | Language models |
|---|---|---|
| Typical inputs | Rows and columns: transactions, sales history, sensor data | Text, documents, emails, conversations |
| Typical outputs | A number, a probability or a category | A draft, a summary, extracted fields, a cited answer |
| Data needed | Labeled history: many past cases with known outcomes | Instructions, examples, source documents and reference answers |
| How quality is measured | Error against held-out history | An evaluation set with scoring rules, plus human review |
| How costs behave | Mostly up front, then cheap per prediction | Little up front, then a cost per call that grows with text and volume |
When do you need both?
When the problem has both kinds of input. Take an illustrative distributor’s demand planning. A statistical or machine learning model produces the weekly forecast for each item from sales history, promotions and seasonality. A language model reads incoming supplier emails and turns delays and quantity changes into structured fields for the planning system. It also drafts the planner’s explanation of why an item’s forecast moved, quoting the forecast model’s own figures and linking the email and the promotion entry behind them. The planner approves every replenishment. Supply chain intelligence is built on this division, and language models in demand planning covers it in depth.
The division has one hard rule.
Compute in code, explain in language.
A language model generates likely text, and to it a number is just more text: plausible, sometimes wrong, and not always the same twice. Given a table, it can misread a column or invent a total. So let the database, the ERP or the forecasting model calculate, pass the results to the language model as facts to explain, and check that every number in its draft matches a number it was given.
What does each need in data, skills and evaluation?
Machine learning needs labeled history: past cases with known outcomes, consistent definitions and enough examples of the rare outcome you care about. The skills are data science plus the engineering to retrain and monitor models. Evaluation means backtesting on periods the model hasn’t seen, then watching for drift as the world changes.
Language models need your documents, a few good examples and a few hundred real cases with reference answers written by the people who own the work. The skills are software engineering around the model: retrieval, prompts, review screens and integrations. Evaluation means an evaluation set with scoring rules, re-run whenever the provider updates the model. Whether to fine-tune is a separate question, covered in RAG or fine-tuning.
Both need a baseline, an owner and monitoring, the groundwork that data and AI foundations covers.
How do you avoid choosing by fashion?
Start from the input and the output, not from the label on the tool:
- What goes in: rows and columns, or language and documents?
- What comes out: a number or a category, or text?
- Is the answer exact? Then it belongs in code, not in any model; see when not to use AI.
- What data do you have: labeled history, or examples and reference answers?
- What would you measure to know it works?
Two traps come up again and again. The first is a chat interface on top of sales data, asked to predict future sales: a language model doing a forecaster’s job. The second runs the other way: a custom classifier trained for months to sort emails that a language model with a good evaluation set could handle from the start. At very high volume, though, a small trained classifier can be cheaper per case, so test both on the same examples before deciding.
Match the tool to the input
Structured history in, a number or a category out: classic machine learning. Language in, language out: a language model. Both kinds of input: split the work, and compute in code. When the answer isn’t obvious, test both approaches on the same cases; AI strategy and discovery makes that choice on your own examples, before anyone builds.
Ask an assistant about this note