# Generative AI or machine learning: which does your problem need?

> In this field note, veridive explains how to tell whether a problem needs classic machine learning or a language model. Forecasting, scoring and anomaly detection on structured data suit machine learning; reading and writing documents and conversations suit language models. It compares data, skills, evaluation and costs, and shows an illustrative system that combines both.

Forecasting, scoring and anomaly detection on structured data are usually classic machine learning. Reading, writing and reasoning over documents and conversations is where language models help. Many good systems use both, each for the part it does well.

## Key takeaways

- Classic machine learning predicts numbers and categories from structured history; language models read and write text, documents and conversations.
- Many good systems use both: one model forecasts or scores, and a language model reads the messages and explains the result.
- Compute in code and explain in language: never let a language model do the arithmetic on a table.
- Machine learning needs labeled history; language models need documents, examples and reference answers to test against.

Two requests arrive in the same planning meeting. The sales director wants generative AI to predict which customers will leave. The service manager wants a machine learning model to answer customer emails. Each has picked the tool that suits the other’s problem.

Forecasting, scoring and anomaly detection on structured data are usually classic machine learning. Reading, writing and reasoning over documents and conversations is where language models help. Many good systems use both, each for the part it does well.

## What is the difference, in plain terms?

Classic machine learning learns from labeled history, rows and columns of past cases with known outcomes, to predict a number or a category: demand for the coming weeks, the chance an invoice is paid late, whether a transaction looks unusual. It is usually trained on your data for one task.

Generative AI, in business mostly large language models, comes pre-trained on vast amounts of general text and is steered with instructions, examples and documents. It reads and produces language: it drafts, summarizes, extracts fields, classifies messages and answers questions from sources.

Both are probabilistic, both can be wrong, and both need evaluation on your own cases. The difference lies in what goes in and what comes out.

## Which problems suit classic machine learning?

Problems with structured inputs, plenty of history and a numeric or categorical answer:

- **Forecasting:** demand, cash, staffing, call volumes.
- **Scoring:** churn risk, lead quality, the likelihood of late payment.
- **Anomaly detection:** unusual transactions, sensor readings, return patterns.
- **Recommendations:** the next product, the next best action.

Quality is measured against history the model hasn’t seen, for example forecast error on past months, and a trained model is cheap to run per prediction. It needs care of its own: a demand model trained before a new sales channel opened knows nothing about it until it is retrained, so retraining is part of the running cost. And some of these problems need neither kind of AI: a moving average or a rule may be enough, and is easier to explain.

## Which problems suit language models?

Problems where the input is language or documents:

- **Reading:** supplier emails, contracts, invoices and call transcripts, with fields extracted into a structure.
- **Classifying:** the intent of a message or the category of a request.
- **Drafting:** replies, summaries, explanations and first versions of reports.
- **Answering:** questions from a document library, with a citation for every claim.

Language models need no training on your data to start, only instructions, examples and sources, but each call costs more than a classic prediction, and its cost grows with the length of the text. Side by side:

| Aspect | Classic machine learning | Language models |
|---|---|---|
| Typical inputs | Rows and columns: transactions, sales history, sensor data | Text, documents, emails, conversations |
| Typical outputs | A number, a probability or a category | A draft, a summary, extracted fields, a cited answer |
| Data needed | Labeled history: many past cases with known outcomes | Instructions, examples, source documents and reference answers |
| How quality is measured | Error against held-out history | An evaluation set with scoring rules, plus human review |
| How costs behave | Mostly up front, then cheap per prediction | Little up front, then a cost per call that grows with text and volume |

## When do you need both?

When the problem has both kinds of input. Take an illustrative distributor’s demand planning. A statistical or machine learning model produces the weekly forecast for each item from sales history, promotions and seasonality. A language model reads incoming supplier emails and turns delays and quantity changes into structured fields for the planning system. It also drafts the planner’s explanation of why an item’s forecast moved, quoting the forecast model’s own figures and linking the email and the promotion entry behind them. The planner approves every replenishment. [Supply chain intelligence](https://veridive.com/solutions/supply-chain-intelligence/) is built on this division, and [language models in demand planning](https://veridive.com/insights/llms-in-demand-planning/) covers it in depth.

The division has one hard rule.

> Compute in code, explain in language.

A language model generates likely text, and to it a number is just more text: plausible, sometimes wrong, and not always the same twice. Given a table, it can misread a column or invent a total. So let the database, the ERP or the forecasting model calculate, pass the results to the language model as facts to explain, and check that every number in its draft matches a number it was given.

## What does each need in data, skills and evaluation?

**Machine learning** needs labeled history: past cases with known outcomes, consistent definitions and enough examples of the rare outcome you care about. The skills are data science plus the engineering to retrain and monitor models. Evaluation means backtesting on periods the model hasn’t seen, then watching for drift as the world changes.

**Language models** need your documents, a few good examples and a few hundred real cases with reference answers written by the people who own the work. The skills are software engineering around the model: retrieval, prompts, review screens and integrations. Evaluation means an evaluation set with scoring rules, re-run whenever the provider updates the model. Whether to fine-tune is a separate question, covered in [RAG or fine-tuning](https://veridive.com/insights/rag-vs-fine-tuning/).

Both need a baseline, an owner and monitoring, the groundwork that [data and AI foundations](https://veridive.com/services/data-ai-foundations/) covers.

## How do you avoid choosing by fashion?

Start from the input and the output, not from the label on the tool:

1. **What goes in:** rows and columns, or language and documents?
2. **What comes out:** a number or a category, or text?
3. **Is the answer exact?** Then it belongs in code, not in any model; see [when not to use AI](https://veridive.com/insights/when-not-to-use-ai/).
4. **What data do you have:** labeled history, or examples and reference answers?
5. **What would you measure** to know it works?

Two traps come up again and again. The first is a chat interface on top of sales data, asked to predict future sales: a language model doing a forecaster’s job. The second runs the other way: a custom classifier trained for months to sort emails that a language model with a good evaluation set could handle from the start. At very high volume, though, a small trained classifier can be cheaper per case, so test both on the same examples before deciding.

## Match the tool to the input

Structured history in, a number or a category out: classic machine learning. Language in, language out: a language model. Both kinds of input: split the work, and compute in code. When the answer isn’t obvious, test both approaches on the same cases; [AI strategy and discovery](https://veridive.com/services/ai-strategy/) makes that choice on your own examples, before anyone builds.

## Frequently asked questions

### What is the difference between generative AI and machine learning?

Classic machine learning learns patterns from labeled history to predict a number or a category, such as demand, churn risk or an unusual transaction. Generative AI, in business mostly language models, reads and produces text: it drafts, summarizes, extracts fields from documents and answers questions from sources. Both are machine learning in the broad sense, suited to different inputs and outputs.

### Can a language model do forecasting?

It shouldn’t produce the numbers. Forecasts come from statistical or machine learning models trained on sales history, which can be tested against past periods. A language model helps around the forecast: reading supplier and customer messages for signals, explaining in plain words why a forecast changed, and answering planners’ questions with sources. Compute in code, explain in language.
