veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesStrategy & leadership

Choose the workflow before the model.

Most stalled AI projects didn’t fail on technology. They failed because nobody picked a specific piece of work to change. Here’s how we choose — and what we measure before writing any code.

veridive5 min read

Many teams have already bought AI. Licenses are active, a few enthusiasts use them daily, and a pilot or two has come and gone. And yet, ask what has changed in the way work gets done, and the room goes quiet.

The problem is rarely the model. Models are capable enough for a remarkable range of tasks. The problem is that the project started with the technology and went looking for a use — instead of starting with a piece of work and asking whether the technology can change it.

Start with the work, not the tool

A workflow is a repeated sequence of steps that someone owns, with inputs, decisions and outputs. “Improve customer service with AI” is not a workflow. “Draft replies to delivery-status emails for the retail support team, using order data and our tone guide” is.

The difference matters because a workflow can be measured. You can count how many cases arrive each week, how long each one takes, how often the answer is wrong, and what that costs. Without those numbers, you can’t tell whether AI helped — and neither can the people who have to approve the budget for the next phase.

A model can’t fix a process nobody owns.

Four questions we ask before anything else

When we sit down with a team, we look for workflows that pass four tests.

  1. Is it frequent? A task that happens twice a year rarely justifies a system. We look for work that happens daily or weekly, in volume.
  2. Is it bounded? The best first workflows have clear inputs and a recognizable good outcome. “Match this invoice to a purchase order” is bounded. “Decide our pricing strategy” is not — at least not yet.
  3. Is the data reachable? If the information lives in approved systems we can access with the right permissions, we can build. If it lives in someone’s head or a folder nobody can find, the first project is a different one.
  4. Does someone own it? Every workflow we change needs a person who cares about the result, can make decisions, and will still be there after launch. No owner, no project.
Four tests applied to candidate workflows Illustrative diagram: a column of candidate workflows passes four gates — frequent, bounded, reachable data, an owner. Fewer candidates pass each gate; those that pass all four become candidates for a first project. Frequent?Bounded?Reachable?Owned? Candidates On the listSet aside at each gate
Fig. 1 — Four tests, one list (illustrative)

A workflow that passes all four is a candidate. Then we measure.

Measure the baseline first

Before we write a line of code, we record how the work is done today: volume, time per case, error rate and cost. We take a sample of real cases — usually a few hundred — and turn it into an evaluation set: examples of inputs with the outcome a good expert would reach.

That evaluation set becomes the most important document in the project. It tells us which model performs best on your examples, not on a public benchmark. It tells us when the system is good enough to go live. And it tells us, months later, whether a model update has quietly made things worse.

What we measureWhy it matters
Volume per weekSizes the opportunity
Time per caseShows where the hours go
Error or rework rateDefines “good enough”
Cost per caseKeeps the business case honest

Choose the model last

Once the workflow and the evaluation set exist, choosing a model becomes an experiment instead of a debate. We run the candidates — often a large commercial model, a smaller cheaper one, and an open-weight model that can run on your infrastructure — against the same examples, and compare quality, cost and speed.

Sometimes the most capable model wins. Often a smaller one is good enough for most cases, with a larger model or a person handling the rest. The point is that the choice is made on evidence from your work, and it can be revisited when prices or models change.

Design where the person decides

The most important design decision in an AI workflow is not the prompt. It’s where a person approves, edits or overrides. We decide that early: which actions are reversible, which cost money, which touch customers, and which need a second pair of eyes.

A good first workflow keeps people firmly in control and removes the tedious parts of their work: gathering information, drafting, checking against policy. It doesn’t remove their judgment.

What a good first workflow looks like

  • Narrow and frequent: one team, one type of case, hundreds of cases a month.
  • Draft, don’t decide: the system prepares, a person confirms.
  • Evidence attached: every recommendation shows its sources.
  • Measured weekly: the same metrics as the baseline, reviewed with the owner.

It’s less dramatic than a company-wide transformation. It’s also how transformations actually start. One workflow that works — and that people trust — does more for an organization’s AI ambitions than ten pilots that impressed in a demo and disappeared.

Start here

If you want to try this in your own organization, pick three candidate workflows and run them through the four questions. Write the owner’s name next to each one. If you can’t, that’s your answer. For a longer list, score the candidates on evidence; for a single one, check its readiness.

And if you’d like a second opinion, we’re happy to help. That’s what a Discovery Sprint is for.

Ask an assistant about this note

StrategyWorkflowsEvaluation

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

How do you choose the first AI use case?

Pick a workflow that passes four tests: it is frequent, it is bounded with clear inputs and a recognizable good outcome, its data lives in approved systems you can reach, and a named person owns it. Then measure how the work runs today before building anything, so you can tell whether AI actually helped.

Why do AI pilots stall?

Usually because nobody picked a specific piece of work to change, not because the model was weak. Without a named workflow, an owner and a baseline, a pilot can impress in a demo but has no way to prove its value or earn the budget for the next phase.