What your software engineers need to learn to build AI systems.
Good software engineers already have most of what it takes. The new parts are an evaluation mindset, output that varies between runs, retrieval and data handling, and judgment about where a person decides, learned fastest by shipping one workflow with someone who has done it.
veridive6 min read
Your best backend engineer already knows most of what it takes to build an AI system: integrations, data modeling, testing, observability, being on call. What trips good engineers up is smaller and stranger. A function returns a different answer each time it runs. The test suite can’t say what “correct” means. And a product decision hides inside every prompt.
The gap comes down to four things: an evaluation mindset, output that varies between runs, retrieval and data handling, and judgment about where a person decides. Engineers learn them fastest by shipping one real workflow next to someone who has done it.
What carries over from ordinary software engineering?
More than people expect. In a typical AI workflow, the model call is a small part of the code. The rest is ordinary software: connectors to the ERP or CRM, queues and retries, authentication, a review screen, logging, deployment. Integration and permissions are where projects usually get hard, and good engineers already know how to do both.
The habits carry over too: version control, code review, automated tests in the pipeline, staged rollouts, least privilege, actions that are safe to run twice, and a written postmortem after something breaks. Teams that drop these because “AI is different” regret it quickly.
What is genuinely new?
Seven areas, each with an old skill underneath it:
| Area | What carries over | What is new | How to practice it |
|---|---|---|---|
| Testing | Unit and integration tests | Evaluation sets that score language, reported by case type | Build a set with the workflow owner before writing a prompt |
| Behavior | Deterministic functions | Output that varies between runs; schemas and validation | Run one case many times and measure the spread |
| Data | Schemas, queries, pipelines | Retrieval quality and permissions at query time | Check whether the right passage came back, separately from the answer |
| Change | Releases and rollbacks | Prompts, sources and models as versioned artifacts | Tie every prompt change to an evaluation run |
| Security | Input validation | Prompt injection through the text the system reads | Add adversarial cases to the evaluation set |
| Product | Requirements and UX | Deciding where a person approves, and the review screen | Design the approval step with the people who will use it |
| Cost | Infrastructure budgets | Cost per task, driven by usage and review time | Measure cost per case during the pilot |
The “change” row surprises engineers most: the model provider can change the model under you. The note on LLMOps and MLOps covers what that means for operations.
Why is evaluation the core skill?
Because without it, nobody can say whether a change helped. Prompt edits turn into superstition, and model choice becomes a matter of taste. An engineer with an evaluation habit samples real cases with the owner, gets the reference answers from the people who own the work rather than writing them, uses exact checks for fields and short rubrics for text, holds part of the set back, reports results by case type and checks any automated grader against people.
Without an evaluation set, every prompt change is a guess.
Then the set runs on every change, like a test suite. The note on testing LLM applications shows how it fits with unit and end-to-end tests. Evaluation is also a specialty in its own right, and some teams treat it as a role of its own.
What do they need to know about data and retrieval?
When a document assistant gives a wrong answer, look at retrieval before the model: often the right passage never came back. Engineers need to understand how documents are split, indexed and searched, why keyword search still matters next to semantic search, and how to measure whether the right passage came back before judging the answer.
Four more things matter at work. Permissions must be enforced when the search runs, so no one sees a document they couldn’t open themselves. Sources change, so the index needs versions and a refresh routine. Scans, tables and mixed Turkish and English text need care before indexing. And personal data needs a plan for masking, logging and retention, agreed with the data protection officer.
How should they learn: courses or projects?
Courses give vocabulary; projects give judgment. Pairing on one real workflow with someone who has shipped one teaches more than a stack of course credentials, because the hard parts only appear on real data with a real owner.
In an illustrative case, two in-house engineers pair with an experienced team on drafting ERP entries from supplier invoices.
- Evaluation first. They sit with the accounts payable team and build the evaluation set with the owner before anyone writes a prompt.
- Build. They implement structured extraction with schema validation and the purchase order lookup, and join the designer and the accountants on the review screen.
- Shadow mode. They read the disagreements with the owner and learn to sort model errors from unclear rules.
- Handover. They take over the runbook, the dashboards and the prompt repository.
- First incident. A large supplier changes its invoice layout, and the override rate for that supplier rises. They add the new layout to the evaluation set, fix the extraction, re-run the set and deploy, without calling anyone.
The fifth step is the real graduation.
How do you know the team is ready to build alone?
Watch for behavior, not credentials:
- They write the evaluation set before the first prompt.
- They explain failures by case type, never with a single average.
- They design the review step with its users, and argue for where a person decides.
- They treat a prompt change like a code change: versioned, reviewed and evaluated.
- They can explain cost per case and what drives it.
- They have handled an incident from alert to written postmortem.
The end state is the one in the handover checklist: your team can change a prompt, re-run the evaluation set and roll back a model without calling the supplier. For who else belongs on the team, see AI project team roles.
Learn on a live workflow
Pick one workflow with a named owner and put your engineers on it from the first day, not at handover. Pairing during a Pilot to Production engagement, part of our custom AI software work, is the most direct route; AI enablement covers role-specific training for everyone around them.
Ask an assistant about this note