veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesEnablement

What your software engineers need to learn to build AI systems.

Good software engineers already have most of what it takes. The new parts are an evaluation mindset, output that varies between runs, retrieval and data handling, and judgment about where a person decides, learned fastest by shipping one workflow with someone who has done it.

veridive6 min read

Your best backend engineer already knows most of what it takes to build an AI system: integrations, data modeling, testing, observability, being on call. What trips good engineers up is smaller and stranger. A function returns a different answer each time it runs. The test suite can’t say what “correct” means. And a product decision hides inside every prompt.

The gap comes down to four things: an evaluation mindset, output that varies between runs, retrieval and data handling, and judgment about where a person decides. Engineers learn them fastest by shipping one real workflow next to someone who has done it.

What carries over from ordinary software engineering?

More than people expect. In a typical AI workflow, the model call is a small part of the code. The rest is ordinary software: connectors to the ERP or CRM, queues and retries, authentication, a review screen, logging, deployment. Integration and permissions are where projects usually get hard, and good engineers already know how to do both.

The habits carry over too: version control, code review, automated tests in the pipeline, staged rollouts, least privilege, actions that are safe to run twice, and a written postmortem after something breaks. Teams that drop these because “AI is different” regret it quickly.

What is genuinely new?

Seven areas, each with an old skill underneath it:

AreaWhat carries overWhat is newHow to practice it
TestingUnit and integration testsEvaluation sets that score language, reported by case typeBuild a set with the workflow owner before writing a prompt
BehaviorDeterministic functionsOutput that varies between runs; schemas and validationRun one case many times and measure the spread
DataSchemas, queries, pipelinesRetrieval quality and permissions at query timeCheck whether the right passage came back, separately from the answer
ChangeReleases and rollbacksPrompts, sources and models as versioned artifactsTie every prompt change to an evaluation run
SecurityInput validationPrompt injection through the text the system readsAdd adversarial cases to the evaluation set
ProductRequirements and UXDeciding where a person approves, and the review screenDesign the approval step with the people who will use it
CostInfrastructure budgetsCost per task, driven by usage and review timeMeasure cost per case during the pilot

The “change” row surprises engineers most: the model provider can change the model under you. The note on LLMOps and MLOps covers what that means for operations.

Why is evaluation the core skill?

Because without it, nobody can say whether a change helped. Prompt edits turn into superstition, and model choice becomes a matter of taste. An engineer with an evaluation habit samples real cases with the owner, gets the reference answers from the people who own the work rather than writing them, uses exact checks for fields and short rubrics for text, holds part of the set back, reports results by case type and checks any automated grader against people.

Without an evaluation set, every prompt change is a guess.

Then the set runs on every change, like a test suite. The note on testing LLM applications shows how it fits with unit and end-to-end tests. Evaluation is also a specialty in its own right, and some teams treat it as a role of its own.

What do they need to know about data and retrieval?

When a document assistant gives a wrong answer, look at retrieval before the model: often the right passage never came back. Engineers need to understand how documents are split, indexed and searched, why keyword search still matters next to semantic search, and how to measure whether the right passage came back before judging the answer.

Four more things matter at work. Permissions must be enforced when the search runs, so no one sees a document they couldn’t open themselves. Sources change, so the index needs versions and a refresh routine. Scans, tables and mixed Turkish and English text need care before indexing. And personal data needs a plan for masking, logging and retention, agreed with the data protection officer.

How should they learn: courses or projects?

Courses give vocabulary; projects give judgment. Pairing on one real workflow with someone who has shipped one teaches more than a stack of course credentials, because the hard parts only appear on real data with a real owner.

In an illustrative case, two in-house engineers pair with an experienced team on drafting ERP entries from supplier invoices.

  1. Evaluation first. They sit with the accounts payable team and build the evaluation set with the owner before anyone writes a prompt.
  2. Build. They implement structured extraction with schema validation and the purchase order lookup, and join the designer and the accountants on the review screen.
  3. Shadow mode. They read the disagreements with the owner and learn to sort model errors from unclear rules.
  4. Handover. They take over the runbook, the dashboards and the prompt repository.
  5. First incident. A large supplier changes its invoice layout, and the override rate for that supplier rises. They add the new layout to the evaluation set, fix the extraction, re-run the set and deploy, without calling anyone.

The fifth step is the real graduation.

How do you know the team is ready to build alone?

Watch for behavior, not credentials:

  • They write the evaluation set before the first prompt.
  • They explain failures by case type, never with a single average.
  • They design the review step with its users, and argue for where a person decides.
  • They treat a prompt change like a code change: versioned, reviewed and evaluated.
  • They can explain cost per case and what drives it.
  • They have handled an incident from alert to written postmortem.

The end state is the one in the handover checklist: your team can change a prompt, re-run the evaluation set and roll back a model without calling the supplier. For who else belongs on the team, see AI project team roles.

Learn on a live workflow

Pick one workflow with a named owner and put your engineers on it from the first day, not at handover. Pairing during a Pilot to Production engagement, part of our custom AI software work, is the most direct route; AI enablement covers role-specific training for everyone around them.

Ask an assistant about this note

EnablementEngineeringEvaluation

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What AI skills do software engineers need?

Mostly the skills they already have: integration, data modeling, testing, observability and incident handling. The new ones are building evaluation sets from real cases, handling output that varies between runs, retrieval and permission-aware search, treating prompts and models as versioned artifacts, defending against prompt injection, and designing the step where a person reviews and decides.

Should engineers learn AI from courses or from projects?

Both have a place, but projects teach what matters most. Courses give vocabulary; shipping one real workflow next to someone who has done it gives judgment about evaluation, failure modes and where a person decides. Pair on a workflow with a named owner and real data, from the evaluation set through the first incident after launch.

How do you know an engineering team is ready to build AI systems alone?

Look for behavior, not credentials. A ready team writes the evaluation set before the first prompt, explains failures by case type instead of quoting an average, designs the review step with its users, treats prompt changes like code changes, and has handled an incident on its own: noticed it, fixed or rolled back, re-run the evaluation set and written it up.