veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesEngineering

Prompts are code: how to version, review and test them.

A prompt change is a behavior change. Keep prompts in version control, review them like code, tie each version to evaluation results, and never edit them live in a dashboard.

veridive5 min read

The assistant’s answers changed on a Tuesday, and nobody had deployed anything. Someone in operations had edited the prompt in a vendor dashboard to fix a complaint. The edit went straight to production, and the old version was gone.

A prompt change is a behavior change, often a bigger one than a code change. Keep prompts in version control, review them like code, tie each version to evaluation results, and never edit them live in a dashboard.

Why do prompts need the discipline of code?

One sentence in a prompt can change behavior across every case type, including ones far from the problem it was meant to fix. Without versions, nobody can answer basic questions: what did the prompt say when this answer was given, who changed it and why, and what else changed at the same time? Prompts are also part of what a system’s owner must be able to hand over, maintain and audit, alongside the code and the evaluation set. And they are often edited by people who don’t think of themselves as developers, which is exactly why the process has to be easy to follow.

Where should prompts live?

In the repository, next to the code that uses them, as plain template files with clearly marked variables. Not as strings buried in code, and not only in a dashboard. Keep one file per task and per model, because a prompt tuned for one model rarely suits another. Every release pins a prompt version, and every call logs it.

If business users need to change wording, such as tone guidance, give them a reviewed path: a change request that runs the same checks as any other change. A prompt management tool can help, as long as it keeps every version, records who changed what and follows the same review and release path. The rule is about the process, not the tool.

What makes a production prompt readable?

Four parts, in a consistent order:

  1. Role and task. Who the model acts as and the single job of this step, in two or three sentences.
  2. Rules. Short, numbered and grouped, each with its reason when the reason isn’t obvious: “Never state a refund amount; amounts come from the order system.”
  3. Examples from real cases. A few masked inputs with approved outputs, chosen for the tricky patterns, not just the easy one.
  4. Output format. The exact structure or schema expected, with one example.

Keep examples consistent with the rules. Examples teach patterns, and when one contradicts a rule, the model often follows the example. A short header completes the prompt: the owner, the purpose, the evaluation set that covers it, and known limitations.

What should a prompt review check?

  • The reason. Which behavior is being fixed, and which evaluation cases show it?
  • The evidence. Did the full evaluation run pass on every important case type, not just the target?
  • Conflicts. Does the new rule contradict an existing one?
  • The right home. Would a threshold, calculation or permission be safer in code?
  • Testable wording. “Keep it brief” is a wish; “no more than five sentences” can be checked.

Every approved version gets a change-log entry: what changed, why, who approved it, and a link to its evaluation run. The set itself is the one described in evaluation sets are the new requirements document.

A prompt edit in a live dashboard is an untested release with no way back.

How do you release and roll back a prompt?

Promote the same versioned file through environments: development, a test run against the full evaluation set, shadow or staging on live traffic, then production. Risky changes can go to one queue or one case type first. The previous version stays deployable, so rollback is a configuration switch, not an edit under pressure. Because every answer is logged with its prompt version, an odd reply found weeks later can be traced to the exact text that produced it. Where prompt tests sit among the other layers is covered in how to test an LLM application.

An illustrative case shows the pattern. Customers complain that a support assistant apologizes too much, so an engineer adds a rule: “Don’t apologize; state the facts.” On the complaint cases, it works. The evaluation run before release shows the cost: on lost-shipment cases, where the tone guide asks for an apology, replies now read as cold, and that case type falls below its threshold. The change doesn’t ship. The next version narrows the rule to “apologize once, and only when the company made a mistake” and adds one real lost-shipment example. Both case types pass, and the change log links both runs.

How do you stop prompts from sprawling?

Sprawl costs money as well as clarity, because every rule is sent with every call. Watch for the signs: a new rule after every incident, rules that contradict each other, capitals and “IMPORTANT” used as emphasis, rules nobody can explain, and a prompt that keeps growing while quality stays flat.

Refactor on a schedule, with the evaluation set as the safety net. Merge overlapping rules. Remove a rule and re-run the set; if nothing gets worse, it can go. Move deterministic rules into code, and split an overloaded prompt into separate steps. The change log tells you where each rule came from.

No prompt ships without a run

Move every production prompt into version control, write a change-log entry for its current state and link it to an evaluation run. From then on, no prompt ships without a new entry and a new run. Prompts are documented and handed over with every custom AI software system we build, and in our AI reliability work every prompt change triggers a regression run.

Ask an assistant about this note

EngineeringPromptsChange control

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What is prompt versioning?

Prompt versioning means managing the prompts of an AI system the way you manage code: stored in version control, changed through reviewed requests, tested against an evaluation set, released through environments and rolled back when needed. Each version gets a change-log entry saying what changed, why, who approved it and which evaluation run it passed.

How do you test a prompt change before release?

Run the full evaluation set with the new prompt and compare results with the current version by case type, not just on the cases the change was meant to fix. A change that fixes one complaint can quietly break another kind of case. Ship it only if every important case type stays within its threshold, and record the run in the change log.