# Prompts are code: how to version, review and test them.

> In this field note, veridive explains how to manage prompts like code. A prompt change is a behavior change, so prompts belong in version control, reviewed like code and tied to evaluation results. The note covers the parts of a production prompt, what a review checks, release and rollback, and how to refactor prompts that sprawl.

A prompt change is a behavior change. Keep prompts in version control, review them like code, tie each version to evaluation results, and never edit them live in a dashboard.

## Key takeaways

- A prompt change is a behavior change: keep prompts in version control and never edit them live in a dashboard.
- A readable production prompt has four parts: role and task, rules with reasons, real examples and an output format.
- Link every prompt version to its evaluation run, promote it through environments, and keep the previous version one switch away.
- When every incident adds a rule, refactor: merge rules, delete what no test needs, and move deterministic rules into code.

The assistant’s answers changed on a Tuesday, and nobody had deployed anything. Someone in operations had edited the prompt in a vendor dashboard to fix a complaint. The edit went straight to production, and the old version was gone.

A prompt change is a behavior change, often a bigger one than a code change. Keep prompts in version control, review them like code, tie each version to evaluation results, and never edit them live in a dashboard.

## Why do prompts need the discipline of code?

One sentence in a prompt can change behavior across every case type, including ones far from the problem it was meant to fix. Without versions, nobody can answer basic questions: what did the prompt say when this answer was given, who changed it and why, and what else changed at the same time? Prompts are also part of what a system’s owner must be able to hand over, maintain and audit, alongside the code and the evaluation set. And they are often edited by people who don’t think of themselves as developers, which is exactly why the process has to be easy to follow.

## Where should prompts live?

In the repository, next to the code that uses them, as plain template files with clearly marked variables. Not as strings buried in code, and not only in a dashboard. Keep one file per task and per model, because a prompt tuned for one model rarely suits another. Every release pins a prompt version, and every call logs it.

If business users need to change wording, such as tone guidance, give them a reviewed path: a change request that runs the same checks as any other change. A prompt management tool can help, as long as it keeps every version, records who changed what and follows the same review and release path. The rule is about the process, not the tool.

## What makes a production prompt readable?

Four parts, in a consistent order:

1. **Role and task.** Who the model acts as and the single job of this step, in two or three sentences.
2. **Rules.** Short, numbered and grouped, each with its reason when the reason isn’t obvious: “Never state a refund amount; amounts come from the order system.”
3. **Examples from real cases.** A few masked inputs with approved outputs, chosen for the tricky patterns, not just the easy one.
4. **Output format.** The exact structure or schema expected, with one example.

Keep examples consistent with the rules. Examples teach patterns, and when one contradicts a rule, the model often follows the example. A short header completes the prompt: the owner, the purpose, the evaluation set that covers it, and known limitations.

## What should a prompt review check?

- **The reason.** Which behavior is being fixed, and which evaluation cases show it?
- **The evidence.** Did the full evaluation run pass on every important case type, not just the target?
- **Conflicts.** Does the new rule contradict an existing one?
- **The right home.** Would a threshold, calculation or permission be safer in code?
- **Testable wording.** “Keep it brief” is a wish; “no more than five sentences” can be checked.

Every approved version gets a change-log entry: what changed, why, who approved it, and a link to its evaluation run. The set itself is the one described in [evaluation sets are the new requirements document](https://veridive.com/insights/evaluation-sets-are-the-new-requirements/).

> A prompt edit in a live dashboard is an untested release with no way back.

## How do you release and roll back a prompt?

Promote the same versioned file through environments: development, a test run against the full evaluation set, shadow or staging on live traffic, then production. Risky changes can go to one queue or one case type first. The previous version stays deployable, so rollback is a configuration switch, not an edit under pressure. Because every answer is logged with its prompt version, an odd reply found weeks later can be traced to the exact text that produced it. Where prompt tests sit among the other layers is covered in [how to test an LLM application](https://veridive.com/insights/testing-llm-applications/).

An illustrative case shows the pattern. Customers complain that a support assistant apologizes too much, so an engineer adds a rule: “Don’t apologize; state the facts.” On the complaint cases, it works. The evaluation run before release shows the cost: on lost-shipment cases, where the tone guide asks for an apology, replies now read as cold, and that case type falls below its threshold. The change doesn’t ship. The next version narrows the rule to “apologize once, and only when the company made a mistake” and adds one real lost-shipment example. Both case types pass, and the change log links both runs.

## How do you stop prompts from sprawling?

Sprawl costs money as well as clarity, because every rule is sent with every call. Watch for the signs: a new rule after every incident, rules that contradict each other, capitals and “IMPORTANT” used as emphasis, rules nobody can explain, and a prompt that keeps growing while quality stays flat.

Refactor on a schedule, with the evaluation set as the safety net. Merge overlapping rules. Remove a rule and re-run the set; if nothing gets worse, it can go. Move deterministic rules into code, and split an overloaded prompt into separate steps. The change log tells you where each rule came from.

## No prompt ships without a run

Move every production prompt into version control, write a change-log entry for its current state and link it to an evaluation run. From then on, no prompt ships without a new entry and a new run. Prompts are documented and handed over with every [custom AI software](https://veridive.com/services/custom-ai-software/) system we build, and in our [AI reliability](https://veridive.com/services/ai-reliability/) work every prompt change triggers a regression run.

## Frequently asked questions

### What is prompt versioning?

Prompt versioning means managing the prompts of an AI system the way you manage code: stored in version control, changed through reviewed requests, tested against an evaluation set, released through environments and rolled back when needed. Each version gets a change-log entry saying what changed, why, who approved it and which evaluation run it passed.

### How do you test a prompt change before release?

Run the full evaluation set with the new prompt and compare results with the current version by case type, not just on the cases the change was meant to fix. A change that fixes one complaint can quietly break another kind of case. Ship it only if every important case type stays within its threshold, and record the run in the change log.
