veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notes · Topic

Engineering.

The language model is one component of an AI system, and rarely the one that fails first. These notes are for engineers and technical leads who build the software around it: the tools, checks, tests and screens that make its output safe to use.

10 notes

Notes on this topic

  1. Evaluation sets are the new requirements document.

    When the output is language, “done” is hard to define. A few hundred real examples with expert-approved answers make quality measurable — and turn model choice into an experiment instead of an opinion.

    Engineering6 min

  2. Where agents should — and shouldn’t — act on their own.

    Agents are powerful when actions are reversible, inexpensive and well logged. For everything else, they should prepare and a person should decide.

    Engineering8 min

  3. Do you need an AI agent, or a workflow with a few model calls?

    Most business processes are better served by a fixed workflow with model calls at specific steps. Agents earn their place where the path can’t be known in advance, and they cost more to test, run and explain.

    Engineering6 min

  4. Prompt injection: which defenses actually hold.

    You can’t filter your way out of prompt injection. Limit what an injected instruction could achieve instead: separate instructions from data, scope permissions, validate outputs and put a person in front of consequential actions.

    Engineering6 min

  5. What are AI guardrails? A map of where each one lives.

    “Guardrails” covers very different controls: checks before the model sees a request, limits on what the system can do, validation of what it produces, and people at approval points. The most dependable are ordinary software, not instructions in a prompt.

    Engineering5 min

  6. How to design a review screen people can decide from in one look.

    The review screen decides whether AI saves time. Show the recommendation, the evidence behind it, what is uncertain and what will happen on approval, so a person can decide without redoing the work.

    Engineering6 min

  7. Getting reliable structured data out of a language model.

    Treat model output like any untrusted input: define a strict schema, constrain generation, validate every field against business rules, and send what fails to a person instead of retrying forever.

    Engineering6 min

  8. How to test an LLM application, layer by layer.

    Test the deterministic parts like ordinary software, the language parts with evaluation sets, and the joins between them. A single end-to-end accuracy number hides where things actually break.

    Engineering6 min

  9. Prompts are code: how to version, review and test them.

    A prompt change is a behavior change. Keep prompts in version control, review them like code, tie each version to evaluation results, and never edit them live in a dashboard.

    Engineering5 min

  10. How to design the tools an AI agent is allowed to use.

    An agent is only as safe and reliable as its tools. Make each tool narrow, clearly described, scoped to minimum permissions, safe to call twice and informative when it fails.

    Engineering6 min

Want a second opinion on your first workflow?