# Getting reliable structured data out of a language model.

> In this field note, veridive explains how to get reliable structured data out of a language model by treating its output as untrusted input. It covers schema design with fixed lists and explicit nulls, constrained generation, layered validation against business rules and master data, retry limits with a human queue, and field-level accuracy on an evaluation set.

Treat model output like any untrusted input: define a strict schema, constrain generation, validate every field against business rules, and send what fails to a person instead of retrying forever.

## Key takeaways

- Treat model output as untrusted input: a strict schema, constrained generation and validation in code before anything reaches another system.
- Design schemas with fixed lists, an explicit null for “not in the document”, and units and formats spelled out.
- Validate types, business rules, cross-field sums and master-data lookups; retry a little, then send failures to a person.
- Measure accuracy field by field on the evaluation set, separating wrong values, missed values and values invented where none exist.

The extraction demo returned clean JSON for every invoice the team tried. Then a supplier who writes amounts as 3.480,00 sent its first invoice. The amount came back as text, a parser read it as three point four eight, and the ERP accepted the draft. Nothing crashed. The model had read the invoice correctly; the contract between the model and the ERP had never said what an amount looks like.

Structured output works when you treat model output like any untrusted input: define a strict schema, constrain generation, validate every field against business rules, and send what fails to a person instead of retrying forever.

## Why does free text break integrations?

Downstream systems need exact values: a date, a decimal amount, a code from a list. Free text such as “net 30 from receipt”, “mid-March” or “see attached” has to be interpreted by someone, and an integration can’t do that.

Language models are fluent, not exact. They produce plausible values in plausible formats, and plausible is the dangerous kind of wrong: a well-formed date that is the delivery date instead of the invoice date passes every format check. Output also varies between runs and between models, so a format that held in testing can drift after a model update.

Before extracting anything, check whether the data already exists as data. Structured e-invoices, for example, need parsing rather than reading; see [where e-Fatura ends and AI begins](https://veridive.com/insights/e-fatura-invoice-automation/).

## How do you design the schema?

With the people who key the documents by hand, one schema per document type:

- **Fixed lists instead of free text.** Document type, currency, tax rate and unit of measure come from lists that match your master data.
- **An explicit null for “not in the document”.** Every optional field allows null, and the instructions say to use it rather than guess. A separate flag marks values that are present but unreadable.
- **Units and formats spelled out.** Dates as year-month-day, amounts as plain decimals with a dot and no thousands separator, currency in its own field, quantities with their unit.
- **One-line field descriptions.** Where the value usually appears and which look-alike to ignore: “invoice date, not delivery date or due date”.
- **Evidence.** For key fields, the page and the text the value came from, so reviewers and validators can check it.

Keep schemas flat and small. Deep nesting and dozens of optional fields invite omissions.

## How do you constrain and validate the output?

Many model APIs and serving stacks can constrain generation to a JSON schema. Use that where available: every response then parses and has the right types. That settles the shape, not the truth, so validation follows in layers of ordinary code:

1. **Types and formats.** Required fields present, dates valid, amounts numeric, codes from the lists.
2. **Business rules.** The due date falls after the invoice date, amounts are positive except on credit notes, and the tax rate is one your business uses.
3. **Cross-field checks.** Line items add up to the net total, net plus tax equals gross within a rounding tolerance, and quantity times unit price equals each line amount.
4. **Lookups against master data.** The supplier’s tax number exists in the vendor master, the purchase order is open, the bank account matches the vendor record.
5. **Evidence checks.** Identifiers such as the invoice number actually appear in the document text.

Each layer is cheap, deterministic and testable, which is more than can be said for another instruction in the prompt.

## What happens when validation fails?

Retry once or twice, feeding back the specific error: “the line items don’t add up to the net total; re-read the totals section”. Retries fix format slips and careless reads. They don’t fix a document that is ambiguous, damaged or new.

So set a hard limit, then send the document to a person with the extracted fields filled in and the failed checks highlighted. Without a limit, retries burn money, and sooner or later one output passes validation by chance.

> Retrying until the output passes validation is a way of manufacturing a plausible wrong answer.

The person’s corrected record becomes an evaluation case. Track the failure rate by supplier and document type: a sudden rise usually means a new layout or an upstream change, not a worse model.

## How do you handle unknown and missing values?

Distinguish three states for every field: present and read, not in the document (null), and present but unreadable or ambiguous (flagged, with a reason). Never let the model fill a gap from general knowledge (“payment terms are usually thirty days”) or from a neighboring field.

Decide per field whether missing is acceptable. A purchase-order number may be routinely absent for some suppliers; a missing total should always stop the document. Defaults belong in code, applied visibly, never invented by the model. If documents arrive as scans or photos, many errors start before the model reads anything; [reading Turkish documents with AI](https://veridive.com/insights/turkish-ocr-documents/) covers that stage.

## How do you measure extraction quality?

Field by field, on an evaluation set of real documents with reference values. For each field, count exact matches after normalization, values missed, and values invented where the document has none, which are the most dangerous. Break results down by document type, supplier group, language and scan quality, and track the straight-through rate: documents that passed every check and needed no correction.

An illustrative example, with invented numbers: fifty supplier invoices from an evaluation set.

| Field | Exact matches | Main error | Fix |
|---|---|---|---|
| Supplier tax number | 49 of 50 | A digit misread on a scan | Checksum and vendor-master lookup |
| Invoice date | 46 of 50 | Delivery date taken instead | Sharper field description, two examples |
| Net total | 50 of 50 | None | None |
| Line items | 42 of 50 | Lines split across pages | Page-joining step, sum check |
| Purchase-order number | 45 of 50 | Invented when absent | Explicit null rule, open-order lookup |

A single document-level score would call this system nearly ready. The field table shows exactly where to work, and which errors validation would have caught anyway.

## Schema first, prompt second

Pick one document type and write its schema with the people who key it by hand: every field, its format, where it appears and what “missing” means. Write the validation rules before the prompt. Extraction with validation and a review queue is a typical first project in [ERP and enterprise workflows](https://veridive.com/solutions/erp-enterprise-workflows/), and a standard part of the [custom AI software](https://veridive.com/services/custom-ai-software/) we build.

## Frequently asked questions

### How do you get reliable JSON output from an LLM?

Define a strict schema with fixed lists, explicit nulls and spelled-out formats, and use the model’s structured-output or schema mode where available so every response parses. Then validate every field in code: types, business rules, cross-field sums and lookups against master data. Retry once or twice with the error message, and send anything that still fails to a person.

### How do you measure LLM extraction accuracy?

Measure it field by field on an evaluation set of real documents with reference values, not by eyeballing whole documents. For each field, count exact matches, missed values and values invented where the document has none, and break results down by document type, supplier and scan quality. Also track how many documents pass every check with no correction.

### What should an LLM do when a value is missing from a document?

Return an explicit null, never a guess. The schema should allow null for every optional field and distinguish “not in the document” from “present but unreadable”, which gets a flag and a reason. Defaults belong in code, applied visibly, and each field needs a rule for whether a missing value is acceptable or sends the document to a person.
