# Clause extraction for legal teams: what AI can and can’t check.

> In this field note, veridive sets out a playbook for contract clause extraction with AI: which tasks suit it, why to start from a clause list and standard positions written by the legal team, how to show cited extractions, how to handle Turkish and English contracts, how to measure accuracy by clause type, and where lawyers always decide.

AI is good at finding, extracting and comparing clauses against your playbook across many contracts; it is not the lawyer. Start with a clause list and a playbook your legal team wrote, cite every extraction, and keep judgment on risk with people.

## Key takeaways

- AI suits finding, extracting and comparing clauses across many contracts; judgment on risk, meaning and negotiation stays with lawyers.
- Start from a clause list and a playbook of standard positions written by the legal team, not from a general prompt.
- Every extraction quotes the clause and cites its page, and “not found” is shown explicitly, because a missing clause is a finding.
- Compare the Turkish and English versions of bilingual contracts, and measure accuracy per clause type on a set built with lawyers.

Finance has a simple question: which of our supplier contracts renew automatically, and how much notice do we have to give? The answer is in the contracts, several hundred of them, some scanned, some bilingual, some on the supplier’s own paper. A lawyer can answer it. It will take weeks, and it will be asked again.

This is where AI earns its place in a legal team: finding, extracting and comparing clauses across many contracts, against positions the legal team already holds. It is not the lawyer. It doesn’t decide whether a deviation is acceptable, what an ambiguous clause means or what to renegotiate. The playbook below starts from a clause list, cites every extraction, measures accuracy clause by clause, and leaves judgment on risk with people. This note is general information, not legal advice.

## Which contract tasks suit AI?

Tasks with a clear question and a checkable answer suit it. Tasks that need facts outside the document, or a view on risk, don’t.

| Suits AI | Stays with lawyers |
|---|---|
| Finding termination, liability, renewal, governing law and payment clauses | Deciding whether a deviation is acceptable |
| Extracting values: caps, notice periods, renewal terms | Interpreting ambiguous or conflicting wording |
| Comparing each clause with your standard position | Judging enforceability |
| Building and updating a register of key terms | Negotiation strategy and anything sent to a counterparty |
| Triaging counterparty paper so the riskiest clauses are read first | Advice to the business |

The usual first tasks are the five clause types in the table’s first row. They come up again and again in portfolio reviews, and the legal team already knows its standard position on each.

## Why start with a clause list and a playbook?

Because “review this contract” has no finish line. A clause list does. For each clause type, the legal team writes three things: what to extract, the standard position, and what counts as a deviation. For limitation of liability, for example, that means whether a cap exists, what it is based on, whom it protects and which carve-outs apply; the standard position in your own words; and flags for no cap, a one-sided cap or a cap below standard.

The playbook also handles what makes contracts hard. Definitions live elsewhere, so a cap based on “Fees” needs the definition of Fees pulled in too. Clauses cross-refer (“subject to clause 14”), and amendments, annexes and side letters change terms after signing, so the system reads the whole contract family and shows which document each value came from. Which document prevails is not its question to answer.

## How should extractions be shown to lawyers?

Every extraction quotes the clause text and cites its page and clause number, with a link to the exact place in the document. Lawyers check the quote, not a paraphrase. Next to it sit the extracted value, the comparison with the standard position and the reason for any flag, the same principle as [answers with receipts](https://veridive.com/insights/answers-with-receipts/).

“Not found” gets the same care. A contract without a liability cap is itself a finding, and the lawyer needs to know whether the clause is absent or the system missed it, so show where it looked. Deviations go to the top of the review queue, and every correction a lawyer makes becomes a new test case.

Take an illustrative example: a legal team reviews its supplier contracts before a renewal round. The standard position is a mutual cap tied to the fees under the contract, with carve-outs for confidentiality and data protection. The system extracts the liability clause from each contract, quotes it with its page and compares it with the standard. Most match. A handful are flagged: one caps only the supplier’s liability, one has no cap at all, one caps at a fixed amount far below the contract’s value, and two bilingual contracts state different caps in their Turkish and English versions. Each flag carries the quoted clause, the page and the reason. A lawyer works through the flagged contracts first and samples the matching ones, to confirm nothing was missed.

> The system finds the clause. The lawyer decides what it means.

## How do you handle Turkish and English contracts?

Expect three forms: Turkish only, English only, and bilingual contracts with the two languages side by side. The clause list needs names and examples in both languages, such as “fesih” for termination, “cezai şart” for a penalty clause and “mücbir sebep” for force majeure, so retrieval finds both. Side-by-side layouts need each column read separately, or the text interleaves line by line.

In bilingual contracts, extract each clause from both versions and compare them. Differences happen: the Turkish text says “otuz gün” (thirty days) where the English says thirty business days, or an amount appears in only one version. Show both quotes side by side, with the contract’s language clause, which often says which version prevails. What the difference means is for the lawyer. Answering questions across the two languages follows the same rules as any [bilingual assistant](https://veridive.com/insights/bilingual-ai-assistants/).

## How do you measure accuracy clause by clause?

Build the evaluation set with the legal team from real contracts: your templates, counterparty paper, old and new, scanned and bilingual. Lawyers mark where each clause is and what the correct extraction and flag would be. Then measure, per clause type and per language:

- **Found:** was the clause located, and nothing else mistaken for it?
- **Extracted:** is the value right, such as the cap basis or the notice period?
- **Cited:** does the quote match the clause and page?
- **Flagged:** does the deviation flag match the lawyer’s view?

Watch missed deviations above all. A false alarm costs a lawyer a minute; a missed uncapped liability costs far more. Agree a threshold for each clause type, because an average across types hides the weak one.

## Where must a lawyer always decide?

Whether a deviation is acceptable, what ambiguous wording means, which document in a contract family prevails, what to renegotiate, and anything that goes to a counterparty or to the business as advice. The roles stay clear:

- **Legal owner:** writes and updates the playbook, sets the thresholds and signs off before the system is used on live work.
- **Reviewing lawyers:** work the flagged queue, correct extractions and decide every deviation.
- **Contract managers:** assemble contract families, check “not found” cases and keep the register current.
- **The system:** finds, extracts, compares, cites and flags. It never decides, and the log records who approved what.

Where contracts are processed, and who can see them, are also questions for counsel and your data protection officer before anything is uploaded.

## Start narrow

Pick one clause type and one contract set, such as liability caps in supplier contracts. Write the playbook entry, have lawyers mark a few dozen contracts, and measure how often the system finds, extracts and flags correctly. Clause extraction and comparison are part of [knowledge and document intelligence](https://veridive.com/solutions/knowledge-document-intelligence/), and the review queue lawyers work in is typical [custom AI software](https://veridive.com/services/custom-ai-software/).

## Frequently asked questions

### Can AI review contracts instead of lawyers?

No. AI can find, extract and compare clauses across many contracts against a playbook your legal team wrote, and flag deviations with the clause quoted and its page cited. Deciding whether a deviation is acceptable, what an ambiguous clause means or what to renegotiate stays with lawyers. The system prepares the review; it doesn’t replace the reviewer.

### Which contract clauses should AI extract first?

Start with clauses that are frequent, well defined and commercially important: termination, limitation of liability, renewal and notice periods, governing law and jurisdiction, and payment terms. For each one, the legal team writes what to extract, the standard position and what counts as a deviation. Measure accuracy clause type by clause type before adding more.

### How should AI handle bilingual Turkish and English contracts?

Extract each clause from both language versions and compare them, because amounts, notice periods and obligations sometimes differ between the two. Show both quotes side by side, together with the contract’s language clause, which often says which version prevails. What a discrepancy means, and what to do about it, is a decision for the legal team.
