# Masking personal data before it reaches a language model.

> In this field note, veridive explains how to mask personal data before text reaches a language model. It covers what to mask, how masking differs from pseudonymization and anonymization, detecting Turkish identity numbers, IBANs, phone numbers and names locally, reversible placeholders, where masking fails quietly, and how to test detection on your own data formats.

Masking reduces what a model provider sees, but it is not anonymization and it fails quietly. Detect locally, use reversible placeholders where the answer needs them, test detection on your own data formats, and don’t treat masking as the whole privacy plan.

## Key takeaways

- Masking reduces what a model provider sees, but it is not anonymization; whether masked data is still personal data is for counsel.
- Detect personal data locally, before the call, with format rules and checksums, name recognition and the customer’s own known details.
- Use reversible placeholders such as CUSTOMER_1 when the answer must mention the person, and fill them back in locally.
- Masking fails quietly through context, free-text notes, scans and attachments, so test recall on your own data formats.

A complaint arrives by email with the customer’s full name, a mobile number and the IBAN they want the refund sent to. The team wants a language model to draft the reply. Before the text leaves your systems, ask whether the name, the number and the IBAN need to go with it. Usually they don’t.

Masking, which replaces personal data with placeholders before the model call, is one of the most useful privacy controls around a language model and one of the easiest to over-trust. It is not anonymization, and when it fails, it fails quietly. So detect locally, use reversible placeholders where the answer needs them, test detection on your own data formats, and don’t treat masking as the whole privacy plan.

## What should be masked, and why?

Whatever the model doesn’t need for the task: a reply to a complaint needs the complaint, not the customer’s identity number. The usual candidates:

- **Direct identifiers:** names, identity and passport numbers, email addresses, phone numbers.
- **Financial details:** IBANs, card and account numbers.
- **Addresses,** down to the apartment number.
- **Combinations that identify:** a birth date, a job title in a small team, a rare event.
- **Sensitive details in free text,** such as health or family information in agent notes. Whether such data may be processed at all is a question for your DPO.

Data that never reaches the provider can’t be logged, retained or exposed there. [Where your data goes when you use a language model](https://veridive.com/insights/llm-data-privacy/) covers the rest of that path.

## What is the difference between masking, pseudonymization and anonymization?

In technical terms:

| Technique | What it does | Can the person be recovered? |
|---|---|---|
| Masking | Removes a value or replaces it with a generic tag such as [NAME] | Not from the masked text |
| Pseudonymization | Replaces each value with a consistent placeholder such as CUSTOMER_1, with the mapping kept separately | Yes, with the mapping |
| Anonymization | Transforms data so nobody can reasonably re-identify the person, even with other information | Not by design |

Most masking for language models is really pseudonymization, because the reply needs the real values back. Data protection law defines these terms its own way, and whether your masked data still counts as personal data is a question for your DPO or counsel, not a technical one; see [the questions to ask them](https://veridive.com/insights/kvkk-gdpr-ai-questions/).

## How do you detect personal data in Turkish and English text?

Locally, before the call: sending text to an external model to find its personal data defeats the purpose. Use layers that catch different things:

1. **Format rules with checksums.** Turkish identity numbers have eleven digits, never start with zero, and end in two check digits derived from the first nine, so a validator rejects most random eleven-digit numbers, such as order numbers. IBANs carry their own checksum, and Turkish ones start with TR and run to twenty-six characters. Phone numbers appear with +90, 0 or no prefix, and with spaces, dots or dashes in different places.
2. **Name and address recognition,** with a named-entity model (one that tags names, places and organizations) running on your own infrastructure. Test it on names with Turkish characters, names typed in capitals or without Turkish letters (“SUKRU” for “Şükrü”), and names that are everyday words, such as Deniz, Umut or Barış. Turkish addresses have their own markers: Mah., Cad., Sok., No: and Daire.
3. **Known values.** The name, email and phone number in the linked CRM record are the most reliable signals; match them exactly.
4. **Labels in the text,** such as “Ad Soyad:”, “IBAN:” or “Name:”, which say what follows.

**Done when:** every entity type has a detector with a measured recall on your own samples. **What goes wrong:** a detector tuned on English text that misses Turkish names.

## When do you need reversible placeholders?

When the output must mention the person or their details. Replace each value with a typed, numbered placeholder such as CUSTOMER_1, PHONE_1 or IBAN_1, keep the mapping on your side for a short time, and put the real values back into the output locally. If it doesn’t, as in routing or classification, mask one way and keep no mapping at all.

Back to the illustrative complaint from the opening. Local detection finds the name through the sender’s CRM record, the mobile number by pattern and the IBAN by pattern plus checksum. The model receives: “My name is CUSTOMER_1. You charged me twice for order ORDER_1. Call me on PHONE_1 or refund the extra charge to IBAN_1.” It drafts an apology to CUSTOMER_1 confirming that the duplicate charge is being checked. Locally, the placeholders are filled back in, the draft is checked for leftover or invented placeholders, and an agent reviews and sends it. The refund itself is a separate step, approved by a person in the payment system. The provider saw the complaint, not the name, the number or the IBAN.

Every placeholder in the output must exist in the mapping, so an invented PHONE_2 or a mangled “Customer 1” stops the draft. Turkish suffixes also follow the real name, not the placeholder: “Ayşe’nin” but “Burak’ın”. Ask the model to avoid suffixes on placeholders, or correct them after substitution.

## Where does masking fail?

> A missed IBAN produces no error. That is why masking needs tests, not trust.

- **Context re-identifies.** “The warehouse manager at our smallest depot who raised the safety complaint” names nobody and identifies someone.
- **Free-text notes** hold health details, family situations and phone numbers typed in odd formats.
- **Scans, images and attachments** carry personal data outside the text that detection reads. OCR them first, or don’t send them; [reading Turkish documents with AI](https://veridive.com/insights/turkish-ocr-documents/) covers the reading side.
- **Over-masking** removes product names, dates or amounts the task needs, so measure answer quality with masking switched on.
- **The mapping leaks.** A mapping table written to application logs undoes the masking. Keep it out of logs, encrypted, and delete it on schedule.

Masking is one control among several, alongside data minimization, provider terms on retention, access control and logging. Our [guardrails](https://veridive.com/approach/#guardrails) start from who can access what and where data is processed.

## How do you test it?

On labeled examples, like any detector:

1. **Build a test set in your own formats:** emails, chats, agent notes, forms and OCR output, real or realistic synthetic, handled in a secure environment. Label every item of personal data.
2. **Measure recall per entity type.** Missed items are the failures that matter.
3. **Measure precision too,** because over-masking hurts quality.
4. **Include the hard cases:** capitals, missing Turkish letters, names that are common words, IBANs split by spaces, mixed Turkish and English.
5. **Re-run on every change** to detectors, formats or channels.

**Done when:** recall per entity type meets a threshold agreed with the DPO, and the test runs automatically.

## Test detection first

Label the personal data in a hundred real messages from one workflow, in a secure environment, and run your detector over them; the misses show where to begin. Mapping processing locations and access roles, and deploying where your data needs to live, are part of our [data and AI foundations](https://veridive.com/services/data-ai-foundations/) work; the legal questions stay with your DPO.

This note is general information, not legal advice.

## Frequently asked questions

### What is PII redaction for LLMs?

PII redaction for LLMs means detecting personal data, such as names, identity numbers, phone numbers and bank details, and removing or replacing it before text is sent to a language model. Replacing values with placeholders lets the model still do its job, and reversible placeholders let your own system fill the real values back in afterwards, without the model provider seeing them.

### Is masked data still personal data under KVKK or GDPR?

That is a legal question for your DPO or counsel, not a technical one. Masking reduces what a model provider sees, but whether data that can be linked back to a person, through a mapping table or through its context, still counts as personal data depends on the law and the details. Describe your masking design to them and let them decide.

## Sources

1. Personal Data Protection Law (Law No. 6698), English text. Personal Data Protection Authority (KVKK). https://www.kvkk.gov.tr/Icerik/6649/Personal-Data-Protection-Law
2. Regulation (EU) 2016/679 (General Data Protection Regulation), official text. EUR-Lex. https://eur-lex.europa.eu/eli/reg/2016/679/oj
3. Guide on Generative Artificial Intelligence and the Protection of Personal Data (in 15 Questions), in Turkish. Personal Data Protection Authority (KVKK). https://www.kvkk.gov.tr/Icerik/8547/uretken-yapay-zeka-ve-kisisel-verilerin-korunmasi-rehberi-15-soruda
