Masking personal data before it reaches a language model.
Masking reduces what a model provider sees, but it is not anonymization and it fails quietly. Detect locally, use reversible placeholders where the answer needs them, test detection on your own data formats, and don’t treat masking as the whole privacy plan.
veridive6 min read
A complaint arrives by email with the customer’s full name, a mobile number and the IBAN they want the refund sent to. The team wants a language model to draft the reply. Before the text leaves your systems, ask whether the name, the number and the IBAN need to go with it. Usually they don’t.
Masking, which replaces personal data with placeholders before the model call, is one of the most useful privacy controls around a language model and one of the easiest to over-trust. It is not anonymization, and when it fails, it fails quietly. So detect locally, use reversible placeholders where the answer needs them, test detection on your own data formats, and don’t treat masking as the whole privacy plan.
What should be masked, and why?
Whatever the model doesn’t need for the task: a reply to a complaint needs the complaint, not the customer’s identity number. The usual candidates:
- Direct identifiers: names, identity and passport numbers, email addresses, phone numbers.
- Financial details: IBANs, card and account numbers.
- Addresses, down to the apartment number.
- Combinations that identify: a birth date, a job title in a small team, a rare event.
- Sensitive details in free text, such as health or family information in agent notes. Whether such data may be processed at all is a question for your DPO.
Data that never reaches the provider can’t be logged, retained or exposed there. Where your data goes when you use a language model covers the rest of that path.
What is the difference between masking, pseudonymization and anonymization?
In technical terms:
| Technique | What it does | Can the person be recovered? |
|---|---|---|
| Masking | Removes a value or replaces it with a generic tag such as [NAME] | Not from the masked text |
| Pseudonymization | Replaces each value with a consistent placeholder such as CUSTOMER_1, with the mapping kept separately | Yes, with the mapping |
| Anonymization | Transforms data so nobody can reasonably re-identify the person, even with other information | Not by design |
Most masking for language models is really pseudonymization, because the reply needs the real values back. Data protection law defines these terms its own way, and whether your masked data still counts as personal data is a question for your DPO or counsel, not a technical one; see the questions to ask them.
How do you detect personal data in Turkish and English text?
Locally, before the call: sending text to an external model to find its personal data defeats the purpose. Use layers that catch different things:
- Format rules with checksums. Turkish identity numbers have eleven digits, never start with zero, and end in two check digits derived from the first nine, so a validator rejects most random eleven-digit numbers, such as order numbers. IBANs carry their own checksum, and Turkish ones start with TR and run to twenty-six characters. Phone numbers appear with +90, 0 or no prefix, and with spaces, dots or dashes in different places.
- Name and address recognition, with a named-entity model (one that tags names, places and organizations) running on your own infrastructure. Test it on names with Turkish characters, names typed in capitals or without Turkish letters (“SUKRU” for “Şükrü”), and names that are everyday words, such as Deniz, Umut or Barış. Turkish addresses have their own markers: Mah., Cad., Sok., No: and Daire.
- Known values. The name, email and phone number in the linked CRM record are the most reliable signals; match them exactly.
- Labels in the text, such as “Ad Soyad:”, “IBAN:” or “Name:”, which say what follows.
Done when: every entity type has a detector with a measured recall on your own samples. What goes wrong: a detector tuned on English text that misses Turkish names.
When do you need reversible placeholders?
When the output must mention the person or their details. Replace each value with a typed, numbered placeholder such as CUSTOMER_1, PHONE_1 or IBAN_1, keep the mapping on your side for a short time, and put the real values back into the output locally. If it doesn’t, as in routing or classification, mask one way and keep no mapping at all.
Back to the illustrative complaint from the opening. Local detection finds the name through the sender’s CRM record, the mobile number by pattern and the IBAN by pattern plus checksum. The model receives: “My name is CUSTOMER_1. You charged me twice for order ORDER_1. Call me on PHONE_1 or refund the extra charge to IBAN_1.” It drafts an apology to CUSTOMER_1 confirming that the duplicate charge is being checked. Locally, the placeholders are filled back in, the draft is checked for leftover or invented placeholders, and an agent reviews and sends it. The refund itself is a separate step, approved by a person in the payment system. The provider saw the complaint, not the name, the number or the IBAN.
Every placeholder in the output must exist in the mapping, so an invented PHONE_2 or a mangled “Customer 1” stops the draft. Turkish suffixes also follow the real name, not the placeholder: “Ayşe’nin” but “Burak’ın”. Ask the model to avoid suffixes on placeholders, or correct them after substitution.
Where does masking fail?
A missed IBAN produces no error. That is why masking needs tests, not trust.
- Context re-identifies. “The warehouse manager at our smallest depot who raised the safety complaint” names nobody and identifies someone.
- Free-text notes hold health details, family situations and phone numbers typed in odd formats.
- Scans, images and attachments carry personal data outside the text that detection reads. OCR them first, or don’t send them; reading Turkish documents with AI covers the reading side.
- Over-masking removes product names, dates or amounts the task needs, so measure answer quality with masking switched on.
- The mapping leaks. A mapping table written to application logs undoes the masking. Keep it out of logs, encrypted, and delete it on schedule.
Masking is one control among several, alongside data minimization, provider terms on retention, access control and logging. Our guardrails start from who can access what and where data is processed.
How do you test it?
On labeled examples, like any detector:
- Build a test set in your own formats: emails, chats, agent notes, forms and OCR output, real or realistic synthetic, handled in a secure environment. Label every item of personal data.
- Measure recall per entity type. Missed items are the failures that matter.
- Measure precision too, because over-masking hurts quality.
- Include the hard cases: capitals, missing Turkish letters, names that are common words, IBANs split by spaces, mixed Turkish and English.
- Re-run on every change to detectors, formats or channels.
Done when: recall per entity type meets a threshold agreed with the DPO, and the test runs automatically.
Test detection first
Label the personal data in a hundred real messages from one workflow, in a secure environment, and run your detector over them; the misses show where to begin. Mapping processing locations and access roles, and deploying where your data needs to live, are part of our data and AI foundations work; the legal questions stay with your DPO.
This note is general information, not legal advice.
Sources
- Personal Data Protection Law (Law No. 6698), English text Personal Data Protection Authority (KVKK) www.kvkk.gov.tr/Icerik/6649/Personal-Data-Protection-Law
- Regulation (EU) 2016/679 (General Data Protection Regulation), official text EUR-Lex eur-lex.europa.eu/eli/reg/2016/679/oj
- Guide on Generative Artificial Intelligence and the Protection of Personal Data (in 15 Questions), in Turkish Personal Data Protection Authority (KVKK) www.kvkk.gov.tr/Icerik/8547/uretken-yapay-zeka-ve-kisisel-verilerin-korunmasi-rehberi-15-soruda
Ask an assistant about this note