veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesGovernance & risk

Where does your data go when you use a language model?

Data travels further than people think: prompts, retrieved passages, outputs, logs, caches, search indexes and test sets. Map each hop, check the contract and settings for each, and design so the most sensitive data never makes the trip.

veridive6 min read

Ask where the data goes when someone uses an AI assistant, and the usual answer is “to the model provider”. That is one stop on a longer journey. A single question can leave copies in the provider’s logs, your application logs, a tracing tool, a search index, a cache and, later, a test set someone builds from real cases.

For LLM data privacy, the useful unit is the hop, not the provider. Map each hop, check the contract and the settings for each, and design so the most sensitive data never makes the trip.

What leaves your systems when someone asks a question?

Follow one illustrative customer email about a damaged delivery through an assistant that drafts replies.

  1. The email arrives in the ticket system with the customer’s name, address, order number and a description of the damage.
  2. The assistant looks up the order in the order system: items, amounts, delivery status.
  3. It searches the returns policy and retrieves three passages. If the index was also built from past tickets, it may retrieve other customers’ words.
  4. It builds a prompt from its instructions, the email, the order fields and the passages, and sends it to the provider’s endpoint in a region.
  5. The provider returns a draft. Depending on the agreement, it may keep the prompt and the draft for a period, for example for abuse monitoring.
  6. Your application logs the full prompt and draft, and a tracing tool, often another company’s service, receives a copy for debugging.
  7. The agent edits and sends the reply. Later, someone copies the case into an evaluation set.

None of this is unusual. It just needs to be decided rather than discovered. Every flow has the same hops:

HopWhat it holdsWhat to check
PromptInstructions, the question, record fields, retrieved textOnly the fields the task needs
Retrieved contextPassages from documents or past ticketsPermissions applied at retrieval
OutputThe draft, often repeating personal dataWhere it is stored, who sees it
Provider logsInputs and outputs, for a periodRetention, region, staff access
Your logs and tracesFull prompts and outputsMasking, retention, access
Search indexText chunks and their embeddingsDeletion reaching the index
Evaluation setsCopies of real casesMasking, access, retention
Fine-tuning dataExamples built into a modelWhether it can ever be removed

What does a model provider typically keep, and for how long?

It depends on the product, the feature and the agreement, and it changes over time, so treat any general answer, including this one, as a reason to check. Patterns you will meet include inputs and outputs held for a limited period for abuse monitoring, reduced or zero retention for some business customers, and files, histories and fine-tuning data kept until you delete them.

Ask each provider, in writing:

  • Retention: what is kept for each feature you use, for how long, and can you shorten it?
  • Training use: is any of your data used to train or improve a model, by default or ever?
  • Region: where are requests processed and stored, including for abuse review and support?
  • Sub-processors: who else handles the data, where, and how will you hear about changes?
  • Deletion: how do you delete stored data, how long does it take, and how is it confirmed?

What stays inside your systems but still needs protection?

The hops you control are often the weakest, because they were built for debugging, not for privacy.

  • Logs and traces often hold full prompts and outputs, readable by more people than the ticket system itself. Mask what you can, keep references instead of copies where you can, and set a retention period.
  • Search indexes hold text and embeddings, the numeric representations used for search. Embeddings are derived from the text, so protect them like it, and make sure deletion reaches them.
  • Caches that store answers can serve one person’s answer to another if they aren’t scoped to the user, the kind of gap that access control for AI assistants should close.
  • Evaluation sets are copies of real cases, often in spreadsheets. Give them an owner, access roles and masked versions where possible.
  • Fine-tuning data is the hardest to take back: a model trained on someone’s data can’t simply forget them.

How do consumer and enterprise terms differ?

A personal account on a chat assistant is a contract between the provider and an individual. The terms are written for consumers, the defaults on retention and training use are the provider’s to set and change, the history belongs to the individual, and the company has no admin controls, audit logs or offboarding.

An enterprise agreement makes the company the customer, with data processing terms, admin controls, retention settings, commitments on training use and a list of sub-processors. The model underneath may be the same. What differs is who controls the data and what was agreed about it, which is why a personal account is a different risk even when the product looks identical.

Which design choices keep sensitive data at home?

The most private data is the data you never send.

  • Send less. Pass only the fields a task needs. A delivery question doesn’t need the customer’s payment history.
  • Mask before sending. Replace names, account numbers and similar identifiers with placeholders and restore them afterwards. Our note on masking personal data covers where masking fails.
  • Process regionally. Choose provider regions that match where the data must stay, and get the region into the agreement.
  • Host in your own cloud when you need your own controls around the model.
  • Run open-weight models on-premises for the most sensitive flows, such as health, HR or legal files, accepting the cost of running them and the need to test quality on your cases.
  • Route by sensitivity. Let a rule send routine cases to a hosted model and sensitive ones to the on-premises model, or to a person.

What should you verify before signing?

  • The agreement itself, not the product page, covers training use, retention per feature, processing regions, sub-processors and deletion.
  • The settings in your account match the agreement, and features you don’t need, such as stored histories, are off.
  • Deletion works: delete a test record and confirm it has left every hop, yours included.
  • Your own logs, traces and test sets have owners, retention periods and access roles.
  • Your DPO or counsel has reviewed transfers and legal basis; our note on KVKK and GDPR questions lists them.
  • The system never shows anyone data they could not open themselves, one of our guardrails.

Data and AI foundations covers the deployment side: keeping each flow where its data needs to live.

This note is general information, not legal advice.

Sources

  1. Personal Data Protection Law (Law No. 6698), English text Personal Data Protection Authority (KVKK) www.kvkk.gov.tr/Icerik/6649/Personal-Data-Protection-Law
  2. Regulation (EU) 2016/679 (General Data Protection Regulation), official text EUR-Lex eur-lex.europa.eu/eli/reg/2016/679/oj
  3. Guide on Generative Artificial Intelligence and the Protection of Personal Data (in 15 Questions), in Turkish Personal Data Protection Authority (KVKK) www.kvkk.gov.tr/Icerik/8547/uretken-yapay-zeka-ve-kisisel-verilerin-korunmasi-rehberi-15-soruda

Ask an assistant about this note

Governance & riskData protectionSecurity

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

Do AI providers use your data to train their models?

It depends on the product and the agreement, not on AI as such. Consumer accounts and business agreements can have different defaults, and defaults can change. Read the agreement for the specific product and features you use, check what it says about training, retention and processing region, and confirm that your account settings match. For sensitive data, prefer a design that sends less.

How long do LLM providers keep prompts and outputs?

It varies by provider, feature and agreement. Some keep inputs and outputs for a limited period for abuse monitoring, some offer reduced or zero retention to eligible business customers, and stored files, histories and fine-tuning data usually stay until deleted. Ask for retention per feature in writing, and remember that your own logs, traces and indexes need retention rules too.