Under the Hood:
What Happens to the Data Provided to a Language Model

Reading time: 4 min

An analysis of the technical functioning of generative models and its implications, including in relation to the GDPR.

The debate over the use of AI with company data tends to polarize into two positions: those who view generative models as ordinary software that can be used without any special precautions, and those who view any input of data into a service like ChatGPT as a violation that should be prohibited. Both positions overlook how the technology actually works. This is where we should begin: most questions about data use are answered on a technical level, and legal assessments stem from those answers.

 

1) A MODEL DOES NOT REMEMBER: IT GENERATES

A language model is not a database. It does not contain records or files from which to retrieve answers: it is a large-scale function that, given an input text, calculates the probability of each possible subsequent text fragment (token), selects one, and repeats the process iteratively.

This has two consequences.

First: Model training and evaluation reward attempts to provide answers more than they do the admission of uncertainty. Even in the absence of information, the model still produces an output with the same degree of apparent certainty. So-called “hallucinations” can be reduced but not eliminated through an update: they stem from the very property that makes the model useful, and the model does not signal when it is generating unfounded content. No provider guarantees accuracy: the terms of use specify that the output may be false, incomplete, or misleading.

Second: the model is unable to explain the reasons behind a response. It can produce a plausible justification, but this is merely a further generation of the response, not a description of its internal mechanisms. On this point, technology intersects with the law: when a decision based solely on automated processing produces legal or similarly significant effects on an individual, the GDPR (Art. 22, in conjunction with Arts. 13–15) grants that individual the right to meaningful information about the logic used.

From an operational standpoint, this means that human oversight remains necessary: the model can be the first reviewer, but never the last.

 

2) WHERE DOES THE DATA END UP?

There are three ways to use an LLM, each with very different implications.

  • Cloud: This is the most common method. The data entered is transmitted to the provider's servers.
  • Self-hosted: The model runs on the organization's own infrastructure. It ensures maximum control, but requires significant costs and expertise.
  • Enterprise: a contractually governed intermediate solution in which the location of the data and the purposes of its processing can be negotiated.

For a European user, there are three key issues.

The first concerns data localization. The transfer of personal data outside the EU is not prohibited, but it requires one of the mechanisms provided for by the GDPR for such transfers, such as an adequacy decision or standard contractual clauses; the simplest and most robust solution is EU data residency, that is, the processing and storage of data on servers located within the EU. When a provider processes personal data on behalf of the user, a Data Processing Addendum (DPA) is also required. The paid version of a service does not necessarily store data in Europe, and not all major providers offer direct EU residency.

The second issue concerns the use of data for training. In consumer versions—whether free or with an individual subscription—most major providers now use data for training by default: OpenAI, Google, Mistral, and, starting in the fall of 2025, Anthropic as well (which has asked each user to make a choice, with the training option set to “on” by default). An opt-out option is available, which must be enabled by the user. In paid business, Enterprise, and API versions, however, training is disabled by default, and this exclusion is stipulated in the contract; free API tiers may allow it, subject to regional exceptions (Google, for example, applies the paid terms to European users even on the free tier). Terms and conditions are subject to change and should be reviewed periodically in the providers’ documentation.

The third concern relates to access to logs. First and foremost, the provider’s internal teams can access them. In the case of U.S. providers, U.S. authorities can as well: the CLOUD Act requires providers subject to U.S. jurisdiction to hand over, upon order of the authorities, data under their control even when it is stored outside the United States. A server in Frankfurt therefore does not offer the same guarantees if the provider is American. For Chinese providers, such as DeepSeek or Qwen, the 2017 Intelligence Law requires cooperation with national intelligence agencies: for sensitive data, their cloud-based versions should be avoided.

Added to this is the risk of data extraction by third parties: through prompt injection, it is possible to trick a model into revealing confidential information, and no vendor is immune to this.

Unless otherwise specified—such as with temporary chats or zero-data-retention agreements—the conversation is saved in the chat history. This is the default behavior, not the worst-case scenario.

 

3) THE TECHNICAL LIMIT OF DELETION

Deleting a conversation is simple. Removing information acquired during training is not: it does not reside in a specific record, but is distributed across the model’s weights. You can verify that the model does not reproduce it in response to certain requests, but there are known techniques for making it emerge in response to different requests. Deletion removes the logs, not what the model has learned.

The right to erasure is therefore subject to a limitation that depends on the system’s architecture, not on the provider’s discretion. According to the European Data Protection Board (Opinion 28/2024), a model trained on personal data is not anonymous by definition: if the data can be extracted, the GDPR continues to apply.

The GDPR also applies to the client company, which is the data controller for customer and employee data. The supplier typically acts as a data processor under the DPA and becomes a data controller if it uses that data to train its own models. In the absence of personal data, the GDPR does not apply, but protections regarding trade secrets remain in effect.

 

4) THE USER REMAINS LIABLE

Responsibility for the output does not transfer to the tool. If a proposal sent to a client contains fabricated data or confidential information belonging to another client, the model’s developer is not liable; rather, liability rests with the person who sent it or the company for which they work. The same applies to defamatory or discriminatory output: liability falls on whoever publishes or commissions it. However, the landscape is evolving: for products placed on the market after December 9, 2026—the deadline for transposition in Member States—the new European Product Liability Directive (2024/2853) includes software and AI systems among products, with strict liability of the manufacturer for damages to natural persons.

 

THEREFORE, CRITERIA FOR USE

European legislation establishes the principles, but it is not enough on its own: it must be supplemented with company guidelines that specify which tools to use and with what data. You don’t need a comprehensive policy to get started. The consumer cloud is suitable for drafts, rewrites, brainstorming, and simple automations—not for personal data or trade secrets. For confidential documents, you need an Enterprise version—preferably with EU data residency—or a self-hosted solution.

Before entering data, perform five checks:

  1. Destination: public cloud, an environment governed by a DPA, or your own infrastructure? Are the servers located in the EU, and what is the provider’s country of origin?
  2. Training: Does the provider use the data for training—by default or unless the user opts out?
  3. Access: Who else can read the logs?
  4. Traceability: If it becomes necessary in the future to demonstrate where the data ended up, or to have it deleted, is there documentation that would allow for this?
  5. Impact: What would be the consequences of releasing this text?

That last question alone is enough to resolve most cases. A lead’s name and an email’s subject line are not the same as a complete CRM export containing ten years’ worth of negotiations: treating them the same way—whether strictly or loosely—is the main mistake.

Generative AI is neither dangerous nor harmless. It is a technology that operates in a specific way, and that way of operating determines what kind of data should be entrusted to it.

Alan Perotti

AI & Digital Sales

The New Sales Edge: Closing Deals at the Speed of AI
10/26/06 | 2 months | ONLINE
Learn More

AI Super Power

10x your output. 0x your stress.
10/27/26 | 4 months | H-FARM
Learn More