To stay GDPR compliant with ChatGPT, NotebookLM or Claude, I anonymize documents before sending them. Presidio, Microsoft's open-source tool, detects names, emails, phone numbers and organizations in a text and replaces them with generic placeholders, entirely on your machine: nothing leaves it. Installation takes 2 Python commands, with the French spaCy model. The AI keeps the useful context without ever seeing the personal data.

The problem: your data goes to the cloud

When you use ChatGPT to summarize meeting notes, when you send a client transcript to NotebookLM, when you paste CRM data into Gemini to build a report, your data leaves your machine. It goes to remote servers, often hosted in the United States, and gets processed by models whose terms of use change regularly.

Using these tools is fine in itself, and they're extremely useful. The trouble starts when the data you share contains personally identifiable information: client names, email addresses, phone numbers, company names, sometimes even social security numbers or bank details.

In France and across Europe, GDPR requires you to protect this data. In practice, many marketing, sales and HR teams send sensitive information to AI tools every day without thinking about it. The usual reason is simple: nobody has shown them another way to work.

The good news is that it comes down to method, and the method is simpler than you might think.

What GDPR says about using AI

GDPR rests on several core principles that apply directly to the use of AI tools:

  • Data minimization: collect and share only the data strictly necessary for the intended processing. If you want an AI to summarize a meeting, it doesn't need to know who the participants are.
  • Purpose limitation: data collected for a specific purpose must not be reused for another one without a legal basis. Sending customer data to an AI tool "to see what it does with it" isn't a valid purpose.
  • Consent and transparency: the people whose data is processed must know how it's used. Do your customers know that their conversations are analyzed by an AI?

So the real question is how to use AI properly. The answer starts with a simple distinction.

Some internal data raises no sharing issue at all: your templates, your guidelines, your internal processes, your anonymous briefs. Anything that contains no information that could identify a natural person can be sent freely to any AI tool.

As soon as a document contains a name, an email address, a phone number or any other identifiable personal data, though, you need to anonymize it before you share it.

The solution: anonymize before you share

Microsoft developed Presidio, an open-source tool built for exactly this need. Presidio analyzes a text, automatically detects personal data and replaces it with generic placeholders.

In practice, a text like this:

Following our call with Marie Dupont (marie.dupont@entreprise.fr,
06 12 34 56 78), here is the recap for the Acme Corp project.

Becomes:

Following our call with <PERSON> (<EMAIL_ADDRESS>,
<PHONE_NUMBER>), here is the recap for the <ORGANIZATION> project.

The key point: all the processing happens locally, on your machine. Nothing leaves your computer. Presidio sends no data to any remote server. That's exactly the behavior you want from an anonymization tool.

You can then send the anonymized text safely to ChatGPT, NotebookLM, Claude or any other cloud AI tool. The AI works with the context it needs and never sees the personal data.

AI tool Processing Personal data Recommendation
ChatGPT (OpenAI) Cloud (USA) Anonymize before sending Use with Presidio
Gemini (Google) Cloud (USA) Anonymize before sending Use with Presidio
NotebookLM (Google) Cloud (USA) Anonymize before sending Use with Presidio
Claude (Anthropic, via API) Cloud (USA) Anonymize before sending Use with Presidio
Claude Code (local) Local + API Code processed locally Suitable for sensitive data
Presidio (Microsoft) 100% local Nothing sent outside Anonymization tool
Ollama / local LLMs 100% local Nothing sent outside Ideal for sensitive data

Tutorial: install and use Presidio step by step

This tutorial builds on Pando Studio's work on anonymization with Presidio. I adapted and expanded it so you can use it right away.

Prerequisite: install Python

Presidio works with Python 3.9 or later. If Python isn't installed on your machine, download it from python.org. On a Mac, you can also install it through Homebrew with brew install python.

Install the Presidio libraries

Open your terminal (or the Command Prompt on Windows) and type:

Terminal
pip install presidio-analyzer presidio-anonymizer

Download the French spaCy model

Presidio uses spaCy language models to detect named entities (names, places, organizations). For French:

Terminal
python -m spacy download fr_core_news_lg

Prepare your file

Put the transcript or document you want to anonymize in a plain text file. For example, create a transcription.txt file containing the raw text of your meeting or client report.

The anonymization script

Create an anonymiser.py file with the following content:

Python
from presidio_analyzer import AnalyzerEngine
from presidio_analyzer.nlp_engine import NlpEngineProvider
from presidio_anonymizer import AnonymizerEngine

# NLP engine configuration for French
configuration = {
    "nlp_engine_name": "spacy",
    "models": [{"lang_code": "fr", "model_name": "fr_core_news_lg"}],
}

provider = NlpEngineProvider(nlp_configuration=configuration)
nlp_engine = provider.create_engine()

# Initialize the analysis and anonymization engines
analyzer = AnalyzerEngine(nlp_engine=nlp_engine, supported_languages=["fr"])
anonymizer = AnonymizerEngine()

# Read the source file
with open("transcription.txt", "r", encoding="utf-8") as f:
    texte_original = f.read()

# Analysis: detect personal data
resultats = analyzer.analyze(
    text=texte_original,
    language="fr",
    entities=[
        "PERSON",
        "EMAIL_ADDRESS",
        "PHONE_NUMBER",
        "LOCATION",
        "ORGANIZATION",
        "IBAN_CODE",
        "CREDIT_CARD",
        "IP_ADDRESS",
        "URL",
    ],
)

# Anonymization: replace with placeholders
texte_anonymise = anonymizer.anonymize(
    text=texte_original,
    analyzer_results=resultats,
)

# Save the result
with open("transcription_anonymisee.txt", "w", encoding="utf-8") as f:
    f.write(texte_anonymise.text)

print("Anonymization complete.")
print(f"File saved: transcription_anonymisee.txt")
print(f"Entities detected: {len(resultats)}")

Run the script

In your terminal, run:

Terminal
python anonymiser.py

The script reads your transcription.txt file, detects all the personal data, replaces it with placeholders and saves the result in transcription_anonymisee.txt. You can then open that file, check the result and send it to the AI tool of your choice.

GDPR compliance is a framework that makes you think about what you share, with whom and why. In practice, it adds 5 minutes per document. Those 5 minutes protect your customers and your company.

Tip: if installing Python or Presidio gives you trouble, you can ask Claude Code to walk you through it step by step. It runs locally on your machine and can execute the installation commands for you.

Need support with GDPR and AI?

The AI Marketing Cockpit: your brand encoded, your tools connected, 54 ready-to-use skills.

Discover the AI Marketing Cockpit

Everyday best practices

Anonymizing a file now and then is good. Building real discipline across the team is better. These are the habits I recommend to my clients.

Set a clear policy. Which data can go into which tool? Create a simple document that lists the data categories (internal with no personal data, containing personal data, confidential) and the tools allowed for each one. It should fit on one page and make sense to everyone.

Create a standard anonymization workflow. Don't let each employee improvise. Define a clear process: the source file goes through Presidio, someone checks the result by hand, then it goes to the AI tool. The workflow can be as simple as a script on the desktop that you double-click.

Check the result by hand. No anonymization tool is 100% reliable. Presidio can miss an unusual proper name or a phone number in a nonstandard format. Take 30 seconds to reread the anonymized text before you send it. The habit sets in quickly.

Never ask an AI to anonymize your data. Doing so means sending it the data before anonymization. It's a common trap: "ChatGPT, can you anonymize this text?" means ChatGPT has received the text with all the personal data in it. Anonymization must always happen locally, before anything is sent.

Favor tools that process data locally when you can. For development and technical tasks, Claude Code runs locally on your machine. For language models, solutions like Ollama let you run LLMs entirely on your computer. The cloud remains essential for the most powerful models, but local processing is enough for many everyday tasks.

For full support with bringing GDPR-compliant AI into your organization, you can start with the free AI marketing diagnostic, which assesses your maturity and identifies your priorities.

Conclusion

GDPR is a framework for AI, and like any framework it gives structure to the practice. The companies that build data protection in from the start are the ones that adopt AI with the most confidence and for the long run.

With a tool like Presidio and a few good habits, anonymizing your data before handing it to an AI takes a few minutes. That's a small investment compared with the risk of a data breach or a GDPR violation, with fines of up to 4% of annual revenue.

Anonymization deals with the data before it leaves. The remaining question is the regulatory framework this processing falls under, which changes on August 2, 2026, when the European AI regulation becomes fully applicable. I detailed the concrete obligations for a marketing team in what GDPR and the AI Act require before you connect AI to your customer data.

If you're new to AI, I also recommend reading how to structure your first AI project at work and the 5 AI tools I recommend in 2026. And to set up a secure local development environment, see my Claude Code guide.

Frequently asked questions

Yes, as long as you comply with GDPR. That means not sending identifiable personal data (names, emails, phone numbers) to remote servers without explicit consent. The simplest and safest solution is to anonymize the data with a local tool like Presidio before sharing it with a cloud AI tool.
Yes. Presidio uses spaCy language models, including fr_core_news_lg for French. Detection of names, places and organizations works well in French. For emails, phone numbers and IBANs, detection relies on regular expressions that work the same in any language.
For audio recordings, first transcribe the file into text with a local tool like OpenAI's Whisper, then run Presidio on the resulting transcript. The anonymized text can then be sent to any AI tool for analysis or a summary. What matters is that both the initial transcription and the anonymization happen locally.