The problem: your data goes to the cloud
When you use ChatGPT to summarize meeting notes, when you send a client transcript to NotebookLM, when you paste CRM data into Gemini to build a report, your data leaves your machine. It goes to remote servers, often hosted in the United States, and gets processed by models whose terms of use change regularly.
Using these tools is fine in itself, and they're extremely useful. The trouble starts when the data you share contains personally identifiable information: client names, email addresses, phone numbers, company names, sometimes even social security numbers or bank details.
In France and across Europe, GDPR requires you to protect this data. In practice, many marketing, sales and HR teams send sensitive information to AI tools every day without thinking about it. The usual reason is simple: nobody has shown them another way to work.
The good news is that it comes down to method, and the method is simpler than you might think.
What GDPR says about using AI
GDPR rests on several core principles that apply directly to the use of AI tools:
- Data minimization: collect and share only the data strictly necessary for the intended processing. If you want an AI to summarize a meeting, it doesn't need to know who the participants are.
- Purpose limitation: data collected for a specific purpose must not be reused for another one without a legal basis. Sending customer data to an AI tool "to see what it does with it" isn't a valid purpose.
- Consent and transparency: the people whose data is processed must know how it's used. Do your customers know that their conversations are analyzed by an AI?
So the real question is how to use AI properly. The answer starts with a simple distinction.
Some internal data raises no sharing issue at all: your templates, your guidelines, your internal processes, your anonymous briefs. Anything that contains no information that could identify a natural person can be sent freely to any AI tool.
As soon as a document contains a name, an email address, a phone number or any other identifiable personal data, though, you need to anonymize it before you share it.
The solution: anonymize before you share
Microsoft developed Presidio, an open-source tool built for exactly this need. Presidio analyzes a text, automatically detects personal data and replaces it with generic placeholders.
In practice, a text like this:
Following our call with Marie Dupont (marie.dupont@entreprise.fr,
06 12 34 56 78), here is the recap for the Acme Corp project.
Becomes:
Following our call with <PERSON> (<EMAIL_ADDRESS>,
<PHONE_NUMBER>), here is the recap for the <ORGANIZATION> project.
The key point: all the processing happens locally, on your machine. Nothing leaves your computer. Presidio sends no data to any remote server. That's exactly the behavior you want from an anonymization tool.
You can then send the anonymized text safely to ChatGPT, NotebookLM, Claude or any other cloud AI tool. The AI works with the context it needs and never sees the personal data.
| AI tool | Processing | Personal data | Recommendation |
|---|---|---|---|
| ChatGPT (OpenAI) | Cloud (USA) | Anonymize before sending | Use with Presidio |
| Gemini (Google) | Cloud (USA) | Anonymize before sending | Use with Presidio |
| NotebookLM (Google) | Cloud (USA) | Anonymize before sending | Use with Presidio |
| Claude (Anthropic, via API) | Cloud (USA) | Anonymize before sending | Use with Presidio |
| Claude Code (local) | Local + API | Code processed locally | Suitable for sensitive data |
| Presidio (Microsoft) | 100% local | Nothing sent outside | Anonymization tool |
| Ollama / local LLMs | 100% local | Nothing sent outside | Ideal for sensitive data |
Tutorial: install and use Presidio step by step
This tutorial builds on Pando Studio's work on anonymization with Presidio. I adapted and expanded it so you can use it right away.
Prerequisite: install Python
Presidio works with Python 3.9 or later. If Python isn't installed on your machine, download it from python.org. On a Mac, you can also install it through Homebrew with brew install python.
Install the Presidio libraries
Open your terminal (or the Command Prompt on Windows) and type:
pip install presidio-analyzer presidio-anonymizer
Download the French spaCy model
Presidio uses spaCy language models to detect named entities (names, places, organizations). For French:
python -m spacy download fr_core_news_lg
Prepare your file
Put the transcript or document you want to anonymize in a plain text file. For example, create a transcription.txt file containing the raw text of your meeting or client report.
The anonymization script
Create an anonymiser.py file with the following content:
from presidio_analyzer import AnalyzerEngine
from presidio_analyzer.nlp_engine import NlpEngineProvider
from presidio_anonymizer import AnonymizerEngine
# NLP engine configuration for French
configuration = {
"nlp_engine_name": "spacy",
"models": [{"lang_code": "fr", "model_name": "fr_core_news_lg"}],
}
provider = NlpEngineProvider(nlp_configuration=configuration)
nlp_engine = provider.create_engine()
# Initialize the analysis and anonymization engines
analyzer = AnalyzerEngine(nlp_engine=nlp_engine, supported_languages=["fr"])
anonymizer = AnonymizerEngine()
# Read the source file
with open("transcription.txt", "r", encoding="utf-8") as f:
texte_original = f.read()
# Analysis: detect personal data
resultats = analyzer.analyze(
text=texte_original,
language="fr",
entities=[
"PERSON",
"EMAIL_ADDRESS",
"PHONE_NUMBER",
"LOCATION",
"ORGANIZATION",
"IBAN_CODE",
"CREDIT_CARD",
"IP_ADDRESS",
"URL",
],
)
# Anonymization: replace with placeholders
texte_anonymise = anonymizer.anonymize(
text=texte_original,
analyzer_results=resultats,
)
# Save the result
with open("transcription_anonymisee.txt", "w", encoding="utf-8") as f:
f.write(texte_anonymise.text)
print("Anonymization complete.")
print(f"File saved: transcription_anonymisee.txt")
print(f"Entities detected: {len(resultats)}")
Run the script
In your terminal, run:
python anonymiser.py
The script reads your transcription.txt file, detects all the personal data, replaces it with placeholders and saves the result in transcription_anonymisee.txt. You can then open that file, check the result and send it to the AI tool of your choice.
GDPR compliance is a framework that makes you think about what you share, with whom and why. In practice, it adds 5 minutes per document. Those 5 minutes protect your customers and your company.
Tip: if installing Python or Presidio gives you trouble, you can ask Claude Code to walk you through it step by step. It runs locally on your machine and can execute the installation commands for you.
Need support with GDPR and AI?
The AI Marketing Cockpit: your brand encoded, your tools connected, 54 ready-to-use skills.
Discover the AI Marketing CockpitEveryday best practices
Anonymizing a file now and then is good. Building real discipline across the team is better. These are the habits I recommend to my clients.
Set a clear policy. Which data can go into which tool? Create a simple document that lists the data categories (internal with no personal data, containing personal data, confidential) and the tools allowed for each one. It should fit on one page and make sense to everyone.
Create a standard anonymization workflow. Don't let each employee improvise. Define a clear process: the source file goes through Presidio, someone checks the result by hand, then it goes to the AI tool. The workflow can be as simple as a script on the desktop that you double-click.
Check the result by hand. No anonymization tool is 100% reliable. Presidio can miss an unusual proper name or a phone number in a nonstandard format. Take 30 seconds to reread the anonymized text before you send it. The habit sets in quickly.
Never ask an AI to anonymize your data. Doing so means sending it the data before anonymization. It's a common trap: "ChatGPT, can you anonymize this text?" means ChatGPT has received the text with all the personal data in it. Anonymization must always happen locally, before anything is sent.
Favor tools that process data locally when you can. For development and technical tasks, Claude Code runs locally on your machine. For language models, solutions like Ollama let you run LLMs entirely on your computer. The cloud remains essential for the most powerful models, but local processing is enough for many everyday tasks.
For full support with bringing GDPR-compliant AI into your organization, you can start with the free AI marketing diagnostic, which assesses your maturity and identifies your priorities.
Conclusion
GDPR is a framework for AI, and like any framework it gives structure to the practice. The companies that build data protection in from the start are the ones that adopt AI with the most confidence and for the long run.
With a tool like Presidio and a few good habits, anonymizing your data before handing it to an AI takes a few minutes. That's a small investment compared with the risk of a data breach or a GDPR violation, with fines of up to 4% of annual revenue.
Anonymization deals with the data before it leaves. The remaining question is the regulatory framework this processing falls under, which changes on August 2, 2026, when the European AI regulation becomes fully applicable. I detailed the concrete obligations for a marketing team in what GDPR and the AI Act require before you connect AI to your customer data.
If you're new to AI, I also recommend reading how to structure your first AI project at work and the 5 AI tools I recommend in 2026. And to set up a secure local development environment, see my Claude Code guide.