Two phrases do a lot of work on the privacy pages of AI CV tools, and neither one means what it is being asked to mean. The first is "European AI". The second is "anonymised".
Where the AI runs, and who built it
"Our AI runs in Europe" sounds like one claim. It is two, and they come apart more often than the phrase suggests.
The first question is where the inference happens and who operates the endpoint. Your CV text goes somewhere, that somewhere sits in a jurisdiction, and a company controls it. This is the question data protection actually cares about.
The second question is who trained the model. Not who serves it. Who built it, on whose data, with whose idea of what a good answer looks like.
For most of the last few years these two collapsed into one, because the good models were only available from the labs that made them. Calling GPT meant calling OpenAI. The answer to both questions was the same answer.
What open weights changed
A growing number of strong models are published under licences that let anyone download and run them. Meta's Llama family. Alibaba's Qwen. DeepSeek. OpenAI's own gpt-oss, released under Apache 2.0 in August 2025.
Once weights are public, hosting is separate from origin. A French provider can serve an American model from a Paris datacentre, and the lab that trained it never sees a single request. That is not a loophole. It is a real and significant privacy outcome: no data leaves Europe, no American company processes your CV, no transfer mechanism is needed.
So a European CV tool can truthfully write "all our models run in France, at European providers", and be describing a system where the model doing most of the work was trained in California.
Both halves of that sentence are true. Neither is the whole picture.
Why origin still matters, and where it does not
Let us be precise about this, because the case is often overstated.
For data protection, the second question barely matters. If the weights run on a European provider's hardware and nothing is sent to the originating lab, your data is as safe as the host makes it. GDPR asks where personal data goes and who processes it. It does not ask who trained the matrix multiplication. Anyone telling you an EU-hosted American model is a privacy problem is selling something.
Three other things do follow from origin, though.
A model carries the conventions it was trained on. This is not abstract for a CV tool. A model trained overwhelmingly on English-language text has absorbed an American idea of what a resume is: no photo, no date of birth, no marital status, one page, achievements front-loaded. A French CV routinely carries a photo and a "Profil recherché" section. A Spanish one often carries both. Ask a US-trained model to improve a French CV and its instinct is to make it more American, quietly, in a hundred small edits nobody flagged as a choice.
You hold what was released, and nothing more. Apache 2.0 weights you have downloaded cannot be taken back, so continuity for a given version is genuinely safe. What you have no claim on is the next version. A lab that stops publishing owes you nothing, and the roadmap of a tool built on somebody else's open releases is a roadmap somebody else controls.
And there is the industrial question, which is a matter of opinion rather than fact, so take it as ours. If Europe's answer to AI sovereignty is that we run American models on European servers, then Europe is a hosting layer. The money goes to GPUs and datacentres, and the part that compounds, the research, the training runs, the talent, happens elsewhere. Sovereign infrastructure without sovereign models is a landlord's position.
The other word this category uses loosely
"European AI" is not the only phrase carrying more than it can hold. "Anonymised" is the other, and it fails the same way: a term with a precise legal meaning, used as though it were a synonym for something weaker.
Under GDPR, anonymous data is data that no longer relates to an identifiable person, by any means reasonably likely to be used, by anyone. It falls out of the regulation entirely. That is a very high bar. Pseudonymised data, where identifiers are replaced but the link can still be made, remains personal data and stays fully in scope. The two are one word apart in marketing copy and a world apart in law.
Now apply that to a CV. Several tools strip your name, email, phone number and address before sending anything to a model, which is a genuinely good design and more than most of the category does. But look at what is left: employers, job titles, dates, city, university, certifications. That is not an anonymous document. An employment history is close to a fingerprint, and the standard test asks whether you can single someone out, link records, or infer who they are. A dated list of who somebody worked for fails all three comfortably.
Stripping identifiers is still worth doing. It is real data minimisation, it shrinks what a provider could log, and it limits what shows up in an error trace. It is simply not anonymisation, and calling it that promises a legal status the technique does not deliver.
We were doing this too, and fixed it the day we noticed
Our public stats page said page views were anonymised. The visitor fingerprint was a SHA256 of the IP address, the browser and the date, unsalted. IPv4 is about four billion addresses and browser strings carry very little entropy, so anyone holding that table could have worked backwards to the addresses. CNIL and the EDPB both treat hashing an identifier as pseudonymisation.
We salted it, which closes that path for anyone who only has the data, and then changed the word to pseudonymous in all fourteen languages, because we hold the salt and could still reverse it ourselves. The honest version is a weaker-sounding sentence and a true one.
What we found across 32 CV tools
We went looking for a privacy policy on every AI CV tool we could find, in four languages, and recorded what each one actually says rather than what it markets. Thirty-two tools. Not all of them have one.
On where the model runs:
- 15 name an American model provider and state no European region at all
- 13 do not name a model provider anywhere in what they publish
- 1 names an American provider and does state a European region
- 3 use European providers
On who trained the weights:
- 16 use American labs
- 13 do not disclose enough to say
- 2 use a European lab
- 1 uses both
The two lists are not the same list, and the difference is the entire point of this article. One tool sits in the "European provider" column and the "American lab" column at the same time, which is a perfectly coherent position and is invisible if you only have one column.
That tool, incidentally, publishes the most detailed technical disclosure in the whole sample. It names every model with its version slug, explains its pipeline step by step, and volunteers the limits of its own anonymisation, which nobody else does. It is not being evasive. The vocabulary simply does not exist yet, so the honest version of the claim has nowhere to go.
Thirteen tools disclose nothing at all. That remains the headline finding, and no amount of nuance about weights origin changes it.
What we do, and what it costs
Job-Agent uses Mistral models, from Mistral, served on Scaleway in Paris. European on both questions. Of the 32 tools we read, two are.
We should be straight about the price of that. The open frontier is largely American right now, and the largest European models are not the largest models. Choosing European weights means accepting a smaller model than the best available, and we spend real engineering effort making up the difference: the analysis pipeline splits work across many small focused calls rather than asking one enormous model to do everything, and a chain of deterministic checks catches the things a smaller model gets wrong. That is a design constraint we chose, not a free lunch we found.
We think it is worth it. You might not, and if what you want is the strongest possible model and you are comfortable with where it comes from, that is a legitimate position and there are good tools for it.
A simple rule
For both phrases the question is the same: who would have to be unable to do something for the word to be earned?
For "anonymised": if the company could still work out who you are, it is pseudonymous, whatever the page calls it. Usually the person who would have to be unable is the one writing the sentence.
For "European AI": ask whether the inference is European, the model is, or both. All three answers are defensible. Only one is currently being asked.
Ask.
