AI crawlers and French websites: who decides the answer

Before reading a page, an AI crawler reads a file that says what it may do with it. That file exists on most websites. This survey looks at what it answers, and above all at who wrote it.

Published on 3 September 2026 10,301 .fr websites, surveyed on 3 September 2026

97 %of websites name no AI crawler in their robots.txt
2.2 %turn at least one away, by naming it
3.1 %close their root to every crawler, search engines included

The file is already there, it is simply speaking to someone else

85.3 % of the websites surveyed publish a robots.txt. That file has a history: it was written for search engines, at a time when the question was which pages would be indexed. Half of these files hold a single group of rules, and that group addresses everyone at once.

AI crawlers read the same file and look for their own name in it. They do not find it on 97 % of websites. This is not a refusal, and it is not a considered permission: it is a question that has not been asked yet, on a file ten years older than the question.

What the parc turns away, crawler by crawler

Each crawler has its own group in a robots.txt, so a website answers OpenAI, Google and Common Crawl separately. The table reads those answers one by one.

Crawler Operated by What it does with the page Websites turning it away
CCBot Common Crawl Trains a model 2.0 %
GPTBot OpenAI Trains a model 2.0 %
Bytespider ByteDance Trains a model 1.8 %
ClaudeBot Anthropic Trains a model 1.8 %
Google-Extended Google Trains a model 1.8 %
Meta-ExternalAgent Meta Trains a model 1.7 %
Applebot-Extended Apple Trains a model 1.6 %
ChatGPT-User OpenAI Fetches the page when someone asks a question 0.3 %
PerplexityBot Perplexity Feeds an answer engine 0.1 %
OAI-SearchBot OpenAI Feeds an answer engine 0.1 %

The gap is clean and it stands on its own: crawlers that train models are turned away by around 2 %, those that fetch a page to answer a question twenty times less. The French parc accepts being cited and refuses being learned from.

The two decisions are not the same

Refusing training protects content a company paid to produce. That is a defensible decision, and 2.2 % of websites have taken it explicitly.

Closing the door on answer engines produces another one: your website stops being a citable source at the very moment your customers put their questions to an AI rather than to a search engine. 3.1 % of the websites surveyed are in that position without having chosen it, through a rule written for every crawler long before these ones existed.

Both decisions are taken in the same file, in three lines. You still have to be able to write it.

On a closed platform, that file is not yours

A robots.txt is served by the machine that serves the website. When that website lives on a closed platform, the platform writes the file, applies it to all its customers at once, and updates it when it decides to.

We reread 2,500 of the files in the survey to see which ones carried the signature of a platform template. The table below gives what each one answers on behalf of its customers.

Platform Websites in the draw Rule groups
Shopify 57 2
WordPress.com 13 1
Squarespace 11 30
Wix 3 0
Webflow 1 0
File with no platform signature 2,410 1

Squarespace serves its customers a file that names about thirty crawlers, including every one in the table above, and turns none of them away. Shopify serves one that names none. In both cases the answer is the same for thousands of websites, and the owner of one of those websites neither wrote that answer nor has the means to change it.

That is what separates a website you own from a website you occupy. The question of AI crawlers arrived in two years; the next one will arrive just as fast, and it will be settled in the same place, in a file someone has to be able to write for you.

Serenity, in plain terms

Serenity is a website created, hosted, secured and kept up to date by Simafri, which has been creating, hosting and maintaining company websites since 2002. The domain name is registered in the client name, and a Simafri Suite account is included.

See the offer and prices

Method

Sampling frame: the 4,590,553 active .fr domains in the open data file published by Afnic Afnic, fichier des domaines .fr actifs. A panel of 15,000 domains is drawn from it, ordered by the sha256 fingerprint of each name, which lets anyone rebuild the same panel from the published seed. The panel stays the same from one survey to the next.

Reading chain: on 3 September 2026, 10,301 of those domains served a page. For each of them, the robots.txt file was requested in the same conversation. 98 % returned a settled policy, that is a readable file or an answer stating there is none, which allows everything. The figures cover those.

What counts as a refusal: a crawler is turned away when its own group of rules forbids it the root of the website. A group addressing every crawler at once is recorded separately, because it is about crawling in general and says nothing about AI. A website whose platform names crawlers inside that general group is therefore not counted as having taken a decision.

What was not read: nothing requiring access to the website, its administration or its hosting. One request for the page, one for the file.

Scope: the panel draws from every active .fr domain, companies, nonprofits and individuals alike, with no qualification by company number, unlike our two other surveys. The figures therefore describe the French parc as a whole. The platform reading covers 2,500 websites drawn among those serving a file.

Source and reuse

The parc figures come from a public index that publishes its method, its dated panel and its weekly series, read on 3 September 2026. The series is served in JSON, crawler by crawler, and replays.

See the index and its method (Stileex)

JSON API

The panel is read every week, which will say whether the parc moves on this question and in which direction.

Take back the say over what your website answers

Simafri has been the technical ally of your company since 2002: we create your website, host it, maintain it and keep it up to date, with your domain name registered in your name. You decide who may learn from your pages and who may cite them, we make the move. Write to us, we will look at your case together.