Before reading a page, an AI crawler reads a file that says what it may do with it. That file exists on most websites. This survey looks at what it answers, and above all at who wrote it.
85.3 % of the websites surveyed publish a robots.txt. That file has a history: it was written for search engines, at a time when the question was which pages would be indexed. Half of these files hold a single group of rules, and that group addresses everyone at once.
AI crawlers read the same file and look for their own name in it. They do not find it on 97 % of websites. This is not a refusal, and it is not a considered permission: it is a question that has not been asked yet, on a file ten years older than the question.
Each crawler has its own group in a robots.txt, so a website answers OpenAI, Google and Common Crawl separately. The table reads those answers one by one.
| Crawler | Operated by | What it does with the page | Websites turning it away |
|---|---|---|---|
| CCBot | Common Crawl | Trains a model | 2.0 % |
| GPTBot | OpenAI | Trains a model | 2.0 % |
| Bytespider | ByteDance | Trains a model | 1.8 % |
| ClaudeBot | Anthropic | Trains a model | 1.8 % |
| Google-Extended | Trains a model | 1.8 % | |
| Meta-ExternalAgent | Meta | Trains a model | 1.7 % |
| Applebot-Extended | Apple | Trains a model | 1.6 % |
| ChatGPT-User | OpenAI | Fetches the page when someone asks a question | 0.3 % |
| PerplexityBot | Perplexity | Feeds an answer engine | 0.1 % |
| OAI-SearchBot | OpenAI | Feeds an answer engine | 0.1 % |
The gap is clean and it stands on its own: crawlers that train models are turned away by around 2 %, those that fetch a page to answer a question twenty times less. The French parc accepts being cited and refuses being learned from.
Refusing training protects content a company paid to produce. That is a defensible decision, and 2.2 % of websites have taken it explicitly.
Closing the door on answer engines produces another one: your website stops being a citable source at the very moment your customers put their questions to an AI rather than to a search engine. 3.1 % of the websites surveyed are in that position without having chosen it, through a rule written for every crawler long before these ones existed.
Both decisions are taken in the same file, in three lines. You still have to be able to write it.
A robots.txt is served by the machine that serves the website. When that website lives on a closed platform, the platform writes the file, applies it to all its customers at once, and updates it when it decides to.
We reread 2,500 of the files in the survey to see which ones carried the signature of a platform template. The table below gives what each one answers on behalf of its customers.
| Platform | Websites in the draw | Rule groups |
|---|---|---|
| Shopify | 57 | 2 |
| WordPress.com | 13 | 1 |
| Squarespace | 11 | 30 |
| Wix | 3 | 0 |
| Webflow | 1 | 0 |
| File with no platform signature | 2,410 | 1 |
Squarespace serves its customers a file that names about thirty crawlers, including every one in the table above, and turns none of them away. Shopify serves one that names none. In both cases the answer is the same for thousands of websites, and the owner of one of those websites neither wrote that answer nor has the means to change it.
That is what separates a website you own from a website you occupy. The question of AI crawlers arrived in two years; the next one will arrive just as fast, and it will be settled in the same place, in a file someone has to be able to write for you.
Serenity is a website created, hosted, secured and kept up to date by Simafri, which has been creating, hosting and maintaining company websites since 2002. The domain name is registered in the client name, and a Simafri Suite account is included.
Sampling frame: the 4,590,553 active .fr domains in the open data file published by Afnic Afnic, fichier des domaines .fr actifs. A panel of 15,000 domains is drawn from it, ordered by the sha256 fingerprint of each name, which lets anyone rebuild the same panel from the published seed. The panel stays the same from one survey to the next.
Reading chain: on 3 September 2026, 10,301 of those domains served a page. For each of them, the robots.txt file was requested in the same conversation. 98 % returned a settled policy, that is a readable file or an answer stating there is none, which allows everything. The figures cover those.
What counts as a refusal: a crawler is turned away when its own group of rules forbids it the root of the website. A group addressing every crawler at once is recorded separately, because it is about crawling in general and says nothing about AI. A website whose platform names crawlers inside that general group is therefore not counted as having taken a decision.
What was not read: nothing requiring access to the website, its administration or its hosting. One request for the page, one for the file.
Scope: the panel draws from every active .fr domain, companies, nonprofits and individuals alike, with no qualification by company number, unlike our two other surveys. The figures therefore describe the French parc as a whole. The platform reading covers 2,500 websites drawn among those serving a file.
The parc figures come from a public index that publishes its method, its dated panel and its weekly series, read on 3 September 2026. The series is served in JSON, crawler by crawler, and replays.
See the index and its method (Stileex)
The panel is read every week, which will say whether the parc moves on this question and in which direction.
Simafri has been the technical ally of your company since 2002: we create your website, host it, maintain it and keep it up to date, with your domain name registered in your name. You decide who may learn from your pages and who may cite them, we make the move. Write to us, we will look at your case together.
Simafri
Let's talk about your project
Tell us what you need in a few words: we will get back to you quickly.
Prefer email? Write to us at support@simafri.com.
The form is not displaying? Write to us directly:
support@simafri.com