An llms.txt is an invitation addressed to AI assistants: a short file at the root of a site, saying who you are and where to look. This reading counts the sites that have put it up, and reads what they tell those same crawlers in the file next door.
An llms.txt is a Markdown file placed at /llms.txt. It introduces the site in a few lines and points to the pages that matter, so an AI assistant can read them directly instead of piecing a whole site together. The format asks for one thing: a level-one heading at the top, and the rest is free.
It answers a new question. A search engine crawls a site page by page and ranks what it finds; an assistant reads fast, cites little, and needs to be told where the substance is. 7.9% of answering .fr sites have already taken a position on that question, two years after the format was proposed.
This is the trap in this reading, and it changes everything: many sites return their homepage for any unknown address. Asking such a site for /llms.txt yields a 200 and an HTML page, never an llms.txt. A count that stops at the HTTP status therefore adds the first two rows of the table together and publishes an adoption figure five times too high.
| What the site answers at /llms.txt | Share of sites |
|---|---|
| A file opening on a Markdown heading | 7.9 % |
| A 200, but something else (usually the homepage) | 36.6 % |
| An answer saying the file does not exist | 55.5 % |
The reading looks at the body of the response and counts as an llms.txt only what opens on a level-one heading, the single element the format requires.
A site that puts up an llms.txt invites AI assistants to come and read. The robots.txt file, three lines away, tells those same crawlers whether they may come in at all. Both are read the same morning, on the same domain.
Among sites publishing an llms.txt, 2.8% forbid their root to at least one AI crawler elsewhere. Across the whole .fr estate that share is 2.2%. Putting up the invitation, in other words, does not come with opening the door: the sites that invite turn crawlers away slightly more often than the rest.
None of that is contradictory in itself: declining to feed a model and wanting to be read well by an assistant are two separate decisions, and they can sit together perfectly well. What the figure shows is that these two files are rarely written together, even though they answer the same visitor.
The median size is 3,878 bytes, roughly four kilobytes: a map of the site, not a copy of it. That matches what the format asks for, and it is also what makes the file sustainable over time.
An llms.txt is worth what its freshness is worth. It names pages, and a page that changes address leaves a line in the file that leads nowhere. It is a file to revisit at every redesign, exactly like a sitemap.
Serenity is a website created, hosted, secured and kept up to date by Simafri, which has been building, hosting and maintaining business websites since 2002. The domain name is registered in the client name, and a Simafri Suite account is included.
Sampling frame: the 4,590,553 active .fr domains in the Afnic open data file Afnic, fichier des domaines .fr actifs. A panel of 15,000 domains is drawn from it, ordered by the sha256 hash of each name, which lets anyone rebuild the same panel from the published seed. The panel stays the same from one reading to the next, so a move in the series is a decision taken by the estate.
Reading chain: /llms.txt is requested once a week, in the same pass that already opens every domain in the panel. The figures cover the sites that answer the question, that is 98% of those serving a page: those serving a file, those serving something else, and those answering that the file does not exist.
What counts as an llms.txt: a body opening on a level-one Markdown heading. The published size is the median of the files served, in bytes. The robots.txt is read on the same domain the same morning, and an AI crawler counts as turned away when its own group of directives forbids it the root.
What was not read: nothing requiring access to the site, its administration or its hosting. One request for the page, one for each file.
Scope: the panel draws from all active .fr domains, businesses, associations and individuals alike, with no filtering by company registration. The figures therefore describe the French estate as a whole, and the share of sites turning away an AI crawler comes from the same weekly reading.
The figures come from a public index that publishes its method, its dated panel and its weekly series, read on 4 September 2026. The series is served as JSON and can be replayed.
See the index and its method (Stileex)
The panel is read every week, which will show how fast the estate puts this file up, and whether those who put it up also open their door.
Simafri has been the technical ally of businesses since 2002: we build your site, host it, maintain it and keep it up to date, with your domain name registered in your name. Your llms.txt and your robots.txt say what you decided, and are revisited when your pages move. Write to us and we will look at your case together.
Simafri
Let's talk about your project
Tell us what you need in a few words: we will get back to you quickly.
Prefer email? Write to us at support@simafri.com.
The form is not displaying? Write to us directly:
support@simafri.com