Research · 6 min read
Study: none of 35 large French companies blocks AI crawlers
We measured the robots.txt, llms.txt and structured data of 40 large listed French companies. Results, method and full data.
Published October 6, 2026 · by Alexandre
In short
Out of 40 large listed French companies measured on October 6, 2026, none of those with a readable robots.txt (35) blocks AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended…). Only 4 publish an llms.txt file, and 21 have structured data on their homepage. Large groups let AI read their sites, but few make it easy.
AI assistants like ChatGPT, Claude, Perplexity or Gemini read websites to answer. Companies can stop them in their robots.txt file. What do France's largest companies do? We measured, without sampling, the sites of 40 large listed groups on October 6, 2026.
Key findings
- None of the 35 companies with a readable robots.txt blocks AI crawlers, neither search crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot) nor training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot).
- 4 out of 40 publish an llms.txt file: Capgemini, Dassault Systèmes, Safran, Veolia.
- 21 out of 40 have structured data (JSON-LD) detected on their homepage.
- 5 sites return an anti-bot protection page instead of their robots.txt: they're not counted in the first figure.
What it means
Large groups made a clear choice: being readable by AI rather than shielding themselves from it. Blocking AI search crawlers would mean disappearing from their answers, at a time when a share of searches ends in ChatGPT or Google's AI Overviews.
However, very few go further: llms.txt remains rare (it isn't an official standard, see our article on llms.txt), and half of the homepages don't explain who the company is in structured data. That's exactly where a small business can get ahead.
What about a small business?
- Check that its robots.txt doesn't block AI search crawlers: the AI crawler checker does it in seconds.
- Describe the business in Organization or LocalBusiness structured data: the structured data generator creates the code.
- Publish an llms.txt, at low cost, with the llms.txt generator.
- Measure whether AI actually cites it, with the free audit.
Full data
| Company | AI crawlers (robots.txt) | llms.txt | Structured data |
|---|---|---|---|
| Accor | All allowed | No | No |
| Air Liquide | All allowed | No | Yes |
| Airbus | Not measurable | No | No |
| ArcelorMittal | All allowed | No | No |
| AXA | All allowed | No | No |
| BNP Paribas | All allowed | No | Yes |
| Bouygues | All allowed | No | Yes |
| Bureau Veritas | All allowed | No | No |
| Capgemini | All allowed | Yes | Yes |
| Carrefour | All allowed | No | No |
| Crédit Agricole | Not measurable | No | Yes |
| Danone | All allowed | No | Yes |
| Dassault Systèmes | All allowed | Yes | Yes |
| Edenred | All allowed | No | No |
| Engie | All allowed | No | Yes |
| EssilorLuxottica | All allowed | No | No |
| Eurofins | All allowed | No | Yes |
| Hermès | All allowed | No | No |
| Kering | All allowed | No | Yes |
| Legrand | All allowed | No | Yes |
| L'Oréal | All allowed | No | Yes |
| LVMH | All allowed | No | No |
| Michelin | Not measurable | No | Yes |
| Orange | Not measurable | No | Yes |
| Pernod Ricard | All allowed | No | No |
| Publicis | All allowed | No | No |
| Renault | All allowed | No | Yes |
| Safran | All allowed | Yes | No |
| Saint-Gobain | All allowed | No | No |
| Sanofi | All allowed | No | Yes |
| Schneider Electric | All allowed | No | Yes |
| Société Générale | All allowed | No | Yes |
| Stellantis | All allowed | No | No |
| STMicroelectronics | Not measurable | No | No |
| Teleperformance | All allowed | No | Yes |
| Thales | All allowed | No | No |
| TotalEnergies | All allowed | No | No |
| Unibail-Rodamco-Westfield | All allowed | No | Yes |
| Veolia | All allowed | Yes | No |
| Vinci | All allowed | No | Yes |
Method
- Sample: 40 large listed French companies, close to the CAC 40 composition, main site of each group.
- AI crawlers: robots.txt read the way each crawler would (its own User-agent group if any, otherwise the “*” group), for OAI-SearchBot, ChatGPT-User, GPTBot, PerplexityBot, Claude-SearchBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot and Bingbot.
- llms.txt: presence of a text file at /llms.txt.
- Structured data: JSON-LD blocks present in the homepage as served to a bot.
- Limits: a firewall may serve bots a different version; those sites are marked “not measurable”. The measurement covers a single day.
- The data from this study can be reused freely, citing CoolSEO and this page.
Frequently asked questions
Do large French companies block ChatGPT?
No, according to our measurement on October 6, 2026: none of the 35 companies with a readable robots.txt blocks OpenAI's crawlers (GPTBot, OAI-SearchBot) or those of other AIs.
How many large companies have an llms.txt file?
4 out of 40 in our sample: Capgemini, Dassault Systèmes, Safran, Veolia.
Can I reuse these figures?
Yes, freely, citing CoolSEO and a link to this page.
Go further
Does AI cite your business?
The free audit checks your site on Google and asks Gemini a real question about your market, in a minute.