Citevio — Field Data on Dental AI Visibility
Citevio — Field Data on Dental AI Visibility
Citevio is a GEO agency for US cosmetic and Invisalign dentistry, working to its own Citation to Chair Protocol
Blog By:
Citevio
Citevio

Your robots.txt says yes. Your server can still say no. We tested 499 dental sites.

8/17/2026 5:20:00 AM   |   Comments: 0   |   Views: 65
Short answer: We took 499 US dental practice websites whose robots.txt already permitted AI crawlers and whose homepages already loaded for an ordinary browser request, then asked each server for the same homepage under real AI crawler identities. On 22 July 2026, 72 of those 499 sites — 14.4% — refused at least one of them. The published policy and the delivered response did not always agree.

Before anything else: this measures access, not visibility. We did not measure whether removing a refusal changes whether an assistant ever names a practice. Being readable is a prerequisite, not a lever with a promised outcome, and nobody should sell it as one.

What the test actually did

Most advice on this topic stops at the policy file. Ours started after that question had already been answered, because a file is a statement of intent and a server response is what actually happens.

We drew 550 sites at random from domains whose robots.txt allowed AI crawlers. Each homepage was then requested eight times in a fixed order: once as an ordinary browser, six times under AI crawler identities, and once as Googlebot as a control. Any site whose plain browser request failed was dropped, because a site that is down for everyone tells us nothing about crawler treatment. That left the denominator of 499.
Design elementValue
Sites drawn at random550
Requests per homepage8, in a fixed order
— plain browser1
— AI crawler identities6
— Googlebot control1
Sites dropped (plain request failed)Removed before analysis
Final denominator499
Measurement date22 July 2026
The denominator is worth stating in full whenever these percentages are quoted: sites whose robots.txt allows AI crawlers and whose homepage loaded for a plain browser request. Quoting 14.4% without that sentence attached makes it a different and much wilder claim.

What came back

Observation, 22 July 2026n of 499%
Refused at least one AI crawler at the server7214.4
— of those: an explicit block status (403 or 429)6913.8
— of those: a bot challenge page30.6
Refused an AI crawler while serving Googlebot normally6312.6
Refused Googlebot61.2
Answered every crawler but sent a different body to at least one418.2
The Googlebot control is the row that changes the interpretation. A server that refuses every automated request has a general access problem, which is a different and more visible issue. A server that hands Googlebot the page and refuses an AI crawler is applying a rule that depends on who is asking. That was 63 of the 499 sites, or 12.6%, on the day we measured — and only 6 sites, 1.2%, refused Googlebot itself.

The last row is the quietest and the easiest to miss in a status-code-only audit. Forty-one sites, 8.2%, answered every identity successfully and still returned a different page body to at least one of them. Nothing in a status column flags that. If your check only records "200 OK," you would score all forty-one as clean.

Which identities got refused

Refusals were not spread evenly across crawler identities, and the two control identities sit at the bottom of the same table.
Crawler identityn of 499%
ClaudeBot469.2
GPTBot408.0
OAI-SearchBot173.4
PerplexityBot102.0
ChatGPT-User71.4
Bingbot (control)61.2
Googlebot (control)61.2
We did not measure why one identity was refused more often than another. A user-agent string is easy to match on and easy to copy between configurations, so the ordering above describes what happened, not a hierarchy of intent.

The view from the request side

Counting sites and counting requests answer different questions, and mixing the two produces a scarier number than the data supports. Across 499 sites and 6 AI crawler identities each, the request-level denominator is 2,994 — not 499. On that base, 2,841 requests, 94.9%, returned HTTP 200. Blocks were 79 requests at 403, 2.6%, and 38 at 429, 1.3%.

So the honest summary carries both frames at once. Most individual requests succeeded. Yet refusals clustered enough that roughly one site in seven turned away at least one AI crawler. A per-site problem and a per-request rate are not interchangeable, and a report that quotes only the friendlier of the two is selecting its own headline.

Why can the file and the server disagree?

We did not measure the cause, and this is the point where most write-ups start guessing. Our data does not record whether a hosting rule, a security product, a content delivery network, a plugin, or a human decision produced any given refusal. Naming one would convert an observation into a diagnosis the study cannot support.

What the study does establish is where the explanation is not located. The public robots.txt rule cannot explain a refusal that occurs after that file has already said yes. Whatever made the decision sits somewhere further down the stack, and on the sites we tested, it was frequently a layer the practice had never been told existed.

What this finding does not mean

Four limits travel with every one of the numbers above.

It is one day's photograph, taken on 22 July 2026. Hosting and CDN rules change without notice, and a repeat run on another date will not reproduce this file exactly. That is expected behaviour, not an error.

It carries no causal claim. What we measured is access, not citation. We did not measure whether removing a refusal makes an assistant more likely to mention a practice.

The cause layer is unknown, as described above.

And it must not be blended with the separate, much larger studies that read robots.txt declarations across thousands of domains. Those read what a file says. This one measures what a server does. Different instruments, different denominators — the rates are not substitutes for one another and must never be averaged together.

How to run the same check on your own site

Ask whoever manages your website to request your homepage three times: once under an ordinary browser identity, once under the AI crawler identity you care about, and once as Googlebot. Have them record two things for each request, not one — the returned status code and the returned page body. A challenge page can arrive wearing a perfectly successful status, which is exactly how the three challenge cases in our sample would have been missed by a status-only check.

Then compare the three responses. If they differ, ask to see the exact rule that produced the difference. A visible rule can be reviewed, kept, or removed on purpose. A reassurance without a result cannot be checked by anyone.

Finally, record what you found before you change anything. If you later want to know whether a change mattered, you will need the earlier state, the same questions, and the same run rules — otherwise any difference afterwards is a story, not a measurement.

Common questions

Is an open robots.txt enough?
No. In this 22 July 2026 study, 72 of 499 sites with permissive policy files — 14.4% — still refused at least one AI crawler at the server.

Does a refusal prove the practice intended to block AI?
No. We measured the response, not who configured it or why. Intent was never in the dataset.

Should every AI crawler be allowed?
That is a policy decision for the practice to make deliberately. This post argues only that the published choice and the enforced choice should match.

Will fixing this get my practice cited?
We have not measured that, and no one should promise it. A successful fetch shows the page can be served. It does not show that any engine will index, cite, or name the practice.

Where can the method and counts be checked?
The aggregated counts, denominator definitions, method, and limitations are published at citevio.com/data.

About the author and disclosure

Muhammed Veysel Erin is the founder of Citevio. Citevio is a vendor, not a dental practice. Citevio ran the server-response measurement discussed here and publishes the aggregated method and counts openly. This post offers no service and makes no placement promise.
You must be logged in to view comments.
Total Blog Activity
997
Total Bloggers
13,451
Total Blog Posts
4,671
Total Podcasts
1,788
Total Videos
Sponsors
Townie Perks