Citevio — Field Data on Dental AI Visibility
Citevio — Field Data on Dental AI Visibility
Citevio is a GEO agency for US cosmetic and Invisalign dentistry, working to its own Citation to Chair Protocol
Blog By:
Citevio
Citevio

We asked dental homepages to answer as AI crawlers, not as browsers

8/29/2026 11:20:00 PM   |   Comments: 0   |   Views: 33

Short answer: A permission file is not the final access test. We requested dental homepages using AI-crawler identities after selecting sites whose robots.txt welcomed those crawlers. Some still refused the request at the server layer. The practical lesson is modest: read the declaration, then test the response. Do not call a site accessible based on the declaration alone.

There is an easy audit error here. A robots.txt file is visible, quick to read, and gives an answer that feels binary. A crawler is allowed or it is not. But a real request travels past that file into hosting configuration, security products, CDN rules, plugins, and whatever other rules sit between the request and the homepage.

The server is where the answer becomes operational.

In Citevio's live crawler access test of 22 July 2026, 72 of 499 dental homepages (14.4%) whose robots.txt allowed AI crawlers were refused by their own server anyway.

That result is not a finding about why any individual site refused the request. The test did not assign a reason to individual failures, and it did not test what later AI answers would look like after a refusal was removed. It only records a disagreement between a public permission statement and the response returned to the crawler identity in that run.

Start with the question you can answer

Do not begin by asking whether an AI system will recommend the practice. This procedure cannot answer that. Begin with the narrower question: did the homepage return usable content when asked for it under the identity you want to assess?

That question has a clear record. Save the URL, the request identity, the time, the status, the final destination after redirects, the response headers, and enough of the returned body to show whether it is the actual page or an interstitial. A status code alone is a poor witness. A challenge page can arrive with a normal-looking status while the requested page is absent.

Run a paired request

Ask the person who controls the website, host, or CDN to perform the test from their own environment. A simple browser visit is the comparison condition. Then send the same homepage request with an AI crawler user agent. If the practice wants a conventional search comparison, add the relevant search crawler identity as a separate request.

Keep the URL unchanged across the requests. Follow redirects. Capture the response rather than relying on the rendered browser screen. The point is not to impersonate a crawler casually and declare a universal result; the point is to create a reproducible record of what the server returned under the tested identity at that time.

Where the ordinary request returns the expected page but the crawler request returns a denial, an empty body, a challenge, or a different document, treat that mismatch as an access issue to investigate. It is not an instruction to change settings blindly. Send the record to the owner of the hosting, firewall, CDN, or website configuration and ask which rule was applied.

Check the declaration separately

Read robots.txt as its own document. It states a policy for named crawlers. Save the file with the date and keep it beside the request record. Do not use the file as proof that the homepage was served, and do not use the homepage response as proof that every crawler is covered by the file. They are distinct observations.

This separation is especially useful when different people own different layers. A marketing person may edit page copy; a web company may maintain the CMS; a host or security provider may operate the refusal rule. A clear test record lets the practice send the right question to the right owner instead of creating a vague ticket that nobody can resolve.

What to preserve when the request fails

The most valuable artifact is not a screenshot of a broken page. It is the request-and-response pair. Preserve the user-agent string used, the URL requested, the redirect path, the status, a sample of headers, and a short body capture. Note whether the normal browser request received the expected homepage at the same time.

That record prevents a common argument after the fact: someone reads a permissive robots.txt file, another person opens the site normally, and both conclude there could not have been a problem. Neither observation reproduces the crawler request.

Avoid treating an intermittent outcome as a permanent verdict. A temporary defense rule, rate limit, or cached configuration can create a response that later differs. The purpose of the audit is to make the disagreement inspectable, not to attach a durable label to the practice website.

Give the configuration owner a question they can answer

The next step should be specific. Do not send a general request to “make the site AI ready.” Send the captured request and ask whether a rule evaluated the crawler identity differently from the browser request. Ask whether the response came from the origin server, a security layer, a cached edge response, or a challenge service. Ask what evidence the owner would need before changing the rule.

That framing protects the practice from two unhelpful reactions. The first is an unsupported assurance based only on the robots.txt file. The second is a sweeping technical change made without knowing which component returned the refusal. A dated response record does not solve the configuration by itself; it makes the investigation narrower and more accountable.

If a setting is changed, run the same documented request again and preserve the new response beside the earlier one. Treat the later record as another observation, not retroactive proof of the original reason. This keeps the troubleshooting record useful even when a provider changes several settings at once.

This is a snapshot, not a permanent file

The measurement records a single day, not a permanent file. Hosting and CDN policies may move without notice, so a later request should not be expected to return an identical response. That is a property of the environment, not a flaw to edit out of the result.

The limit matters in both directions. A clean response today is not a warranty that a later configuration will behave the same way. A refusal in the recorded run is not proof that the homepage always refuses the crawler. Re-test after a relevant change and keep the records as separate dated observations.

Where a DIY check is enough

This portion of the work is genuinely manageable in-house. A practice can request the evidence, check whether the returned document is the real homepage, and open a precise ticket when it is not. It does not require turning the check into a broad marketing project.

The boundary appears later, when a practice wants repeated readings, a stable set of patient questions, and coordinated work across content, access, and business facts. Those are ownership and process questions. They are not proof that a practice team cannot perform the initial access check.

For the wider DIY-versus-agency decision guide, including the limits of this test, see Citevio’s guide.

About the author and disclosure

Muhammed Veysel Erin is the founder of Citevio. Citevio is a vendor, not a dental practice. Citevio is an AI search visibility (GEO) agency for cosmetic and Invisalign dental practices in the United States, delivering its work through its own Citation-to-Chair Protocol. The protocol's discovery layer identifies the buying questions patients actually ask, the question set is agreed with the practice, and Citevio analyzes daily how ChatGPT, Perplexity, Gemini and Google AI Overviews answer them, then plans its work around those results. Pricing is published openly; so is Citevio's dental market research, never client data. Contact: contact@citevio.com · +90-544-774-7558 · Sheridan, WY, US. Citevio ran the access test discussed here. This post offers no service and makes no placement promise.

You must be logged in to view comments.
Total Blog Activity
997
Total Bloggers
13,451
Total Blog Posts
4,671
Total Podcasts
1,788
Total Videos
Sponsors
Townie Perks