Short answer: We split 6,497 US local business domains into four bands by review count and read what each one's robots.txt declared about AI and data crawlers. In the lowest band, 13.0% of 1,648 domains blocked at least one of the 13 AI and data crawlers we checked. In the highest band, 8.7% of 1,624 domains did. That is an association in a snapshot taken on 22 July 2026 — not evidence that reviews change crawler access, and not a reason to buy reviews as a technical fix.
The interesting part of this dataset is not the headline. It is what happens to the pattern when you change which crawlers you count.
What the dataset is, precisely
The scan began with 7,632 rows. Domains we could not reach at all were dropped — 1,135 of them — because an unreachable domain has no readable policy file and cannot support either answer. That leaves a published denominator of 6,497 domains, measured on 22 July 2026 and recounted on 30 July 2026.
One definitional choice does a lot of work here. A domain counts as blocking if at least one crawler column came back blocked. A column marked partial is not counted as blocked. That rule is deliberately conservative: it means every figure below is a floor rather than a dramatised total, and a report using a looser definition would produce larger numbers from the same file.
The band comparison
| Review count band | Domains (n) | Blocked =1 of 13 AI and data crawlers | Blocked =1 of the crawlers used by ChatGPT, Claude or Perplexity |
|---|
| 0–25 | 1,648 | 13.0% | 7.8% |
| 26–122 | 1,605 | 12.8% | 9.0% |
| 123–315 | 1,620 | 10.9% | 6.7% |
| 316+ | 1,624 | 8.7% | 7.0% |
Read the third column alone and there is a tidy story: as review counts rise, blocking falls, from 13.0% down to 8.7%. It is the kind of line that ends up in a slide deck with an arrow drawn on it.
Now read the fourth column. Restricted to the crawlers actually used by the three assistants most people are asking about, the ordering falls apart. The highest rate in that column belongs to the second band, at 9.0% of 1,605 domains — not the lowest band. The most-reviewed band sits at 7.0%, above the 6.7% of the band below it. The neat descent exists in one instrument and not in the other.
That is the single most useful thing in this table, and it is an argument against trusting the table.
One more property of this table is easy to skim past: the four bands are close to the same size, at 1,648, 1,605, 1,620 and 1,624 domains. The boundaries land on unrounded values like 122 and 315 rather than tidy ones, and the four groups came out close to equal in size. Two things follow. No band's rate is resting on a handful of domains, so the instability in the fourth column is not simply a small-sample artefact. And it is worth checking this property in any banded table you are shown: if one band holds a few hundred rows and another holds several thousand, their rates are not equally solid, and putting them in the same column invites a comparison the data cannot carry.
Two labels, one very easy mistake
The two columns are not interchangeable, and swapping their labels is the most common error we see in this category. "13 AI and data crawlers" is a wider set that includes crawlers no consumer assistant reads answers from. "The crawlers used by ChatGPT, Claude or Perplexity" is a narrower set of seven columns.
Quoting the larger, more alarming number from the wide set while describing it with the narrow set's qualifier produces a sentence that is technically assembled from real data and is nonetheless false. If you see a statistic about dental sites "blocking ChatGPT," ask which crawler list produced it before you react to the size of it.
Why this is not a review strategy
The study did not assign practices more reviews, hold everything else constant, and recheck their policy files afterwards. It read two fields that already existed on the same day. That design can describe a pattern. It cannot establish what produced one.
Plenty of ordinary explanations sit unmeasured behind a correlation like this. A practice with 400 reviews has usually been operating longer, may have a larger budget, is more likely to have a maintained website, and is more likely to be on hosting where someone has looked at the configuration this decade. Site age, platform, agency involvement, and technical upkeep all travel with review count in the real world. We collected none of them, which is exactly why the finding stays descriptive.
What we did not measure
We did not measure who wrote each robots.txt file, or whether a platform default produced the rule without anyone choosing it. We did not measure whether any rule changed before or after our scan. And we did not measure whether either review count or crawler policy changes the probability that an assistant names a business.
That last one matters most. A practice can have many reviews, a wide-open policy file, and still never be mentioned. The reverse also appears in a snapshot. Neither case would establish a cause.
The denominator closest to your practice
The bands above are drawn from all local business domains in the scan. If you want the row that sits nearest to a dental practice, the dentist subgroup is a better fit: among 3,929 dentist domains, 12.5% blocked at least one of the 13 AI and data crawlers, and 8.3% blocked at least one of the crawlers used by ChatGPT, Claude or Perplexity.
For context, across all 6,497 domains the same two figures are 11.4% and 7.6%. Dentist domains sat slightly above the general population on both measures on the day of the scan. That is a small difference, in one snapshot, and it should not be inflated into a claim about the profession.
What to actually do with this
Separate the two jobs, because they answer to different owners and different evidence.
Reviews are patient-facing evidence about experience. They should be earned legitimately and should reflect current care. That case does not need an AI argument propped underneath it, and it is weakened by one.
Crawler access is technical policy. Audit it directly: read the actual robots.txt rules on your own domain, note which crawler identities are named, and separately test what your server returns when those identities ask — a permissive file and a permissive server are not the same thing, and the file is only half the check.
What this dataset does not support is buying reviews as an access fix. Nothing in these four bands makes that a measured intervention.
Common questions
Do more reviews make a site easier for AI crawlers to reach?
We did not test that. The band difference in this 22 July 2026 scan is an association, and it does not survive intact when the crawler set is narrowed.
Are all crawler blocks bad?
No. A practice may refuse a crawler on purpose, and that is a legitimate choice. The problem is a rule nobody chose and nobody can explain.
Can this table tell me why my practice is missing from an AI answer?
No. It measures published crawler rules on a given day. It does not measure why an engine names one business and not another.
Why is partial not counted as blocked?
Because it is ambiguous, and counting ambiguous cases as blocked would inflate every figure here. We chose the conservative reading.
Where are the full definitions?
The denominator rules, the crawler set, and the limitations are published at citevio.com/data.
About the author and disclosure
Muhammed Veysel Erin is the founder of Citevio. Citevio is a vendor, not a dental practice. Citevio collected the aggregated data described in this post. No product, ranking, or outcome is offered here, and the association reported above is not presented as a cause.