Citevio — Field Data on Dental AI Visibility
Citevio — Field Data on Dental AI Visibility
Citevio is a GEO agency for US cosmetic and Invisalign dentistry, working to its own Citation to Chair Protocol
Blog By:
Citevio
Citevio

A failed measurement is not a finding. It is a missing one.

8/20/2026 8:56:00 PM   |   Comments: 0   |   Views: 29
Short answer: When we scanned 527 US dental practice websites between 6 and 26 June 2026, not every check produced a usable reading for every site. Page speed could only be measured on 256 of them. One directory check returned a valid reading for 291 sites; another for just 146. We dropped each failed reading from that row alone rather than scoring it as a bad result — and that single decision changes almost every percentage a study like this can publish.

A failed request tells you about your measurement tool. A completed check that finds nothing tells you about the thing you measured. Combining them produces a cleaner-looking number and a worse claim.

One study, several honest denominators

The temptation is to publish everything against a single round number, because 527 appears in every row and the table looks authoritative. Here is what the same study looks like when each row keeps only the readings it actually obtained.
Check, 6–26 June 2026Sites with a usable readingSites with no usable reading
Full sample scanned527
Page speed (First Contentful Paint)256the rest returned a measurement error
Bing Places presence291236
Foursquare presence146381
Google review data5234 practices had none
Look at the Foursquare row. A valid reading existed for 146 sites and did not exist for 381. Whatever that row says, it is a statement about 146 practices. Publishing its result as a share of 527 would be reporting mostly about sites we never successfully checked.

Within its own denominator, that row reads: 66 of 146 came back OK, 45.2%; 47 were WARN, 32.2%; and 33 were FAIL, 22.6%. The Bing Places row, on its 291 valid readings, was 256 OK, 88.0%, and 35 WARN, 12.0%. Those two checks look nothing alike — and part of the difference may be the checks themselves rather than the practices, which is precisely why each keeps its own base.

Even the review row moves. Four practices had no review data at all, so the review figures rest on 523, not 527. It is a difference of four, it changes no conclusion, and we publish it anyway, because a study that rounds away its small inconveniences has already shown you what it does with the large ones.

The row that is neither a win nor a loss

The clearest case we have of this rule sits in a different measurement entirely. In a separate panel of 16 buyer-intent questions producing 64 valid answers, a third engine was included in the original design and returned a quota error on all 16 attempts.

That column is not a defeat. It is a blank. Reporting it as "we did not appear on that engine" would be false in a way that flatters nobody and misleads everybody — it would convert an infrastructure failure on our side into a verdict from the engine. "We do not know" is the finding, and it stays in the report under that name.

What the mistake does to a headline

The distortion runs in both directions, which is why it survives so long unnoticed.

Fold errors into the negative class and invisibility inflates. Every timeout, rate limit, and authentication failure quietly becomes another business the assistant "ignored," and the resulting percentage is larger, more quotable, and partly manufactured. This is the shape of most alarming statistics in this category.

Fold them the other way and a broken measurement system disappears. If a scan silently drops everything it could not read and never reports how much that was, nobody can tell whether the study checked 90% of its sample or 28% of it. The Foursquare row above is the honest version of that disclosure.

The same discipline applies before a scan even begins. In a separate crawler study we started with 7,632 rows, could not reach 1,135 domains at all, and published the denominator as 6,497. An unreachable domain has no readable policy file. It cannot support a yes, and it cannot support a no.

The question to ask any report you are shown

There is a short version of everything above, and it works as a filter on other people's studies as well as ours. Ask what the denominator is for each individual figure, not for the study as a whole. Ask how many attempted readings failed and where those went. Then ask whether the failures are reported anywhere in the document at all.

A study that answers all three is not necessarily right, but it can be argued with. A study that publishes one confident denominator across every row is either unusually lucky with its instruments or is not telling you where its gaps went. In our experience the second is far more common, and it is rarely deliberate — it is what happens when a spreadsheet has a blank cell and someone reads the blank as a zero.

Does a low result explain why one particular practice is absent?

No, and this is where sample-level findings get misused most often. Nothing in this design measures the cause of any individual outcome. We ran no controlled intervention, changed nothing, and re-measured nothing afterwards.

A practice can be perfectly reachable and still absent from an answer. It can be unreachable during one check and reachable an hour later. It can appear on a later run for reasons no one recorded. Each of those conditions needs its own label, and none of them is interchangeable with the others.

How to record your own checks

If you are going to check whether an assistant mentions your practice — and you should, rather than paying someone else for a screenshot — the recording rules matter more than the tool.

Fix the exact question wording before you start, and do not adjust it once you have seen an answer you dislike. Record the engine, the date, and whether the answer completed at all. Then keep three states in separate columns, never two: named, not named, and error.

Decide the number of runs in advance and keep every one of them, including the disappointing ones. If you check again after making a change, use the same question, the same engine, and the same run count, or the comparison means nothing.

And do not let a single appearance become a trend in your own notes. One run is one observation. Our own rule asks for the same result on two consecutive runs before anything gets called a pattern, and we hold our own reporting to it.

What we did not measure

We did not measure why any individual check failed — whether the cause sat with the site, the network, the API, or the tool. We did not measure whether a site that failed one check would have passed on a different day. And we did not measure how long any change takes to appear in an AI answer, which is why no interval in this post can be turned into a schedule or a promise.

Common questions

Is a failed request the same as being invisible?
No. It is an unresolved measurement, not a negative answer, and it belongs in its own column.

Why do different rows in the same study have different denominators?
Because a practice missing one reading is dropped only from that row. Page speed rests on 256 readings, Bing Places on 291, Foursquare on 146, and review data on 523 — from the same 527-site scan.

Can two engines' results be averaged into one rate?
Not when they have different readable bases. Separate instruments stay separately reported.

Does an omission prove the website has a technical problem?
No. The measurement does not isolate why a name was absent, and a technical explanation is only one of several the data cannot distinguish between.

Where are the denominator rules documented?
They are published with the data, alongside the method and limitations, at citevio.com/data.

About the author and disclosure

Muhammed Veysel Erin is the founder of Citevio. Citevio is a vendor, not a dental practice. Citevio produced the measurements discussed here and publishes the failed-reading counts alongside the successful ones. Nothing in this post claims that an observed association causes an AI outcome.
You must be logged in to view comments.
Total Blog Activity
997
Total Bloggers
13,451
Total Blog Posts
4,671
Total Podcasts
1,788
Total Videos
Sponsors
Townie Perks