How We Check AI Search Advice Before We Act On It
The advice about how to show up in AI search is being written faster than anyone can check it. A lot of it is wrong, and some of it is wrong in the same document that contains the evidence proving it wrong.
We read every AI search source twice, with two different models, working independently. Then we line the two readings up and label every finding: confirmed if both models found it, single source if only one did, conflict if they disagree. Only the conflicts get a human. That last category is the reason the process exists, because published guidance about AI search contradicts itself often enough that acting on it unchecked will cost you a quarter.
I spent a decade in rooms where a wrong number was a public problem. Recovery.gov during the stimulus, a Pentagon spending tracker, the Census Bureau. The habit those rooms build is not skepticism exactly. It is a reflex to ask where a number came from before you repeat it, because the moment you repeat it, it is yours.
AI search guidance is currently produced with none of that reflex. It is fast, it is confident, it is mostly summaries of other summaries, and almost nobody is checking it against the evidence sitting inside the source itself.
What we found in a well regarded source
A widely watched course video on answer engine optimization says out loud that YouTube is the most cited domain. On screen in that same video, a chart titled "100 most cited domains in ChatGPT" lists Reddit first at 4.38 million mentions, Wikipedia second, and YouTube fifth at 214,000.
Neither statement is a lie. YouTube does lead in Google AI Overviews. Reddit does lead in ChatGPT. The video simply blurs the two platforms into one claim, and the spoken version is the one people remember. A content team that heard it and built a video program would have spent a quarter aiming at the wrong engine, using a source that had the correct answer on screen the entire time.
We caught it because two models read that video separately and came back disagreeing. One followed the audio. The other read the chart. The disagreement was the signal.
The method
Two models, same source, no contact between them. One reads the transcript and the rendered frames. The other watches the source natively. Their findings then get aligned line by line and every line gets a label.
Confirmed means both readings independently surfaced it. We use it. Single source means only one did. We keep it and flag it, because a fact one careful reader found is usually real but has not been double checked. Conflict means the readings disagree, and we go back to the original source and resolve it before anything downstream uses that line.
A single model reading a source alone gives you a confident summary with no way to tell which lines are solid. Two models disagreeing tell you exactly where to look.
The practical effect is that verification stops meaning read everything again and starts meaning read the three lines that are actually in dispute. That is what makes it something you can run on every source instead of only on the ones you happen to be suspicious of.
Then we check whether the page can act on it
Verification decides what advice is true. A separate check decides whether a page can benefit from it. We score pages 0 to 100 on citation readiness across six things: hedging language, whether a direct answer sits at the top, declarative density, a freshness signal, an author or entity signal, and structured data.
Two of those came back interesting on real sites. Pages that read perfectly well to a human routinely have no declarative answer in the first screen of extractable text, because the top of the page is a warm up. And sites that rebuild themselves weekly frequently carry no freshness signal at all, which matters more than it sounds: in the research we verified, roughly three quarters of the pages ChatGPT cites most were refreshed within thirty days.
The part most methodology pages leave out
The first version of our own scoring flagged twelve of thirteen sites on one check. That is not thirteen bad websites. That is a broken check, because a signal that fires on everything tells you nothing. We had picked the threshold by intuition instead of measuring the distribution, so we measured it and moved the threshold to where it actually separated strong pages from weak ones.
Then we widened the underlying detector to fix a different error, and the correction pushed five of eight sites to a perfect score, which was just as useless in the opposite direction. So we threw the change out the same hour and re-derived the threshold against the new detector, because a cutoff derived under one measurement is meaningless under another.
Every threshold we use now carries a note saying whether it was measured or asserted, and the asserted ones say so plainly. We would rather hand you a number with its provenance attached than a confident score built on a guess nobody wrote down.
What we do not claim
Statistics from third party research stay attributed to whoever published them, and we do not present them as ours or as replicated. Thresholds tied to observed citation gain are still accruing, and until that history is long enough to mean something, we say so rather than dressing an opinion up as a finding. If a number here is an operator judgment, it is labeled an operator judgment.