Five weeks of daily AI citation readings on our own domain, and the three specific ways a GEO tracker will mislead you if you read it like a rank tracker.

A generative engine optimization checker tells you how often an AI assistant pulled your pages into an answer. That is the honest, narrow version of what these tools do. They do not tell you whether the assistant said your name, they mostly do not cover every engine that matters, and at least two of the dashboards we run daily have served us four-day-old numbers while presenting them as current. We know because we have been reading four of them every single day on our own domain since 28 July 2026, and writing every reading down.
That log is the reason for this post. Thirty-three dated readings across five weeks, on one site, on four surfaces, cross-checked by asking the assistants directly. Below is what the trackers agreed on, where they disagreed with each other, and the three specific ways a GEO checker will mislead you if you read it the way you read a rank tracker.
Three different things get sold under the same label, and mixing them up is the first mistake.
Citation counts. Semrush's AI Search widget and Bing Webmaster Tools' AI Performance panel both count how many of your URLs appeared as a source under an AI answer. This is retrieval. The assistant fetched your page and used it.
Brand mentions. A separate metric, counting whether your company name appears in the answer prose. Semrush breaks this out as "Mentions" and it is not the same number.
Prompt-level checks. Some tools run a list of prompts against assistants on a schedule and diff the answers. This is the closest thing to a real rank tracker, and the most expensive to run honestly.
Our own numbers show why the distinction is not academic. Over five weeks our Semrush Cited Pages went from 0 to 19. Over the same five weeks, AI Visibility read 0 and Mentions read 0. Every reading. Nineteen of our pages were being fetched and used to build answers, and not one of those answers named us.

Semrush breaks Cited Pages out by engine. On our most recent fresh dataset the split was ChatGPT 5, Google AI Mode 8, Gemini 7, and Google AI Overview 0.
That AI Overview zero is not a slow start. It is every reading. Thirty-three consecutive readings across five weeks, and the number has never once been anything but 0, while the other three engines each moved off 0 independently and at different times.
If you read only the headline Cited Pages figure, you watched it climb 0 to 19 and concluded your AI visibility was improving. It was, on three engines. On the fourth, the one attached to Google's actual search results page, nothing happened at all. We now log AI Overview as its own line in every report specifically so an aggregate rise driven by the other three can never be reported as if all four were moving.
Ask your tracker for the engine breakdown before you believe the total. If it only gives you a total, it is hiding the most useful thing it knows.
This is the failure mode nobody warns you about, and it is the one most likely to make you draw a wrong conclusion.
On 29 August we read the Semrush widget twice in one day after it came back from an outage. Both reads returned Cited Pages 16 with an identical engine split. The dataset was stamped 28 August. It was yesterday's numbers, rendered as today's, with no visible staleness indicator. Had we logged those as two fresh readings, we would have counted a flat trend that was really a gap.
Bing's AI Performance panel does the same thing differently. Its seven-day view once showed Total Citations 165 while the underlying date range ran only through a date four days earlier. Not wrong, just answering a question about a window that had already closed.
We now record a blocked or stale surface as a gap, never as a zero and never as flat. That single discipline is what stopped us calling three false trends. The rule we settled on after the second incident: never call a trend off fewer than three consecutive readings, and write down in advance which reading would have to flip to change the call.
Semrush's widget covers ChatGPT, Google AI Overview, Google AI Mode, and Gemini. Perplexity is not in it.
Perplexity is where the only real movement of the whole five weeks happened, and we only saw it because we ask the assistants directly every cycle instead of trusting the dashboards.
On 31 August we published a post on computer vision ROI in retail and put its own target query to Perplexity. Our URL came back as source 8 of 15, with our real title on the source card, while the answer prose named other vendors. Retrieved, not cited. The next day, the same query, our name appeared inline in the answer text. The day after that, the page was cited nine times inline against nine distinct claims and the source card had risen to second of fourteen hosts.
Absent, then retrieved, then cited, then cited throughout. Three days. On a surface the tracker does not measure.
Then we ran the honest half of the test, because one query proves nothing. We took a different post, a guide on digital process automation, and put its target query to Perplexity. Complete absence. Fifteen sources, all established vendors, our name nowhere. So the correct read is a page-level win on one low-competition query, not domain-level AI visibility, and we wrote it up that way rather than as "Perplexity now cites us".
That is the shape of most GEO wins. Narrow, page-specific, and invisible to a domain-level dashboard.
Google Search Console has a "Generative AI features" report. Over six readings ours went 561 impressions across 82 pages to 753 across 106. It is first-party data about Google's own AI surfaces, it costs nothing, and every third-party tracker is guessing at what it states directly.
The catch is that it is a UI-only report. We tried pulling it through the Search Analytics API with the searchAppearance dimension and got zero rows back. If you want it, someone opens the page and reads it. There is no automating that part, which is presumably why so few teams look at it.
One more caution learned the hard way: watch the shape of a rise, not just the direction. One of ours was seven new pages entering the report. The next was plus twenty-six impressions on plus one page. Breadth and depth are different results, and merging them into one "rising" story throws away the only interesting part.
We check four assistants by hand each cycle, and the tempting shortcut is to search the rendered page for your brand name and count hits. Do not do that naively.
Three separate times, our Gemini check returned a single body-wide match for "codestreaks" that resolved, on inspection, to a chat title in the conversation sidebar. Not the answer. An earlier session of ours had been named after the site, so the tracker was finding its own history.
Slice the answer region first, then count, then print the surrounding text for every hit and read it. We also count occurrences and classify each one as inline citation, source card, or neither, because "appeared on the page" and "was cited in the answer" are the two states people conflate most.
One more thing worth logging: whether the assistant browsed at all. A ChatGPT answer with zero external links is not evidence you were skipped, it is evidence it never looked. We mark those as unfair tests rather than absences.
Nothing we use covers Perplexity, breaks engines out honestly, stamps dataset freshness on the page, and distinguishes retrieved from cited. Until something does, the setup that has worked for us costs nothing but discipline:
Five weeks of that produced a clearer picture than any single dashboard did, mostly because it caught the dashboards being wrong.
A tool that reports how often AI assistants retrieve or mention your site when answering questions. Most count citations, which is retrieval. Fewer count brand mentions in the answer text, and those two numbers can differ wildly. Ours ran 19 cited pages against 0 mentions.
They are accurate about what they measure, which is narrower than the label suggests. In our log, two dashboards served datasets several days old without saying so, one covered four engines and missed the one where we actually got cited, and the aggregate number concealed an engine that had never moved. Treat them as one input, not a scoreboard.
Daily if you are running an active content programme, but only if you also record the readings. The value came from the log, not from any single reading. Weekly checks with no written history cannot tell a real trend from a stale dataset, which is the specific failure we hit twice.
Not directly, and this is the honest answer nobody likes. Semrush reports an AI Overview cited-pages count and ours has read 0 for every one of 33 readings. Google Search Console's Generative AI features report gives you first-party impression data on Google's AI surfaces, but it is UI-only and does not break out AI Overview separately.
Not reliably. Our 19 cited pages and 753 generative-AI impressions have coincided with a blog click-through rate that has not moved off 0.10%. We wrote up that gap separately in our post on ranking without clicks. Citation and traffic are separate outcomes and should be measured separately.
Written by the Codestreaks team from our own tracking log, docs/seo/geo-tracking.md, which holds 33 dated readings taken between 28 July and 2 September 2026 across the Semrush AI Search widget, Bing Webmaster Tools AI Performance, the Google Search Console Generative AI features report, and direct spot-checks against ChatGPT, Gemini, Perplexity and Claude. Every number here is a reading from that file on codestreaks.com, our own domain, with a 272-URL sitemap. Drafting is AI-assisted and every draft gets a human editing pass against the measured data before it ships. Where a surface was unreachable on a given day we recorded a gap, and this post says so rather than filling it in.
If you are building content for AI answers rather than just for rankings, we covered the on-page side in preparing your SaaS for AI search, and the automation side in running SEO agents in production.
If you want a second pair of eyes on your own measurement setup, our AI consulting engagements usually start with exactly this kind of audit: what you are counting, what it actually means, and which number you should be acting on. Book a free 30-minute scoping call through start a project and we will come back within two business days. We take two engagements a quarter, so if it is not a fit we will say so on the call.