Last updated: 13 August 2026
Every entry in the AI Visibility Logs follows the same procedure. This page describes it once, so the logs themselves don’t have to repeat it.
What these tests measure
A log records what AI engines say about a company when asked, and which sources they cite to say it.
Browsing is left on. That makes these retrieval tests, not recall tests, they measure what the engines surface about a brand right now, not what’s encoded in the models. That’s deliberate: it’s the condition an actual prospect is in when they ask about your business.
The account
Every run uses a separate account created for testing, with no custom instructions, no saved memory, and no chat history.
This matters more than it sounds. Private browsing is not a clean test:
- Browser incognito controls cookies and local history. It does nothing to the settings on a logged-in account.
- ChatGPT’s Temporary Chat stops a conversation being saved or used for training, but it still applies custom instructions and personalization.
An early version of this series was run in Temporary Chat and returned an answer about a third-party company that referenced my own framework by name. A dedicated clean account is the only condition I trust.
Before each session, every personalization surface is opened and verified individually, not assumed: custom instructions, saved memory, reference chat history, reference recent conversations, and any account-level personalization. Each engine is then asked what we discussed recently; a clean account has no answer to that.
Two settings have leaked in this series despite appearing to be off. Both changed a result.
One test account per engine, reused across logs. Creating a new account each time would introduce differences in tier and default settings. What matters is that it stays clean: no other use, no follow-ups, a fresh session for every run.
The settings
Free tier, default settings, throughout. Pinning a specific model on most platforms requires a paid plan, so model selection is left on whatever each engine chooses by default, and interface modes are left as they arrive.
This is a deliberate choice, not a shortcut: it describes what a default user sees, not what each engine can do at its best. Where an engine’s model or mode is visible, it’s recorded in that log’s conditions table.
Free-tier accounts can silently upgrade a query. Perplexity includes an allowance of deeper Pro searches; when one fires, it is noted at the time, because the interface does not record it afterwards.
The query
One query per log, asked verbatim and unchanged in a fresh session on each engine. No follow-ups, no clarifications, no rephrasing between engines.
Queries are written flat. A question that assumes an answer, what single framework does this company specialise in, forces each engine to pick something, and then you observe them picking differently. That’s the question manufacturing the disagreement you set out to measure. Queries are phrased the way a prospect would type them.
What gets recorded
- The full response from each engine
- Every source URL the engine displayed, with tracking parameters stripped
- Model and mode where the interface shows them
- Date of the run
Sources are the primary data. Which page an engine pulled is usually more informative than what it concluded from it.
Retrieved and cited are different numbers. Where an interface distinguishes them, both are recorded. Perplexity reports results retrieved separately from sources cited. Gemini’s citation chips do not survive copy-paste and must be opened individually. Claude and ChatGPT display cited sources only. Source counts are not comparable across engines without saying which is being counted.
Where a missed fact is under investigation, the subject’s own pages are audited for whether it appears at all.
Verification
Contested claims are checked against primary sources before publication. Where an engine states something inaccurate, incomplete, or invented, it’s recorded as such, with the correct version and a link.
Where an engine misses something, the subject’s own site is checked before that counts as a failure. A fact that exists only on a page nothing links to is not the same as a fact an engine overlooked, and the difference changes what the result means. What gets recorded: whether the fact appears in the title, meta description, headline or body of the pages an engine would plausibly read, or only somewhere deeper.
Limitations
Every log publishes its own limitations alongside its results. Common ones:
- Run count. A single run can’t distinguish a genuine engine difference from ordinary variance. Where a log has one run per engine, it says so, and its findings are directional.
- The web is a variable. Retrieval results describe a date. Sources change; answers move with them.
- One subject at a time. A company with a large content library and entries in several databases will produce a different pattern from a business with a thin site.
- Publishing is a feedback loop. These logs are themselves pages about the companies they test.
Corrections
Logs are updated in place rather than replaced, and every change is dated in the entry’s revision history. Earlier results stay on the page, what changed between two dated runs is the itself data.
The standard here is that a result is publishable when it’s honest about its conditions, not when it’s airtight.
Spot an error in a log, or run the same test and get a different result? Email us – support@vishnuharshan.com