Ask four AI engines about the same company. Do they agree? (log 001)

AI Visibility Logs

what AI says about a company
Last updated: 12/08/2026

Tested on Reforge, which describes itself as a professional education platform teaching product and growth to experienced tech operators, across Claude, ChatGPT, Gemini, and Perplexity.

Bottom line: an AI engine reads a handful of your pages, and those pages decide what it says about you. Three of the four tested here still call Reforge independent in most runs, four months after it was acquired.

The query

Asked verbatim, unchanged, in a fresh session on each engine: What is Reforge, and what are they best known for?

No URL. No context. No follow-up. The way a prospect would type it.

Test conditions

  • Log reference ID: AIVL-001 | Run 1 | 30/07/2026
  • Topic category: AI Retrieval & Source Selection
  • Target subject: Reforge (brand name only, no URL supplied)
  • Engines tested:Claude (Anthropic), ChatGPT (OpenAI), Gemini (Google), Perplexity
  • Model versions: Claude: Sonnet 5. Gemini: Flash. ChatGPT and Perplexity: auto model selection.
  • Modes: Perplexity- Search (default). All engines on default interface settings.
  • Test environment: Separate account, no custom instructions, no saved memory, personalization off
  • Runs per engine: 1
  • Date of runs: 30 July 2026
  • Status: Updated 12 August 2026 — ChatGPT re-run with memory disabled; acquisition figures revised.

Sources cited

Engine Citations Pages Mix Sources
Claude 6 6 1 owned / 5 databases reforge.com · Crunchbase · Tracxn · PitchBook · CB Insights · ZoomInfo
ChatGPT 5 2 2 owned reforge.com (×4) · Brian Balfour’s profile page
Gemini 4 4 2 owned / 2 secondary reforge.com · Brian Balfour’s profile page · MOGE product overview · a personal blog post on Balfour’s ideas
Perplexity 5 5 1 owned / 4 third-party reforge.com · PitchBook · LinkedIn · VentureCapitalTracker · a 2022 Series B press release

Every URL above is what the engine itself displayed. Nothing was added.

I assumed owned sources would beat third-party ones, but Claude’s mix was more database-heavy than Perplexity’s, and its answer was far better, with the acquisition, founding year and headquarters all correct. And ChatGPT produced the most detailed answer of the four off just two distinct pages.

What the source class actually determines is which part of the question an engine can answer. Investor databases carry firmographics: founded, based, funded, and acquired. Crunchbase and CB Insights track ownership changes as a matter of course, which is the likeliest reason Claude had the Miro deal, while carrying nothing about what a company teaches, which is why Claude was thin on frameworks and dropped a co-founder. Owned pages carry the opposite.

Perplexity’s weakness wasn’t third-party sourcing. It was a weak set within that class: a stale funding press release and a low-authority tracker doing work a maintained database would have done properly.

Source class shapes the blind spots. Source quality shapes the accuracy. Neither is about which model you asked, or how many pages it read.

The finding I didn't expect

Gemini cited Brian Balfour’s profile page on reforge.com. So did ChatGPT. It’s where ChatGPT found the Miro acquisition. Gemini read the same page and didn’t report it.

That’s the most useful thing in this test. Retrieval is not comprehension. An engine can fetch the page carrying your most important fact and still not surface it. Which means the gap between what you have published and what gets said about you isn’t only a publishing problem; you can put the fact on the page, watch an engine cite that exact page, and still get an answer that omits it.

There is no lever that fixes this from your side. Publishing well is necessary. It is not sufficient. (The second round below revises this.)

Where all four agreed

  • Category. Education and career development. Four for four.
  • Audience. Experienced, mid-to-senior technology professionals. Four for four.
  • No single flagship framework. Not one engine named one, even unprompted.

Where they diverged

  • The Miro acquisition: two of four. Claude and ChatGPT have it. Gemini and Perplexity don’t.
  • Andrew Chen as co-founder: one of four. Only ChatGPT.
  • Founding year: three of four. Perplexity omitted it.
  • Detail depth: wide. ChatGPT named a dozen frameworks; Gemini named several and listed instructor companies. Perplexity gave four sentences with no founder, no year and no frameworks.
  • Category wording varied. “Career development and education company” (Claude), “professional learning platform” (ChatGPT), “elite career accelerator and executive education platform” (Gemini), “professional education and career development company” (Perplexity).

One fabrication

Gemini referred to Reforge’s frameworks as the “Reforge Playbooks” in quotation marks, as though it were the company’s own term.

It isn’t. Reforge uses “playbook” freely as a common noun and has a Growth Playbooks collection on its blog, but the branded name for its member assets is Artifacts. Gemini named Artifacts correctly two paragraphs later.

So it invented a brand term while simultaneously knowing the real one. No reader would catch it. That’s what makes it worth logging.

Possibly relevant: two of Gemini’s four sources were an AI-generated product summary and a personal blog paraphrasing Balfour’s ideas. Neither is the company. Loose paraphrase upstream is a plausible route to an invented label downstream, though with one run, I can’t establish that, only note it.

Fact-check

I verified the contested claims rather than trusting any engine.

  • The Miro acquisition is real. Announced 24 March 2026. Balfour joined Miro as Chief Growth Officer.
  • Reforge Learning does continue as a standalone brand, ChatGPT’s version was the precise one.
  • Andrew Chen co-created the first Growth Series. The engines that credited Balfour alone were incomplete.
  • 2016, San Francisco: correct.

No engine flagged that its picture of the company might be out of date.

What this means if you are the brand being asked about

Two source classes describe you, and you need both to be maintained. Your own pages supply what you do and who for. Company databases supply the facts of the entity, founded, based, funded, and owned. Engines draw on whichever they reach, and an engine reading only one class inherits that class’s blind spots.

Your database listings are marketing assets. Three engines cited Crunchbase, PitchBook, ZoomInfo, CB Insights, or Tracxn. Almost nobody maintains those. They are describing you as they last knew you, and they get quoted back to your prospects.

Nothing propagates evenly. Half the engines missed the single biggest development in the company’s history, four months on. If you’ve rebranded, been acquired, or changed what you sell, assume some engines still have the old version.

There is no single answer about you. The earlier attempt at this test, run on a personalized account, returned an answer about Reforge that was partly shaped by who was asking. Your prospects all have their own history, instructions, and saved context. “How do we show up in ChatGPT?” doesn’t have one answer; it has as many as there are people asking.

What that adds up to. Go and look at your own database listings: Crunchbase, PitchBook, ZoomInfo, CB Insights, Tracxn. Most are claimable and editable, and most companies have never touched theirs. Then check that the facts you would want quoted are written plainly on your own pages, team bios included. That’s the part you control.

The part you don’t: an engine can cite the right page and still not surface what’s on it. Publishing well is necessary. It isn’t sufficient.

Limitations - Run 1

  • One run per engine. Non-determinism is real. I cannot distinguish “Gemini doesn’t know about the acquisition” from “Gemini didn’t mention it this time.” Two engines already showed me the same query returning different answers on different runs. (Resolved in the second round below)
  • Model tiers aren’t matched, by design. Pinning a specific model on ChatGPT or Perplexity requires a paid plan, so both were left on automatic selection — which is what an unsubscribed user gets. Gemini ran on Flash for the same reason. The tradeoff: Gemini was also the engine that fabricated a brand term and missed the acquisition, so some of what reads as “Gemini” here may be “Flash.” And because auto selection can route differently per query, the model itself may vary between runs without appearing in the interface. These tests describe what a default user sees, not what each engine is capable of at its best.
  • All four engines retrieved. Every engine cited sources, so nothing here measures unaided recall. This is a retrieval test throughout.
  • Retrieval measures the web on a date. These runs describe 30 July 2026. Reforge publishes; the answers will move.
  • One subject. A company with a strong owned content library. A business with a thin site would likely produce a different pattern entirely.
  • Publishing this is itself a variable. This page is now a source about how engines describe Reforge.

Runs 2–4 — 6 August 2026

Three runs per engine, one week after the first. Same query, same clean account, same default settings.

The point of repeating it was to find out which of the original findings were real and which were single-run noise. Two didn’t survive. One got considerably stronger.

Test conditions

  • Log reference ID: AIVL-001.1
  • Query: Unchanged from Run 1
  • Runs per engine: 3 (12 total), plus 3 further ChatGPT runs on 12 August
  • Date of runs6 August 2026; ChatGPT re-run: 12 August 2026
  • Models: Claude: Sonnet 5. Gemini: Flash. ChatGPT and Perplexity: auto.
  • Test environment: Same clean account. Fresh session per run. ChatGPT memory found enabled on 6 August; disabled and re-run on 12 August.

What held, across all four runs

  Claude ChatGPT Gemini Perplexity
Education / career development category 4/4 4/4 4/4 4/4
Mid-to-senior tech audience 4/4 4/4 4/4 4/4
Growth and product as subject matter 4/4 4/4 4/4 4/4

Every engine, every run, same answer on who the company is and who it’s for. No engine invented a different category. No engine got the audience wrong.

What didn't

  Claude ChatGPT Gemini Perplexity
Miro acquisition reported 4/4 1/3 (see below) 0/4 0/4
Founded 2016 4/4 1/4 2/4 1/4
Andrew Chen credited 2/4 2/4 0/4 0/4
San Francisco 4/4 0/4 0/4 1/4

The acquisition row is the striking one. Claude has it every time. Gemini and Perplexity have never had it, across four runs and eight days.

ChatGPT’s number needs an explanation, and it turned out to be the most useful result in this round.

The memory problem, and what it showed

After the 6 August runs I checked the test account by asking each engine what we had discussed recently. Three said they had no access to previous conversations. ChatGPT summarised the earlier sessions. Its memory had been on the whole time.

That matters because the account had been asked about Reforge four times across two dates. ChatGPT had been told about the Miro acquisition in those sessions, and had stored it. So its 4/4 wasn’t retrieval; it was partly recall of my own earlier tests.

I turned memory off, verified it with the same probe, and re-ran the identical query three times on 12 August. The acquisition dropped to 1 of 3.

Then the useful part. Here is what each clean run retrieved:

Run Pages retrieved Acquisition reported
Run 1 reforge.com, Growth Foundations course, courses index, AI Product Leadership course No
Run 2 Brian Balfour’s profile page, Growth Foundations course, courses index, a Reddit thread Yes
Run 3 reforge.com, Growth Foundations course, courses index, AI Growth course, pricing page No

Same engine. Same query. Same day. Same session type. The one run that reached the founder profile page is the one run that reported the acquisition. The two that read only course and pricing pages did not.

That’s the finding of this round, demonstrated inside a single engine rather than inferred across four. A fact appears in the answer if, and only if, retrieval happens to touch a page carrying it.

Everything else remains unstable. ChatGPT gave the founding year on 30 July and in one run since. Claude omitted Andrew Chen on 30 July and named him in two of three runs after. Perplexity produced a founder, a year and a city in exactly one run out of three.

Sources cited — 6 and 12 August

Engine Unique sources across 3 runs Recurred in all 3 Reached an about, news or founder page?
Claude 3 (run 1 only; runs 2 and 3 displayed none) No
ChatGPT (6 Aug, memory on) 6 reforge.com root, Brian Balfour’s profile page, one Reddit thread Yes — founder profile, all 3 runs
ChatGPT (12 Aug, memory off) 8 Growth Foundations course, courses index Once — founder profile, 1 of 3 runs
Gemini 12 reforge.com root No
Perplexity 7 cited, from 15 retrieved a16z funding announcement (2 of 3) No

The last column is the whole result.

Every source, run by run — 6 and 12 August 2026

Claude

ChatGPT — 6 August, memory on

ChatGPT — 12 August, memory off

Gemini

  • Run 1: reforge.com (×3) · /courses/product-led-growth · /courses/mastering-product-management · /courses/product-strategy · /teams
  • Run 2: reforge.com · /teams · /course-categories/career-development · /courses/product-management-foundations · uxcel.com course review
  • Run 3: reforge.com (×3) · /course-categories/growth · /courses/growth-series/details · yourstory.com (×2) · techscaler.co.uk

Perplexity — cited across the three runs

Perplexity — the full 15 results retrieved, of which 5 were cited

The finding that changed

Run 1 said: two engines cited the same page, and only one reported the acquisition from it. That looked like an extraction failure; the fact was on the page, one engine surfaced it, one didn’t.

Four runs of source data say something different, and more useful.

Gemini’s twelve unique sources across three runs are course pages, course-category pages, the teams page, a training aggregator and two company-profile blogs. reforge.com root recurs; nothing else owned does. No about page, no news post, no founder profile.

Perplexity retrieved fifteen results and cited five. Of the fifteen: PitchBook, ZoomInfo, Tracxn (twice), LinkedIn, VentureCapitalTracker, Craft, Owler, Built In, SignalHire, Levels.fyi, a 2020 a16z funding announcement, two blogs, and exactly one owned page, reforge.com root.

So neither engine failed to extract the acquisition. Neither one ever retrieved a page that mentions it. They aren’t misreading the company’s corporate pages; they’re not reaching them at all. Gemini reads the catalogue. Perplexity reads the company registry.

That’s why the miss is stable rather than random. It isn’t a bad run. It’s where the engine looks.

And ChatGPT, once memory was off, reproduced the same pattern within itself: the run that reached the founder profile reported the acquisition; the two that stayed in the course catalogue did not.

Claude is the one engine this can’t be checked against. It reported the acquisition in every run, but displayed sources in only one of the three on 6 August, and that run’s sources were a blog, a marketing post, and CB Insights, none of which is where the news lives. Whether it retrieved the fact or already held it, the interface doesn’t say.

Two smaller results

The fabrication didn’t recur. Gemini’s “Reforge Playbooks” presented in Run 1 as though it were a branded term, appeared in one run out of four and never again. Recording it as a one-off, not a pattern.

Perplexity retrieved a different company with the same name. One of its fifteen results was a Tracxn profile for an unrelated business also called Reforge. Nothing in the answer suggests it was noticed. For any company without a distinctive name, that’s worth knowing about.

What this adds for a brand

Your directory listings aren’t supplementary. For one of these four engines they were fourteen of fifteen retrieved results. PitchBook, ZoomInfo, Owler, Built In, Levels.fyi – most companies have never looked at theirs.

Product pages don’t carry corporate facts. Gemini read Reforge’s course catalogue thoroughly and accurately, and consequently described a company that had been acquired four months earlier. If your news lives only on a news page, engines anchored to your product pages will never see it.

Don’t test your AI visibility on your own account. ChatGPT reported the acquisition in every run while its memory was on, and in one run out of three once it was off. It had learned the fact from my earlier sessions and was repeating it back. If you check what AI says about your company while logged into the account you use every day, you will get a more flattering answer than your prospects get.

Stability cuts both ways. The engines that had the acquisition through retrieval had it consistently. The ones that didn’t – didn’t, for eight days running. A wrong answer about your company is not usually a fluke you can wait out.

Limitations - Run 2

  • ChatGPT’s 6 August runs were contaminated. Memory was enabled without my knowing, on an account that had already been asked about Reforge. Those runs are superseded by the 12 August re-runs with memory off. Claude, Gemini and Perplexity were checked with the same probe and showed no carryover. This is the second personalization setting to leak in this series; both had appeared to be off.

  • Three runs is enough to separate stable from unstable, not enough to quantify either. 2/3 and 3/3 are different; 2/3 and 2/3 are not meaningfully different.

  • Source visibility is not comparable across engines. Perplexity distinguishes results retrieved from sources cited. Claude and ChatGPT show inline citations only. Gemini’s citation chips do not survive copy-paste and had to be opened individually. Two Claude runs displayed no sources at all, which may mean no search fired or may mean it wasn’t shown.

  • Eight days is not a trend. Two dated rounds show what changed between two dates.

What I got wrong the first time

I ran this test once before and didn’t publish it. The question was leading; it asked which single framework Reforge specialises in, which forced each engine to pick one and made the picks look like disagreement. And the session was contaminated: it ran in ChatGPT’s Temporary Chat, which still applies custom instructions, so an answer about Reforge came back partly addressed to my own positioning (more on why).

Fix both, and most of the disagreement evaporates. A less dramatic finding than the one I started with, and the one the evidence supports.

What's next for this log

Log 002 is live. Does AI answer differently when you give it a URL instead of a name? – The same question, asked with the URL in place of the company name. One engine’s retrieval changed almost entirely; another’s didn’t move at all.

Revision history

30 July 2026 — Published. One run per engine, brand name only.

6 August 2026 — Added three further runs per engine. Two findings from the first round did not survive repetition: Andrew Chen’s attribution was run variance rather than an engine difference, and Gemini’s “Reforge Playbooks” appeared once in four runs. The same-page finding was superseded by source data showing that the two engines which miss the acquisition never retrieve a page carrying it. Original 30 July results retained above.

12 August 2026 — ChatGPT’s memory was found enabled during the 6 August runs, on an account that had already been asked about Reforge four times. Re-ran the same query three times with memory disabled: the acquisition figure fell from 4 of 4 to 1 of 3, and the single run reporting it was the one that retrieved the founder profile page. ChatGPT’s 6 August figures are superseded and retained above. Claude, Gemini and Perplexity showed no carryover.

Part of the AI Visibility Logs, an open record of how AI systems find, interpret, and describe brands. How AI Visibility tests are run.

Scroll to Top