CitedWell

Half the time a brand makes the list, it does not become the pick

Our scoring engine tags every brand mention with a type. Some brands get named as the explicit recommendation. Others show up in a numbered or ordered comparison list without being called out as the winner. We re-scored 270 live audit panels this session and pulled every mention that landed in a list. Across all four engines, only 49.1% of those listed appearances turned into the actual recommendation. The rest just made the list.

What we measured

Our scorer, engine/scorer.ts, classifies every brand mention in a response into one of four types: recommended (explicitly called out as the pick), compared (named at a specific position in a list, but not the pick), mentioned (named with no list position), or absent. We recomputed this classification fresh this session across 270 live panels (non-fixture, real grounded engine calls across ChatGPT, Claude, Gemini, and Perplexity) in three B2B SaaS categories: project management, customer support, and HR software.

We pulled every mention that carried a list position, meaning the response put that brand at a specific rank inside an enumerated comparison, and checked what fraction of those were the recommended type rather than just compared. That gave us 1,675 list-position mentions total: 823 recommended, 852 compared but not chosen.

Making the list converts to being the pick about half the time

Category List-position mentions Became the recommendation Conversion rate
HR software 160 107 66.9%
Project management software 781 400 51.2%
Customer support software 734 316 43.1%
Overall 1,675 823 49.1%

HR software converts at nearly two-thirds. Customer support converts at under half. A brand that lands in a numbered comparison in customer support software is more likely to be passed over for the actual pick than to become it. The same brand landing in an HR software list has real odds of walking away with the recommendation. Category changes what a list slot is worth, not just how often you get one.

The pattern holds inside the two engines that produce enough list-style answers to check

Not every engine answers in numbered lists often enough to make this a fair per-engine comparison. Of the 1,675 list-position mentions, 1,087 came from ChatGPT and 464 from Gemini. Claude produced only 119 and Perplexity just 5, because those two engines mostly answer in prose without an explicit ranked list, even when several brands are named in the same response. We are not reporting a Claude or Perplexity conversion rate here; the sample is too thin to mean anything at the per-engine level.

Engine List-position mentions Became the recommendation Conversion rate
ChatGPT 1,087 616 56.7%
Gemini 464 110 23.7%

Where we do have enough volume to compare, the gap is large. On ChatGPT, more than half of listed brands go on to be the pick. On Gemini, fewer than a quarter do. A brand that only tracks its overall list-appearance rate would read these two outcomes as the same kind of win. They are not. Landing in a Gemini comparison list is a much weaker signal of an eventual recommendation than landing in a ChatGPT one.

Appearing in an AI's numbered comparison is not the finish line. Across two engines with enough volume to check, less than a quarter to a bit over half of those list appearances actually convert to being the recommendation, and which end of that range you land on depends on the category and the engine, not on being on the list at all.

What this means for tracking visibility

A visibility score that only counts whether a brand shows up in a list is measuring the wrong finish line. Two brands with an identical list-appearance rate can have very different outcomes if one converts at 67% and the other at 43%. The gap between compared and recommended is where the real buying signal lives, and it is invisible unless the tracking distinguishes the two mention types instead of lumping every list appearance together as a win.

Methodology note

Data recomputed this session by running engine/scorer.ts against results.jsonl from all 270 live (non-fixture) teaser audit panels in ops/state/engine-runs, three B2B SaaS categories: project management, customer support, and HR software, June to July 2026. Mention type classification (recommended / compared / mentioned / absent) and list position are both computed fields in the scorer, not persisted separately, so the count above reflects a fresh scoring pass rather than a stored value.

Want to know how many of your AI list appearances are actually converting to the recommendation?

Get an AI visibility audit -- $490