One AI engine goes null on agency buyers 8 times out of 10. Another almost never does.
An earlier post on this blog measured the overall null rate: how often an AI names none of the six brands a panel tracks on an organic, buyer-intent prompt. The overall number was 7.4% across 6,616 real responses. That number hides which engine is actually producing it. Split by engine, the gap is not small. ChatGPT goes null 0.5% of the time. Claude goes null 12.8% of the time, over 25 times more often, and on one specific prompt qualifier the gap opens even wider.
The null rate is not one number, it is four
Same definition as before: a null response is an organic (non-branded) prompt where zero of the panel's six tracked brand names, the target plus five named competitors, appear anywhere in the answer. We re-scored every real, search-grounded response across our 270 live audit panels (project management, customer support, and HR software; ChatGPT with web search, Gemini with grounding, Perplexity Sonar, and Claude with web search) using CitedWell's production scoring logic, and grouped the result by the `engine` field on each scored response.
| Engine | Organic responses | Null | Null rate |
|---|---|---|---|
| ChatGPT (OpenAI) | 973 | 5 | 0.5% |
| Gemini | 1,779 | 41 | 2.3% |
| Perplexity | 1,815 | 179 | 9.9% |
| Claude | 2,049 | 263 | 12.8% |
ChatGPT almost never returns a response naming none of the six tracked brands. Claude does it more than 1 in 8 times. That is a real spread, but it is not spread evenly across categories either. It concentrates almost entirely in one category and, inside that category, in one specific kind of prompt.
The gap is a project management software problem
| Engine | PM software | HR software | Customer support |
|---|---|---|---|
| ChatGPT | 0.0% | 0.0% | 1.2% |
| Gemini | 0.7% | 6.1% | 6.0% |
| Perplexity | 12.6% | 4.0% | 2.0% |
| Claude | 19.7% | 0.0% | 0.5% |
Claude's overall 12.8% null rate is almost entirely a project management number. It is exactly 0.0% in HR software and 0.5% in customer support, but 19.7% in project management. Perplexity shows the same pattern at a smaller scale: 12.6% in PM versus 2-4% elsewhere. Gemini runs the opposite way, slightly worse in HR and customer support than in PM. ChatGPT stays low everywhere. Whatever is causing Claude and Perplexity to go null lives specifically inside project management prompts, and an earlier post already identified the likely cause at the category level: one use-case qualifier.
Inside project management, it is one qualifier and two engines
Project management panels run recommendation-intent prompts in three forms: plain, "...for remote teams," and "...for agencies." Splitting the null rate by both qualifier and engine shows the category number is really two engines' number on one qualifier.
| Engine | Plain | "...for remote teams" | "...for agencies" |
|---|---|---|---|
| ChatGPT | 0.0% | 0.0% | 0.0% |
| Gemini | 0.3% | 0.3% | 0.3% |
| Perplexity | 1.8% | 12.4% | 35.2% |
| Claude | 0.0% | 0.0% | 79.6% |
Claude's plain and "for remote teams" null rates are both a flat 0.0% (0 of 350 and 0 of 362 responses). Its "for agencies" null rate is 79.6% (261 of 328 responses). Every single one of Claude's project management nulls in our entire dataset comes from that one qualifier. Gemini and ChatGPT show no such cliff. Gemini sits at 0.3% across all three forms, essentially flat. ChatGPT is 0.0% across all three. Perplexity shows the same direction as Claude, a real jump from 1.8% plain to 35.2% on "for agencies," but nowhere near as steep.
Why this changes where you look first
If a brand's own audit shows a null-heavy result on an agency-qualified prompt, the fix is not "our AI visibility has a hole in it." It is "one or two of the four engines we measure route this specific question to a different competitive set entirely, and the other engines do not." A brand fixing this by publishing more generalist agency-use content and then re-checking a blended, four-engine score will see the number barely move, because ChatGPT and Gemini were never the problem and Claude and Perplexity are chasing a different market's retrieval sources, not a content gap on the brand's own site.
The practical read: before spending on a fix for a null-heavy prompt, split the result by engine first. A null rate concentrated in one or two engines calls for a different, narrower response than a null rate spread evenly across all four, and a blended score cannot tell you which situation you are in.
See your own null rate broken down by engine, not just blended into one number.
Get an AI Visibility Audit, $490Methodology
Data drawn from 270 live, search-grounded audit panels (project management, customer support, and HR software brands), each run across four AI engines: ChatGPT with web search, Gemini with grounding, Perplexity Sonar, and Claude with web search. We re-scored every successful (non-error) response using CitedWell's production scoring logic (engine/scorer.ts), restricted to organic (non-branded) prompts, and grouped the result by the engine field CitedWell records on every scored response. A response counted as null when it contained zero mentions of any of the panel's six tracked brand names. 6,616 organic responses total across all four engines. No development-rail or fixture data is included; all responses came from live engine calls. Data collected June-July 2026.