The number looks good until you inspect the answer. In roughly 40% of cases where an AI engine cited a brand, the dashboard counted a win while the reader saw somebody else's name or no name at all.

A healthy-looking miss.

That gap matters because AI visibility has three rungs. Cited means the engine linked your page. Named means your brand showed up in the answer a reader saw, which is a different event and a much rarer one. Recommended means the answer told the reader to use you. We can measure the gap from cited to named. Nobody can measure named to recommended across the market yet.

Every tool in this category makes it easy to read rung one and stop there. Ours included.

Kevin Indig coined the term "ghost citation" in Growth Memo, across 3,981 domains, 115 prompts, 14 countries and four engines. Nikki Lam at NP Digital and I ran a separate study of roughly 16 million brand appearances for Search Engine Land, checking whether cited brands were also named in the generated answer.

Published ghost rates vary a lot between studies. Prompt sets differ, engine mixes differ, and the denominator differs most of all. Ours came out at 40%, and I'll only defend that number with the method attached to it. Treat any single ghost rate, including this one, as a reading rather than a constant.

A citation without a name is evidence, not recognition.

Across the study, roughly 40% of AI citations didn't name the cited brand in the answer. Put another way, a brand could supply the evidence and remain invisible in the prose built from it.

Two answers can cite the same page and land completely differently on the reader.

Named: Writesonic's analysis of roughly 16 million brand appearances found that approximately 40% of AI citations didn't name the source brand.

Ghosted: One analysis found that approximately 40% of AI citations didn't name the source brand.

Both could link to the same page, and both are accurate, but only the first one tells the reader who did the work.

A study about ghost citations can get ghost-cited too.

A citation still has value. It tells you the engine found your page relevant enough to use, and it puts a link within reach of anyone who checks sources. It just doesn't prove recognition or preference.

Many teams roll citations and mentions into one celebratory number, then count a retrieval event as recognition without checking what the answer told the reader. That's rung inflation: reporting rung one as if it were rung three.

My read: Recognition is the result. An invisible citation is unfinished work.

At Writesonic, we track two separate co-primary metrics with two different denominators. AI Visibility is the percentage of tracked answers that mention your brand. Citation Share is the percentage of tracked answers that cite at least one page from your domain. We kept them apart because they're different events. This study is the evidence at scale for how far apart they are.

Each week I take one finding from our own data and work through the reporting or content decision it changes. Subscribe to get the next one.

A 19% to 52% spread makes one blended visibility score misleading.

When you average the results, you hide a much messier engine-level story. Ghost citation rates ranged from 19% to 52%, which changes what an identical citation count means.

The full order (less tidy than the average) was:

  • Perplexity 52%

  • Google AI Mode 49%

  • AI Overviews 41%

  • ChatGPT 37%

  • Gemini 25%

  • Grok 22%

  • Copilot 19%

Some engines show the source and keep the brand out of the answer. Others put the brand in the prose.

The full study, co-authored with Nikki Lam at NP Digital, is here: The ghost citation problem in AI search. It has the complete engine breakdown.

Blend the engines and you erase that behavior. If your gains came from an engine with a high ghost rate, rung one may be rising much faster than rung two. The headline score improves, but recognition barely moves. The average is where rung inflation hides.

My read: I don't trust a blended score anymore. Not with a 33-point spread underneath it.

The engine split reflects consistent behavior. Perplexity and Google's engines behave more like citers: they link more and name less. Gemini and Copilot behave more like namers: they name more and link less. ChatGPT and Grok sit between those patterns.

The distinction is blunt (on purpose). A citer puts its evidence in links. A namer writes the brand into the sentence.

Citation share and mention share therefore measure different things. On a citer, your job is getting your name into the sentence the engine builds from your page. On a namer, your job is the context it names you in.

Take two brands with 100 cited appearances each and apply the study's published ghost rates. At a 52% ghost rate, 100 citations leave 48 named appearances. At a 19% ghost rate, the same 100 citations leave 81.

If you look only at citations, you call this a tie. Readers encounter one brand in the answer 48 times and the other 81 times, so the input metric matches while recognition differs by 33 appearances.

One of those brands is losing and can't see it.

This connects to the positioning problem I wrote about in issue 01. If AI names what your company used to be, readers get the wrong identity, so it doesn't count as a win.

This example is arithmetic, not another observed sample. It shows how the engine rates change the meaning of a familiar metric. Count cited appearances. Then count how many became named.

My read: Citation totals aren't results. They're the denominator.

Your own domain can't carry the whole strategy.

The next mistake is assuming you can close the gap through your own site alone. In a separate analysis of 2,007 brands, YouTube had 3.19x citation influence, Reddit 3.17x, Gartner 3.15x, Amazon 1.46x, Wikipedia 1.42x, and a brand's own website 1.30x.

In our post, we put citation influence in plain language: a mention on Reddit makes a brand about 3.17x more likely to appear in an AI-generated answer. We didn't break that lift down by engine, so those multipliers don't explain the citer-versus-namer split.

The pattern is uncomfortable. We only counted a source if at least 600 of the 2,007 brands cited it, which left fourteen that were broadly trusted rather than quirks of one category. A brand's own website scored below all fourteen, marketplace listings included.

Reddit was cited for 1,984 of the 2,007 brands in that sample. Nearly every brand in the study got cited from a place they don't own. That's why an owned-site-only plan falls short.

Your website remains the canonical source for what you do. We rebuilt ours because owned pages need speed and control, as I explained in issue 02, but a better site can't create independent support.

A good site is the floor, not the ceiling.

A company can call itself useful. When customers and independent sites repeat the claim, engines have stronger grounds to name it, and that's the part no company can manufacture alone.

My read: The strongest case is the one you don't have to make alone.

Being named still leaves one whole rung before recommendation.

This study measures the gap between cited and named. It doesn't measure whether a named brand was recommended. A named appearance isn't a recommendation. We don't yet have data across the market for that third rung, and I don't want to smuggle a recommendation claim into a mention metric.

Nobody can tell you your recommendation rate yet. Including us.

A name can appear in several ways: the engine may list a brand as one option, describe it without judgment, warn against it, or choose it for the reader. Those outcomes all count as named, but the brand only reaches recommended when the engine picks it.

Rung three requires reading the recommendation, its context, and the reason the engine gave. Rung inflation makes its biggest leap here. A source link or brand mention becomes implied buyer preference even though the data doesn't support that jump.

We can inspect recommendation quality manually for important prompts and track whether the answer positions a brand for the right use case, but we can't publish a sound recommendation rate across the market from this dataset.

I'd rank the work by the rung it improves, not the metric it flatters.

Here's what I'd change in our own reporting first. Start with measurement, then do the source work that can change what engines retrieve and say.

  1. Split every report by engine. Show cited appearances, named appearances, and the conversion between them. A blended score can sit at the bottom, but it shouldn't drive the decision.

  2. Audit the largest ghost gaps first. Find prompts where your page gets cited and your brand stays absent. The engine already trusts the page. Your name just isn't on it.

  3. Inspect the passage the engine used. Make the cited passage connect your brand to the category claim and the evidence behind it. The name has to survive the jump from retrieval into prose.

  4. Build support beyond your domain. Start with YouTube at 3.19x and Reddit at 3.17x. Then pick the channels and subreddits your buyer reads.

  5. Review recommendation prompts manually. Read the full answer. Record whether you're absent, named, shortlisted, or explicitly chosen, plus the reason. Keep that work separate from the broader metric until the measurement is sound.

The last rung remains open, and we'll measure it when we can do it honestly.

Until then, keep rung inflation out of the report. Show cited, then named, then recommended. Don't report the first rung as the third.

We built the cited-versus-named split into Writesonic's reporting because we kept needing it ourselves. You can run step two by hand without us, and plenty of teams should. What matters is that somebody runs it.

Forward this to whoever on your team owns the AI visibility number, because they're the person who has to defend it in the next review.

PS. Pull up your last citation report and check one thing: can you tell how many of those citations put your name in the answer? Reply and tell me whether your tooling could answer that.

Reply

Avatar

or to participate

Keep Reading