Cited, Mentioned, Recommended: The AI Visibility Ladder
Being cited by an AI and being recommended by it are different outcomes with different causes. The four rungs, what governs each, and why most AI visibility reporting measures the wrong one.
The four rungs
AI visibility is a ladder, not a yes or no. Each rung is governed by something different, which is why a brand can be cited constantly and never recommended.
| Rung | What it means | What governs it |
|---|---|---|
| 1. Retrieved | A model read your page while building an answer and did not cite it | Crawlability and parseable structure |
| 2. Cited | Your page appears as a source in the answer | Content usefulness: structure, statistics, clarity, freshness |
| 3. Mentioned | Your brand is named in the answer text | Entity recognition, plus how the web talks about you |
4. Recommended | Your product is on the shortlist the buyer actually considers | Aggregate web consensus: reviews, forums, analysts, press, video |
There is a fifth state most reporting ignores. On detailed, requirements-heavy questions, models increasingly name products a buyer should avoid for a specific use case, with sources. Weak third-party consensus is no longer just absence from the shortlist. It can be an explicit rule-out.
The short version: citation is earned on your site. Recommendation is earned off it.
Why this distinction costs people money
Most AI visibility tools report rungs 2 and 3 and present the number as “AI visibility.” A brand watches its citation count rise, concludes the programme is working, and never notices that the answers citing it are recommending someone else.
That failure has been measured. Lily Ray analysed 100 B2B “best [category] software” queries across 15 April, 15 May and 8 June 2026. Eighty of them returned an AI Overview, and self-promotional listicles earned 323 citations across those answers. In 224 of them, 69%, Google cited the brand’s own page and left that brand out of the recommendations, pointing buyers to competitors instead.
The mechanism is worth understanding. The model treats your buyer’s guide as a source about the category. It extracts the competitor names, the comparisons and the evaluation criteria you compiled, then makes its recommendation from web-wide consensus, where established players dominate. For an emerging brand, a self-ranked buyer’s guide can function as a vote for your competitors: you did the research that helps the model describe them.
This splits sharply by stage. Established leaders get both outcomes, because analysts and review sites already validate them, so their guides earn citations and their brands get recommended. Emerging brands win the citation and miss the recommendation. That is not worthless, shaping how a model defines your category is real positioning work, but it is not the shortlist placement the tactic promises.
What a recommendation is actually worth
Two behavioural studies put numbers on the gap between rungs.
A June 2026 paper, From Prompt to Purchase, joined opt-in clickstream data to real ChatGPT, Claude and Gemini conversations and compared each user against matched backward placebos. When an assistant recommended a brand to someone with no recent engagement with it, that user’s same-name Google search rose 4.3 percentage points (95% CI 3.1 to 5.5), visits to the brand’s own site rose 2.4 points (1.4 to 3.5), and visits to the brand’s page on a retailer’s site rose 1.0 point (0.3 to 1.7). A passing mention did roughly half the work of a recommendation.
Similarweb tracked real user journeys for seven days after an answer. A brand an assistant recommended was 2.5 times more likely to receive a site visit in the following week, and those visitors engaged about twice as deeply as everyone else.
Both are observational rather than controlled experiments. The arXiv paper is the stronger of the two, because it reports absolute effects with confidence intervals against a matched baseline.
The attribution problem, and why you think AI is not working
Here is the finding that changes how you should read your analytics.
In the Similarweb data, 55.9% of AI-influenced traffic arrived through search, not as an AI referral. The assistant made the recommendation; the visitor then typed the brand name into Google and landed as ordinary branded organic traffic, indistinguishable from anyone else. For visits with no AI influence, the search share was 40.4%.
The arXiv paper measures the same effect from the other end. The largest single response to a recommendation was a branded Google search, at nearly twice the size of the direct site visit it also produced.
AI recommendations are already sending real, engaged buyers, and standard attribution reports almost none of it. If you judge AI visibility by the referral line in GA4, you are looking at the smaller half of the effect and concluding the channel is negligible.
This is also why “AI is only about 1% of traffic” understates the case. That figure counts visible referrals. It does not count the branded searches an AI answer caused.
How to measure the rung you actually care about
No single signal is complete. Three together give a reliable read.
- Prompt tracking. Whether and how you are mentioned or recommended, even when no click lands. Track the framing, not just the count: recommended, neutral, hedged, or recommended against.
- Self-reported attribution. A “how did you hear about us?” field catches buyers whose journey started in an AI chat and arrived via branded search.
- Sales call recordings. Buyers’ own language often reveals that an AI conversation shaped the shortlist long before any form was filled.
Watch branded search volume as a proxy. A sustained lift with no matching campaign is increasingly AI influence appearing under another name.
What earns each rung
| Rung | What moves it |
|---|---|
| Retrieved | Crawler access, server-rendered text, clean structure |
| Cited | Statistics with sources, direct answers, freshness, schema |
| Mentioned | Entity resolution: Wikidata, consistent sameAs, consistent descriptions |
| Recommended | Review platforms, analyst coverage, community discussion, earned media, video |
The test to apply before publishing another self-ranked guide: if a model ignored everything on our domain, would the rest of the web still put us on the shortlist? If not, that gap is the priority, and more content on your own site will not close it.
What this means for our own clients
We report the ladder rather than a single number. A rising citation count alongside a flat recommendation rate is a specific, diagnosable gap: your content is working and the web does not yet corroborate it. The fix is off-site, and reporting that hides the distinction hides the problem.
For requirements-heavy queries in your category we also check whether models recommend against you, and trace the sources they cite when they do. Nobody enjoys that slide. It is the one that changes what gets done next.
Frequently Asked Questions
No. Citation shapes how the model describes your category and its evaluation criteria, which is real positioning. It is just not the same as being on the shortlist, and it should not be reported as though it were.
Not reliably, and attempts to fake consensus through seeded reviews or bought forum posts are the fastest route to a poisoned entity. Recommendation is harder to game than a search ranking ever was, which is the good news for anyone genuinely good at what they do.
If you are the established leader, yes, and publish the definitive one. If you are emerging, publish genuinely useful guides but expect citation rather than recommendation, and put the marginal effort into reviews, communities and press instead.
Most likely you are climbing rungs 2 and 3 while rung 4 stays flat. Check the framing of your mentions rather than the count.