How Perplexity, ChatGPT And Gemini Pick Their Sources

From BloomWiki
Revision as of 15:32, 16 August 2026 by 104.23.172.97 (talk)
Jump to navigation Jump to search

The honest framing first: nobody outside these organisations knows the selection logic, and the systems change without announcement. What follows is drawn from observable behaviour, visible citations and published research, which supports useful generalisations and does not support precision.

There is a specific moment worth picturing. Somebody types a question into an assistant asking who they should use for the thing you sell. A short list comes back. If your name is not on it, you were never in the running, and unlike a search results page there is no second page for them to try.

Keep the raw text of every answer, not just a tally. Six months in, the archive is the most useful thing you own, because it lets you see exactly when a competitor entered the shortlist, which source appeared alongside them, and whether your own description shifted from something a marketer wrote to something a customer would recognise. A score with no working behind it cannot tell you any of that. ai search optimization

No, though the foundations overlap. The measurement, the target surfaces and the emphasis on third party sources are genuinely different, and the Ahrefs overlap data shows the two channels draw from largely separate pools of pages.

And do not let anyone rewrite your entire site in the flat, listicle heavy register that is currently fashionable in this discipline. It reads as machine assembled to human beings, and content that reads that way tends to be treated as low quality by both audiences.

The second divergence is that third party sources carry unusual weight. Review sites, directories, forum threads, comparison articles and press coverage are frequently what an assistant quotes when asked about a category. Your own site is one voice among many, and often not the loudest.

Statistical Caution This field circulates numbers faster than it checks them. A widely repeated referral growth statistic rested on nineteen analytics properties. A frequently quoted conversion comparison came from a company selling the service it flattered.

Where a Real Tension Exists Two places, and they are worth naming honestly rather than pretending everything aligns. The first is the hero section. A large image with six words over it is a legitimate design choice and it gives a machine nothing to work with.

Writing Prompts That Sound Like Customers The foundational skill is deceptively mundane. Somebody has to write the questions your buyers actually ask, in their words, without the category vocabulary your team uses internally.

What Not to Do in the Name of Legibility Hidden text intended only for machines fails on every axis. It is detectable, it violates most guidelines, and it produces exactly the uniform low quality signal you were trying to avoid.

Buy the technical audit if nobody on the team reads server logs, and buy the third party source work unless you already have a functioning public relations capability. Those are the two areas where the learning curve is steep and the cost of getting it wrong is highest.

A false trade off gets invented early in most of these projects. Somebody proposes stripping the design, flattening the copy and restructuring everything around what a crawler finds convenient, and somebody else correctly points out that this would make the site worse for customers.

Keep a dated note of what you observed each quarter, including behaviour that later turned out to be temporary. The value is not in the individual observations, most of which expire, but in noticing how fast they expire. A team that has watched three of its confident conclusions become wrong within a year develops the right amount of scepticism about the fourth.

This variability is the main practical trap. Testing without web access and concluding you are invisible measures the training corpus rather than current retrieval, and the two can disagree sharply. Record which mode you used with every run.

You cannot control those pages, but you can influence them. Claim and complete your listings. Correct factual errors where the platform allows it. Respond to reviews. Give journalists and analysts accurate material to work from. Where a comparison article about your category exists and gets your details wrong, a polite correction is often accepted.

Two implications follow regardless of which system you are studying. Being findable by the underlying search step is necessary, and being worth quoting once fetched is what decides whether you are used. Almost everything actionable sits in those two requirements.

Where It Diverges Sharply Traditional SEO optimises for a ranked list. Generative systems optimise for a synthesised answer, and the sources they pull from are not the same set. Ahrefs studied 15,000 long-tail prompts across four assistants in July 2025 and found that around 80 percent of the pages cited did not rank anywhere for the original query, with only about 12 percent appearing in the top ten. Ranking first does not reserve you a seat.