How ChatGPT chooses which sources to cite (and how to become one)

ChatGPT with search on reads a handful of pages per question and links the ones it used. Here is what happens between your question and those links, which pages tend to make the cut, and the changes that get a site cited.

GuidesBy TangentFlow Team7 min read
Illustration of connected documents feeding into a chat window

When ChatGPT answers "best accountant for freelancers in Leeds" it does not consult a list of accountants. It runs a search, reads a few pages, writes an answer from them, and links the pages it leaned on. Those links are the citations. If your site is one of them you get named, described in your own words, and often clicked. If it is not, you are invisible for that question no matter how good you are.

This guide explains the mechanics in plain terms, then what to change.

Key takeaways

A citation is a page ChatGPT actually read while writing. It reads a small number of pages per question, chosen from a web search it runs first. Pages get picked when they match the question closely, load as plain text, state facts directly, and come from a source the model already treats as reliable for that topic. You can influence all four. You cannot buy your way in.

What a citation is

Turn web search on in ChatGPT and ask a question about a market. The answer arrives with small link markers in the text and a sources panel at the side. Each one is a page the model retrieved and used while composing that specific answer. OpenAI's own announcement of ChatGPT search describes it as answers "with links to relevant web sources".

Three things follow from that:

  • Citations are per answer. The same question asked tomorrow can cite different pages. There is no permanent index of "cited sites".
  • Citations are per question. Being cited for "invoice templates" says nothing about "best invoicing app".
  • A citation is not a recommendation. ChatGPT can cite your page while recommending someone else, because your page happened to list them.

Being recommended is the goal. Being cited is the mechanism, and it is the part you can engineer.

Nobody outside OpenAI knows the exact ranking, but the pipeline is visible from the outside and consistent with what OpenAI has published.

1. The question becomes searches. The model rewrites what you typed into one or more search queries. "Best accountant for freelancers in Leeds" might become "freelancer accountant Leeds", "accountants for self-employed Leeds reviews". Your page has to match at least one of those rewrites, which are usually plainer and more literal than the original.

2. A search engine returns candidates. OpenAI has said ChatGPT search uses third-party search providers alongside its own index, and its crawler documentation lists a dedicated bot, OAI-SearchBot, that builds that index. Which pages come back is close to ordinary search ranking: relevance, freshness, authority, and whether the crawler could read the page at all. A site blocking OAI-SearchBot in robots.txt does not get this far. Our AI bots checker shows what your robots.txt says to each bot.

3. A few pages are read. Not all of the results. A typical answer cites a handful of pages, usually well under ten. Those are the ones that were fetched and passed to the model as text. Pages that need JavaScript to show their content, or hide the answer behind tabs and accordions the fetcher does not open, arrive nearly empty.

4. The answer is written from those pages. The model summarises, compares and names businesses using what it just read, plus what it already knew from training. It then attaches links to the pages that supported each part.

Everything you can do falls into one of those four steps: be findable, be retrievable, be readable, be worth quoting.

Which pages tend to get picked

Across the markets we scan, the cited pages for "best X" and "X vs Y" questions cluster into a few types. The share varies by market, but the types are stable.

Type of page Example Why it gets cited
Roundups and listicles "10 best invoicing apps for freelancers" Directly answers the "best" question with names and reasons
Community threads Reddit, Hacker News, niche forums Plain-text opinions with specifics, heavily represented in ChatGPT citations in every public study we have seen
Reference pages Wikipedia, industry bodies Trusted for definitions and background
Review platforms G2, Capterra, Trustpilot, Google reviews via aggregators Ratings, counts and quotes the model can lean on
Vendor pages Pricing, comparison and docs pages on the product's own site Cited for facts about that vendor: price, features, limits
Directories Category directories with structured listings Cited when the market has no strong roundups

Notice what is missing: homepages. A homepage is rarely the best match for any single question, so it rarely gets read. The pages that get cited are the ones that answer one question completely.

The five properties of a citable page

Comparing cited pages with the ones that lost, the winners tend to share the same properties.

It matches the question literally. The question, or a plain rewrite of it, appears in the title and the first paragraph. "Accountants for freelancers in Leeds: fees, what they handle, how to choose" gets picked over "Welcome to Hartley & Co."

It answers in the first screen. The direct answer sits above the fold in prose or a short list. Long introductions push the answer past what the fetcher passes to the model.

It states facts as sentences. "Fixed fee from £45 a month, includes self-assessment" can be quoted. A pricing table built from icons, or a price only shown after a form, cannot. Our post on the facts AI can't find lists the facts assistants look for, page by page.

It loads as text. The page returns its content in the HTML, not after a script runs. The citability checker fetches your page the way a bot does and scores what came back.

It says who wrote it and when. A visible author, a date, and an organisation behind the page. Not because the model checks credentials, but because pages with those features are the ones search engines already rank, and step two is a search.

What does not work

A few things people try that do not move citations:

  • Adding "ChatGPT" or "AI" to page titles. The model is searching for the customer's question, not for pages about itself.
  • Stuffing the page with the question in twenty variations. Search engines demote it, so it never reaches step three.
  • Blocking competitors' names. If a roundup on your site lists rivals honestly, it is more likely to be cited, and you are named alongside them with your reasons intact.
  • Paying for placement. There is no advertising in ChatGPT's citations. Sponsored slots in directories are marked and, on any directory worth being in, do not change the rankings.

How to become a source, in order

  1. Check you can be read. Robots.txt, JavaScript-only content, and login walls each remove you from step two or three. Fix these first; nothing else matters until they are done.
  2. Find the questions. Write the ten questions your customers ask before they know your name. Ask them in ChatGPT with search on and note which pages are cited. Our guide to checking whether ChatGPT recommends you walks through it.
  3. Build one page per question. Not a blog post that mentions it. A page whose whole job is that question, with the answer in the first screen and the facts as sentences.
  4. Get onto the pages already cited. For each question, two or three third-party pages are cited again and again. A review platform, a roundup, a community thread. Get listed, answer the thread, or ask the author for an update. The sources that decide who gets recommended explains how to find yours.
  5. Re-check weekly. Citations move. A page that was cited in March can be replaced in April by a fresher one. Checking once tells you where you stood; checking weekly tells you whether the work is landing.

What this looks like when it works

Take a made-up but typical case. A photography studio is not named for "wedding photographer in Bristol under £2,000". The cited pages were two directories, one wedding blog roundup, and a Reddit thread. The studio's own site has prices only in a downloadable PDF. Three changes: a pricing page with the packages written as sentences, a listing in both directories, and a reply on the thread from the owner with the same facts. A few weeks later the studio is named in the answer, with the pricing page as a citation and the reason "packages from £1,450 with a second shooter".

That is the whole pattern. Be readable, answer the actual question, be present where the model already looks, and check again.

If you want the list of questions and cited pages for your own market, a free scan asks ChatGPT twelve of them and shows every source it used. Or read the sample report first.

Find out what AI says about you before your customers do.

One URL. About a minute. No account.

Keep reading