Checking ChatGPT once tells you where you stood on one afternoon. It does not tell you whether the pricing page you rewrote did anything, whether the review platform you joined is now being read, or whether a rival just got added to the roundup that decides your market. For that you need the same questions, asked the same way, on a schedule, with the results written down.
This is the method we used by hand before we built the tool, and it still works in a spreadsheet.
Key takeaways
Fix a list of questions and never change it. Ask each in a fresh chat with search on. Record four things per question: named or not, position, the reason, and the cited domains. Compare week to week, not day to day. Automate once the list is stable and you are making changes you want to measure.
Step 1: Fix the question list
Write ten to twenty questions your customers ask before they know your name. Mix the types:
- Finding: "best invoicing app for freelancers", "accountant for contractors in Manchester"
- Comparing: "FreshBooks vs Ledgerly", "Ledgerly alternatives"
- Deciding: "is Ledgerly worth it", "Ledgerly pricing"
- Specific need: "invoicing app with Stripe sync", "dentist open Saturdays in Leeds"
Then freeze the list. The whole value of tracking is asking the same thing repeatedly. Add questions if you must, but never edit an existing one, and keep the wording plain. Our guide to checking whether ChatGPT recommends you has more on choosing them.
Step 2: Ask the same way every time
Three rules that keep the results comparable:
- Fresh chat every question. Earlier messages shape later answers. One question per chat, then close it.
- Web search on. Citations only exist when the model searches. Without search you are measuring training data, which changes on a different clock.
- Same account, no memory. Turn off chat memory or use an account that has never discussed your business. Otherwise the model knows you and the test is meaningless.
Ask each question on the same day each week. The answers vary a little between runs even on the same day; weekly spacing makes real change visible over that noise.
Step 3: Record four numbers per question
Keep a sheet with one row per question per week. Four columns matter:
| Column | What to write | Why |
|---|---|---|
| Named | Yes, caveat, or no | The headline. "Caveat" means named with a "but" or a wrong fact. |
| Position | 1, 2, 3… or blank | First-named brands get most of the clicks and most of the trust. |
| Reason | The phrase attached to your name, or to the winner | Tells you which fact the model is using, and which fact you need to supply. |
| Cited domains | Every source domain, comma separated | This is the column that explains everything else. |
Two optional columns pay off later: the rival names in order, and a wrong fact note if the answer got something wrong about you.
Twenty questions takes about forty minutes a week by hand. That is the honest cost of doing this manually.
Step 4: Turn the rows into three trends
After four weeks you can read trends. Three are worth a chart.
Share of answers. The percentage of questions where you were named. Track yourself and the two or three rivals that appear most. This is the number to report upward; it moves slowly and it is hard to argue with. We publish the same figure for whole markets in the monthly benchmarks.
Cited-domain frequency. Count how often each domain appears across all questions. The top five domains are your market's sources. If a review platform is cited in twelve of twenty answers and you are not on it, that is the week's task. If a domain you just got listed on starts appearing, the listing is being read.
Reason drift. Are the reasons attached to the winner stable ("cheapest") or changing? A stable reason you cannot match is a positioning problem. A changing reason means the sources are in flux and there is room to move.
Step 5: Tie changes to dates
Add a column for what you changed and when. "Rewrote pricing page", "listed on Capterra", "answered r/freelance thread". Then when a question flips from no to yes, look back three to six weeks. That is usually where the cause is.
Without this column, tracking tells you what happened but not why, and you will not know which work to repeat.
Reading the results honestly
A few things that trip people up:
- A single week means nothing. Answers wobble. Two consecutive weeks in the same direction is a signal.
- Position 1 for a branded question is not a win. Of course it names you when you ask about you. Weight the finding and comparing questions.
- Being cited is not being named. Your page can be a source for an answer that recommends someone else. Record both.
- Different assistants, different sources. Perplexity and Gemini cite different pages from ChatGPT. If you have the time, track at least one other assistant; if not, ChatGPT is the right one to start with.
When to stop doing it by hand
Manual tracking is right while you are still choosing questions and deciding whether the work is worth it. It stops being right when any of these is true:
- You are making changes you want to measure and cannot afford to miss a week.
- You want more than one assistant.
- You have more than one business, or clients.
- You want the reasons and citations parsed rather than typed.
At that point a tool pays for itself in the first hour. TangentFlow's free scan runs twelve questions on ChatGPT and records every column above; paid plans re-run the same questions weekly on up to four assistants, chart share of answers against your rivals, and email you the moment a question flips. The sample report shows what the output looks like, including the cited-domain list.
Whichever way you do it, the method is the same: fixed questions, fresh chats, four numbers, weekly, with dates for what you changed. The tools only remove the forty minutes.



