Your buyers ask ChatGPT before they ask you. Most companies have never read the answer.
Not because they don't care. Because there's nothing to look up. Search gave you a rank — a number you could check, report, and argue about in a Monday meeting. AI engines give one answer, name a few vendors, and keep the list of losers to themselves. Nobody gets an alert when the answer stops including them.
So here is the method, in full, for finding out by hand. It's the same one I use in paid audits. There's no tool to buy and no trick in it — just discipline about how you ask and honesty about what came back.
The lookup illusion
First, the mistake almost everyone makes: they open ChatGPT, type their company name, read a flattering paragraph, and relax.
That test is worthless. Typing your name and getting a description proves someone who already knows you can look you up. It's a lookup, not a discovery — the buyer in that transaction had already found you.
The appearance that matters is the other kind: someone types the problem — "best vendors for X," "how do I fix Y" — with no brand name anywhere, and the engine decides who to name. That's where pipeline comes from. And it's precisely the appearance nobody checks, because checking it requires asking questions that don't contain your name.
Most companies that "show up in ChatGPT" are showing up when the prompt hands the engine their name. That's not visibility. That's an echo.
Build the prompt set before you open anything
Write the prompts first, away from the engines, so the results can't tempt you into softening the questions. You want 8–10 for a quick read, 30–40 for a real audit, across five kinds of intent:
- Category discovery — "best [category] vendors for [use case]." No brand names. These are the prompts that matter most.
- Brand-direct — "what does [company] do," "is [company] any good." The baseline.
- Comparison — "[you] vs [competitor]," "alternatives to [competitor]."
- Problem-first — the question a buyer asks before they know your category exists.
- Trust — "is [company] legitimate," "[category] pricing." The ones procurement actually types.
Pull the language from real buyers — sales call notes, support tickets, the phrasing in RFPs — not from your keyword tool. Engines answer questions the way people ask them, and your buyers don't talk in keywords.
Run it clean, or don't bother
This is where home-grown audits quietly go wrong. Engines personalize. If you run the prompts from your everyday account, the results carry your history, your location, and every previous conversation where you mentioned your own company. You will score artificially well, believe it, and fix nothing.
- Fresh session, every prompt. No conversation history, no memory, no signed-in personalization. ChatGPT's logged-out temporary chat is the cleanest instrument available.
- Three runs per prompt. Answers vary between runs of the identical question. One run is an anecdote. Three is a signal.
- Record everything verbatim, with the date. Screenshots included. Paraphrased findings drift toward what you hoped they'd say.
- Note whether browsing was on. An answer built from a live web read and an answer built from training data are different problems with different fixes.
- Watch your location. The engines geolocate the person running the test. I once watched an engine, asked about a business with no city in the prompt, confidently describe a different business with a nearly identical name two thousand miles away.
Score it like you mean it
Three numbers, none of them complicated. Appearance rate: answers naming you, divided by total answers. Share of answer: your mentions divided by all vendor mentions on the same prompts — because being named alongside four competitors is a different fact than being named alone. Citation share: of the answers that cite sources, how many cite something you own.
Then split every number by intent category. The split is the finding.
This month I ran this exact method, unsolicited, on a certified manufacturer with decades in its category — not a client, just a category worth measuring. Named in 16 of 30 answers, which reads like a win. But twelve of those sixteen came from prompts that already contained the company's name. Found on merit, with no brand name in the prompt: four times out of eighteen.
Same company, same day, two numbers. Only one of them grows the business.
What the findings usually say
Having run this across categories, the same four patterns keep coming back:
- The engines cite third parties, not you. Directories, review sites, industry press, Reddit. You can be the subject of the answer without being its source — and then the answer is only as accurate as whoever the engine trusted instead.
- Trust questions go generic. Ask "is this company legitimate" and engines often return supplier-caution boilerplate instead of the company's actual credentials — sometimes after reading the company's own site. Your own site is not corroboration.
- Nobody owns the comparison. Most categories have no credible "X vs Y" page, so the engines synthesize one. Whoever writes the honest version first tends to get cited for it.
- Conflicting facts make engines hedge. If your name, category, and claims differ across your site, your directories, and your LinkedIn, the engine splits the difference or leaves you out.
What you can't do about it
You cannot buy a placement, guarantee an appearance, or prompt-engineer your way into an answer. Nobody controls model outputs, and anyone selling a guaranteed ranking in ChatGPT is selling something they don't have.
What moves the probability: structure your pages so a machine can extract a direct answer, put real schema on your site, get present in the third-party sources the engines demonstrably cite, publish the comparison content they're currently forced to invent, and make your facts identical everywhere. Then re-run the same prompt set in 90 days — the change between two dated runs is the only number that means anything.
The engines are still deciding whose content they trust in most categories. In a lot of them, nobody has claimed the answer yet.
Run the prompts. Write down what came back. It's a Tuesday afternoon, and it replaces every opinion in the building with a number.
Questions people actually ask
How often should I re-run an AI visibility audit?
Quarterly, using the identical prompt set, and always with the run dates recorded. Models update and indexes refresh, so any single run is a snapshot. The number that actually means something is the change between two dated runs — appearance rate moving from one quarter to the next tells you whether the fixes worked.
Why do ChatGPT's answers change between runs of the same prompt?
Generation is probabilistic: the same question can produce different vendor lists, different framings, and different citations across runs, even minutes apart. That's why one run is an anecdote and three runs per prompt is the minimum for a signal. Score across all runs, not the run you liked.
Does this method work for Gemini and Claude too?
Yes — the method is engine-agnostic: clean session, fixed prompt set, verbatim recording, three runs. What changes is hygiene. Each engine personalizes differently, so check what account state you're carrying into the test, and note per engine whether web browsing was on, because a browsed answer and a training-data answer need different fixes.
What's a good appearance rate?
There's no universal benchmark, and anyone quoting one is guessing. The useful comparisons are internal: your category-discovery rate versus your brand-direct rate, and your share of answer versus the competitors named on the same prompts. A company named in 90% of brand-direct answers and 20% of discovery answers doesn't have a visibility problem — it has a discovery problem, which is more specific and more fixable.
Find out what the machines say about you. One page, forty-eight hours, the prompts your buyers actually ask — and exactly what came back, including who got named instead of you.
Get my free snapshot