Skip to content

AILast updated 8 min readBy The PixelCrayons team

AI Visibility Audit: Mentions, Citations
and Evidence

In one answer

AI Visibility Audit: Mentions, Citations and Evidence An AI visibility audit measures brand mentions, citations and recommendations across a defined set of questions, with a recorded method and evidence. It cannot reveal private model reasoning or guarantee future recommendations. Separate prompt groups and repeatable conditions help turn observations into a focused investigation.

An AI visibility audit measures brand mentions, citations and recommendations across a defined set of questions, with a recorded method and evidence. It cannot reveal private model reasoning or guarantee future recommendations. Separate prompt groups and repeatable conditions help turn observations into a focused investigation.

Editorial illustration connecting answer panels to their underlying sources.
Trace the visible answer to its evidence. A conceptual illustration, not a capture of an AI product.

Why might ChatGPT recommend a competitor?

A competitor’s appearance is an observation to investigate, not a diagnosis of your website. The question, conversation context, search mode, available sources and product behaviour can all affect the answer. A single missing mention cannot establish that your brand is unknown or that your SEO has failed.

Start by saving the exact response and question. Did it recommend the competitor, merely mention it, or quote a source that names it? Did it use web search? Did your prompt already name the competitor? Those distinctions change what the observation means.

The audit below is a proposed measurement workflow. It is designed to reveal where a deeper investigation is justified. It is not a universal ranking system or a claim that PixelCrayons has tested every platform.

For a broader diagnostic, start with the four layers of an SEO audit. If an AI answer proposes a tactic rather than names a provider, use our checks for an AI-generated marketing recommendation to assess whether that advice fits your business.

Which buyer questions should you test?

Build a prompt set from genuine buying tasks, using sales questions, site searches and support conversations where available. Group questions by the decision they support. Keep brand-free discovery questions separate from questions that explicitly name a company.

Prompt group Details
Unbranded discovery Example question: Which website agencies work with B2B SaaS marketing teams?
What it can tell you: Whether a brand appears without being supplied
Unbranded selection Example question: Which providers can improve a SaaS marketing website while keeping the existing product team?
What it can tell you: Visibility for a more specific buying constraint
Branded evaluation Example question: What should I verify before choosing the named company for a website project?
What it can tell you: How that named brand is described and supported
Branded comparison Example question: Compare Company A and Company B for the stated buying task
What it can tell you: Accuracy and framing when both names are supplied

Do not pool these groups into a single mention rate. A brand supplied in a question has a different opportunity to appear. For discovery comparisons, score all competitors within the same neutral answers. Swapping brand names creates a separate evaluation test, not a like-for-like discovery comparison.

Write down why each prompt belongs in the sample. Include realistic constraints such as market, business type and buying task. Avoid adding flattering wording for your brand or hostile wording for competitors. Freeze a core set for repeat comparisons and track exploratory questions separately.

How do you make the audit repeatable?

Record the conditions closely enough that another person can understand what you observed. Use a fresh conversation for each standalone question. If you are studying follow-up conversations, keep that as a separate protocol with the complete preceding context saved.

Record What to capture
Audit identity Run ID, date, time, timezone and prompt-set version
Product conditions Platform, displayed model or mode, account tier where relevant, search on/off/automatic
Context Language, locale, fresh or existing session, relevant personalisation settings
Input and output Exact prompt, complete answer and visible source links
Availability Successful answer, refusal, error or incomplete output
Scoring Brand, mention, recommendation, owned-domain citation, third-party citation and accuracy notes

If a model identifier or personalisation setting is unavailable, record “not exposed” rather than guessing. Repeated outputs can vary even when visible settings remain constant. Repeats describe that variation; they do not automatically produce independent observations or a representative estimate of all buyer conversations.

For a small pilot, you could run a fixed set of prompts three times under documented conditions. Label that as your pilot design, not an industry requirement. Report the counts, dates and limitations. Expand the sample according to the decisions you need to make, rather than chasing a large-looking score.

What should you count as a mention, citation or recommendation?

Use a written scoring rule before reading results. Count each brand once per answer for each metric, even if its name appears several times. Keep recommendations and citations distinct because either can occur without the other.

Metric Operational definition
Mention rate Valid answers naming the brand ÷ valid answers in the stated prompt group
Recommendation rate Valid answers explicitly presenting the brand as a suitable option for the task ÷ valid answers in that group
Owned-domain citation rate Valid answers with a visible citation to the brand’s own domain ÷ valid answers in that group
Third-party citation rate Valid answers citing an external page in support of a statement about the brand ÷ valid answers in that group

Define “valid” in advance. Keep refusals, tool errors and incomplete responses in the run log and report them separately. Do not remove an unfavourable answer because it lacks sources or brands. In a successfully returned answer, “no providers named” is a valid observation.

Mark ambiguous recommendations for review. For example, a warning about a provider is a mention, not a positive recommendation. A citation to a general industry page is not automatically evidence about your brand. Save the supporting sentence and source URL so another scorer can check the decision.

What does a completed audit example look like?

The following six rows are invented teaching data. They demonstrate the scoring rules and calculation only. They are not real ChatGPT outputs, competitive findings or a PixelCrayons visibility benchmark.

Valid answer Illustrative scoring
Discovery 01 Mention: yes. Recommended: yes. Own domain cited: no.
Discovery 02 Mention: no. Recommended: no. Own domain cited: no.
Discovery 03 Mention: yes. Recommended: no. Own domain cited: yes.
Discovery 04 Mention: no. Recommended: no. Own domain cited: no.
Discovery 05 Mention: yes. Recommended: yes. Own domain cited: yes.
Discovery 06 Mention: no. Recommended: no. Own domain cited: no.

In this illustration, the mention rate is 3/6, or 50%. Recommendation and owned-domain citation rates are each 2/6, or about 33%. These descriptive fractions say nothing about statistical significance or your share of the whole market. The full working log should also record third-party citations and factual accuracy.

Report results by prompt group and platform. Where a prompt has more repeats than another, disclose the weighting or use a consistently weighted prompt-level summary. Otherwise a heavily repeated question can dominate the headline number.

AI visibility audit workflow: separate prompt groups, hold visible conditions consistent, save complete responses, score defined metrics, investigate evidence and repeat.
A repeatable audit preserves the question, conditions and evidence behind every score.

How do you turn observations into useful work?

Choose an investigation that could confirm or reject a specific explanation. Missing citations are not proof that the model has no supporting knowledge. Recommendations do not reveal a hidden, stable league table.

Observation Details
Brand description is inaccurate Next check: Inspect your page and the visible cited sources
Possible action if confirmed: Correct owned information; request evidence-based corrections elsewhere
Competitors appear on a relevant discovery task Next check: Examine the task fit and cited evidence
Possible action if confirmed: Improve a genuinely missing buyer answer or proof asset
Your public page cannot be retrieved Next check: Check response status, robots rules and access controls
Possible action if confirmed: Repair unintended access restrictions
Brand is mentioned but not recommended Next check: Review the answer’s actual qualification language
Possible action if confirmed: Clarify fit and limitations where the evidence warrants it
Mentions increase, enquiries do not Next check: Inspect referrals, landing experience and qualified outcomes separately
Possible action if confirmed: Investigate acquisition and conversion before claiming commercial impact

Maintain a change log with the observation, supporting source, proposed action, owner and next measurement condition. Keep alternative explanations visible. Changes in platform behaviour or the source landscape can affect a later audit even if you changed nothing.

What do the platforms actually recommend?

Follow each platform’s published requirements. Do not turn advice for one search experience into a claim about every AI system.

Google’s current generative AI search guidance keeps SEO fundamentals central: useful original content, crawlable pages and a good reader experience. It does not require special AI schema or a particular writing format. Meeting requirements does not guarantee inclusion. Check the live guidance and your Search Console settings as features evolve.

OpenAI’s crawler documentation distinguishes OAI-SearchBot, used for search discovery, from GPTBot, used for potential training. Review the relevant crawler controls instead of treating every AI user agent as interchangeable. Allowing access is not a guarantee of citation or recommendation.

Fix inaccurate company information, publish evidence that helps buyers assess fit, and make important answers readable. Do not buy invented reviews, fabricate expert quotes or generate large numbers of near-identical pages. Those tactics do not create trustworthy evidence for readers.

When referrals reach your website but useful enquiries do not follow, start with the conversion measurement checks before interpreting the change as a visibility success or failure.

How should visibility connect to commercial measurement?

Track visibility, visits and business outcomes as separate stages. A citation can appear without a click. A click can arrive without a qualified enquiry. Some journeys will be unobservable in your analytics, so do not invent attribution to fill the gaps.

Start with the audit worksheet, then connect observed referrals to landing-page behaviour and CRM outcomes where your measurement permits. Our AI search optimisation, technical SEO and content marketing services address different parts of that work. Request a proposal with your buyer questions, target markets and current evidence.

Questions

Frequently
asked.

No. It records visible behaviour and helps prioritise investigations. It cannot expose the model’s private reasoning or isolate a single cause from a missing mention.

No. Supplying a brand changes the test. Report discovery, branded evaluation and branded comparison separately, using clear denominators.

No. Structured data should accurately describe visible content where appropriate. Google does not require special AI schema, and citation behaviour cannot be guaranteed by adding markup.

Choose a cadence that supports a real decision. Preserve the core prompt set and settings, record platform changes, and rerun after meaningful content or technical changes. A fluctuating daily score alone is not a useful strategy.

Have a process
to automate?

Tell us the job. A senior engineer scopes a pilot and prices it, with the limits stated up front.

Reply within four business hours (typical) · Free · No retainer required