Start With the Decision the Audit Must Support
An AI visibility audit is a controlled assessment of whether answer systems can find your business, identify it correctly, associate it with the right services and locations, and use it in a relevant answer. It is not a single prompt sent to one chatbot, and it is not a vanity count of mentions. The useful output is a defensible set of decisions about data cleanup, content, partnerships, and measurement.
Write the decision statement first. A regional service company may need to decide whether to repair inconsistent location data before expanding service pages; a national brand may need to decide whether its product evidence is strong enough for comparison queries. Keep the audit scope narrow enough to act on, while covering the markets, services, and customer questions that matter commercially.
- Define the business entity, brands, locations, services, and competitors in scope.
- Name the audiences and buying stages the audit must represent.
- Agree in advance what counts as a material error and who owns its correction.
- 1
Business questions
- 2
Evidence collection
- 3
Findings
- 4
Prioritized fixes
- 5
Retest
Build a Query Set That Represents Real Demand
Build queries from customer language, sales-call notes, support tickets, Search Console themes, and internal site search—not from a list of keywords alone. Include branded prompts, category questions, problem-to-solution questions, comparisons, “best” and “near me” variants, and questions that test qualifications or exclusions. Phrase each query naturally and preserve the wording so later rounds are comparable.
Segment the set by intent and geography. A useful test matrix might include “who is [brand]?”, service-fit questions, provider-selection questions, local variations, and reputation or pricing questions. Do not manufacture demand or treat a model’s suggested prompt as market evidence. Label assumptions, and record which queries have a measurable business counterpart.
- Use at least one branded, non-branded, comparison, local, and risk-reduction query per priority service.
- Include alternate category terms only when customers actually use them.
- Freeze a baseline set; add exploratory prompts to a separate list.
Customer language
Intent groups
Priority services
Locations
Baseline prompts
Control the Test Environment
AI answers vary with model, retrieval mode, account context, location, date, and prompt wording. Document those variables instead of pretending the result is universal. Test the systems that customers can realistically use, and distinguish web-grounded answers from responses generated without current retrieval. If a platform provides citations, capture the cited URLs and the answer exactly as presented.
Run each baseline test consistently, then use a second pass to investigate anomalies. Avoid repeated prompting until the system produces the answer you want; that creates confirmation bias. Store screenshots or exports where permitted, the full prompt, timestamp, locale, logged-in state, and any model or search mode shown by the product. Sensitive customer information should never be included in prompts.
- Keep prompt wording, locale, device, and retrieval settings stable across comparisons.
- Record “not enough information” and incorrect answers rather than treating them as failures to reproduce.
- Use multiple independent observations before declaring a trend.
- 1
Prepare
- 2
Run baseline
- 3
Capture evidence
- 4
Classify result
- 5
Schedule retest
Score Visibility Without Hiding the Details
A scorecard makes a large audit usable, but a composite score should never replace the evidence log. Track at least four dimensions: presence (does the business appear when relevant?), accuracy (are facts correct?), prominence (is it among the useful options?), and citation quality (is there a traceable, authoritative source?). A fifth field for answer usefulness captures whether the response actually helps a buyer.
Use an explicit scale such as not observed, partial, and strong, with written criteria for each. Weight the dimensions according to the decision: a regulated business may weight factual accuracy and source quality above prominence, while a local provider may emphasize location and service-area correctness. Preserve raw observations so stakeholders can challenge a score without rerunning the entire audit.
- Log the exact claim, its status, source URL, and whether the source is first- or third-party.
- Mark hallucinated services, locations, credentials, pricing, and reviews as high-risk errors.
- Separate “not cited” from “not mentioned”; they require different remedies.
Presence
Accuracy
Prominence
Citation quality
Buyer usefulness
Diagnose Entity and Source Problems
Most audit findings become actionable when you trace a statement back to its evidence. Compare the answer’s business name, ownership, services, locations, hours, qualifications, and customer promises with the canonical information on the company site and trusted profiles. Then check whether important third-party references agree. Conflicting facts create ambiguity; missing facts create uncertainty; unsupported claims create risk.
Do not respond to every inaccurate answer by editing the website. First identify the source pattern. A wrong address repeated across directories calls for profile remediation; an unclear service definition calls for a better service page and structured data; an outdated review or article may require a direct correction request. Keep a finding’s severity tied to customer harm and business impact, not embarrassment alone.
- Create one canonical entity record with approved names, addresses, services, and proof links.
- Classify sources as first-party, authoritative third-party, user-generated, or unverified.
- Escalate safety, legal, clinical, financial, and credential errors for human review.
Turn Findings Into an Ordered Remediation Plan
Prioritize fixes by impact, confidence, effort, and dependency. Correcting a wrong phone number or service area is usually more valuable than publishing another broad article, because the error can distort every downstream answer. Establish canonical facts first, propagate them to important profiles, then improve the pages and evidence that explain why the business is a fit.
Assign an owner and acceptance test to each action. “Improve AI visibility” is not an acceptance test; “the service page states the service boundary, provider, location, evidence, and last-reviewed date, and the profile record matches it” is. Where a platform cannot be directly controlled, document the available correction path and set a monitoring date rather than promising a guaranteed change.
- P0: correct harmful or materially false claims and conflicting core identity data.
- P1: clarify priority services, locations, proof, authorship, and customer decision criteria.
- P2: expand coverage for emerging questions after foundational evidence is stable.
Critical factual error
Entity inconsistency
Missing evidence
Content opportunity
Experiment
Measure Change With a Baseline and a Control
Repeat the frozen query set after meaningful changes, using the same documented conditions. Track presence, factual accuracy, cited sources, answer framing, and referral behavior where analytics can identify it. AI referral data may be incomplete, so pair it with branded search trends, assisted conversions, direct inquiries, and qualitative sales feedback rather than claiming that every outcome came from an answer engine.
A comparison period or untreated query group can help separate an intervention from ordinary fluctuation. Record the date of content changes and profile corrections, and allow enough time for crawlers and indexes to update. Report uncertainty plainly: an observed citation is evidence of an appearance in one test context, not proof of universal ranking or causation.
- Report counts and examples together; percentages without the denominator are hard to interpret.
- Review high-risk claims monthly and the broader query set on a defined quarterly cadence.
- Retire prompts that no longer reflect customers, but preserve their history.
Make the Audit an Operating Practice
AI visibility changes as sources, interfaces, models, and business facts change. Establish a lightweight owner-led cadence: content or SEO maintains the query set, subject-matter experts validate claims, operations confirms locations and availability, and analytics reports outcomes. A shared evidence register prevents teams from publishing contradictory descriptions of the same offer.
The audit should improve the customer-facing information system, not encourage gaming. Never seed fake reviews, add unsupported schema, or write claims solely because a model repeated them. The durable strategy is accurate, accessible first-party information supported by credible external references and a disciplined process for finding and correcting errors.
- Define change triggers: rebrand, new location, service launch, leadership change, or regulatory update.
- Keep a changelog linking each correction to its evidence and owner.
- Include legal, privacy, accessibility, and brand review in the release workflow.
