Does the AI recommend our client? An AI-visibility audit for a real local business
The situation
A local business owner came to my class and said her phone had stopped ringing even though her search rankings had not moved — people were asking chatbots for recommendations and she was not in the answer. That is now a job agencies bill for, and it is good for students because it is measurable, local, and cannot be faked. What I assess is methodological discipline under noisy conditions.
Steps
-
Build a prompt set from how customers actually ask
Twenty to thirty natural-language prompts tiered by funnel stage: discovery, evaluation, decision. These come from asking the client what customers actually say on the phone, not from a keyword tool. The prompt set is the instrument.
What you only learn by doing it: Force at least a third of the prompts to omit the client's brand name entirely. Students instinctively write prompts containing the name, which guarantees a mention and produces a meaningless 100% visibility rate. Unbranded prompts are the only ones that measure discovery.
-
Baseline by hand, three runs per prompt per platform
Every prompt runs at least three times on each platform in a fresh session, logging whether the client was mentioned, at what position, and what sources were cited. Running once tells you almost nothing — these systems are stochastic.
What you only learn by doing it: The three-run rule teaches variance in a way no statistics lecture achieves. They will get a mention on run 1 and nothing on runs 2 and 3, and have to decide what to report. Make them log all three and report the fraction — and use logged-out sessions, because their own chat history silently inflates everything.
-
Compute the metrics and read the citations, not just the mentions
A spreadsheet
Visibility rate, share of voice against named competitors, average position, and accuracy of what was said. Then catalogue which domains the assistants cited instead. That citation list is the actual diagnosis.
What you only learn by doing it: The citation audit is where the insight lives. Over and over the answer is not “our client needs more content” — it is that assistants are citing a directory or a review site where the client has a stale listing. That is a different and much cheaper fix, and students only find it by reading sources rather than counting mentions.
-
Check whether what the assistants say is even true
The client, the client's site, the outputs
Where the client is mentioned, verify the substance against the client's own facts — hours, service area, pricing, offerings. Errors get logged with the exact prompt and platform.
What you only learn by doing it: This is often the single most valuable deliverable to the client and students do not expect it. Bring receipts — a screenshot with the prompt visible — because the client will not believe it otherwise and the output will not reproduce on demand.
-
Ship one change, then re-measure identically
The client's CMS or listing profiles; the same protocol
One scoped intervention — a comparison page, a corrected listing, an FAQ rewrite, schema markup — then wait several weeks and re-run the identical protocol.
What you only learn by doing it: Insist on one change. Students want to ship five and then cannot attribute anything. Also set expectations on timing: practitioner accounts talk about roughly 60 days before changes surface, so “no detectable change yet” is a legitimate finding that should earn full marks if the method was sound.
-
Report with the uncertainty attached
Slides or a written brief
Before and after visibility rates with run counts, what changed and what did not, and observed correlation separated from claimed causation.
What you only learn by doing it: Grade the honesty of the causal language above the size of the result. The agency reports students imitate routinely claim revenue attribution from a visibility bump, and that chain has several unmeasured links. A student who writes “we cannot separate this from seasonality” has understood the assignment.
Where this breaks down
This is the least pedagogically documented workflow on the site. The practitioner process comes from agency write-ups and vendor material with a commercial interest in the answer, and we found no peer-reviewed or association-published faculty case study. Treat the metric definitions as a working convention, not a standard.
AI assistants fabricate business details freely, so every factual claim about the client must be verified with the client rather than assumed.
Results are genuinely noisy. Personalisation, session history, geography and model updates all move the numbers, so any single-run finding is close to meaningless and any before/after comparison run on different account states is invalid.
Working with a real client raises real obligations: get written agreement about what students may access and publish, do not let students push changes to a live site without owner approval, and be clear the class cannot guarantee outcomes.
Provenance: practitioner workflow documented, teaching application extrapolated. The eight-step agency process — tiered prompt library, baseline, multiple runs per prompt per platform, visibility and share-of-voice metrics, citation-source analysis, scoped intervention, re-measurement — is described in published agency accounts.