# AI Visibility Score: A Manual Method
Supplier discovery in large language models depends on entity consistency across third-party sources, so AI visibility is measured by asking the questions buyers ask and recording which suppliers get named. A manufacturer can run this by hand with a fixed prompt set, five engines, and a monthly cadence.

# AI Visibility Score: A Manual Method

<div data-answer>

Supplier discovery in large language models depends on entity consistency across third-party sources, so AI visibility is measured by asking the questions buyers ask and recording which suppliers get named. A manufacturer can run this by hand with a fixed prompt set, five engines, and a monthly cadence.

</div>

**The interactive scorer is not built.** This page is the method it would automate, written so a manufacturer can run it today with a spreadsheet and about an hour a month. Everything below is executable by hand.

## What is being measured

Not rankings. The question is whether an assistant names your company when a buyer describes a need in natural language, and what it says about you when it does.

Generative engines cite sources that carry explicit, extractable, attributed statements, and supplier discovery in large language models depends on entity consistency across third-party sources rather than on-site optimisation alone. Those two facts set what this measurement can detect: it reads the output of an entity graph the manufacturer only partly controls. AI Overview presence also reduces click-through rate on informational industrial queries, so a supplier can lose traffic to an engine that never names it, which is a separate failure this method does not measure.

That produces three numbers per engine: whether you appeared, in what position among named suppliers, and whether the description was accurate. The third matters as much as the first, because an assistant that names you and misstates your capability has produced a disqualification rather than a lead.

## Step 1: fix the prompt set

Write between ten and twenty prompts phrased the way a buyer would ask, not the way a marketer would. Each should carry a constraint, because unconstrained prompts return generic answers that measure nothing.

Poor: "best contract manufacturers". Useful: "Find me a CNC machining supplier that can hold plus or minus 0.0005 inch tolerance in Inconel." Useful: "I need an ISO 13485 injection molder for a Class II device. Who should I contact?"

Draw the constraints from your own quote requests. The language buyers actually used is the language to test, and it is the same source that should be driving [keyword research](/manufacturing-seo).

**Freeze the set.** A prompt panel that changes between runs measures nothing over time. Add prompts to a second list if needed, and keep the original panel stable so the trend is real.

## Step 2: fix the engines

Run every prompt against the same engines every time. ChatGPT, Claude, Perplexity, Gemini, and Google AI Overviews is a reasonable panel, and the specific five matter less than using the same five each run.

Use a signed-out or fresh session where possible. A logged-in session carries personalisation and prior context, which contaminates the result with your own history.

## Step 3: score each response

Four fields per prompt per engine.

**Named.** Was your company named at all. Yes or no.

**Position.** If named, where among the suppliers listed. First, second, third, or later.

**Accuracy.** Was what it said about you correct. Score correct, partially correct, or wrong, and record the specific error where there is one.

**Citation.** Did it link a source, and was that source your site, a directory, or a third party.

The citation field is the most useful and the most skipped. It tells you which sources the engine treats as authoritative for your entity, and that is what you can act on.

## Step 4: compute the score

Three figures, all simple.

**Coverage** is the share of prompt-engine pairs where you were named. Ten prompts across five engines is fifty pairs; named in twelve is 24 percent coverage.

**Accuracy rate** is the share of times you were named where the description was correct.

**Own-source rate** is the share of citations pointing at your own domain rather than a directory or third party.

Do not average these into a single number. They move independently and for different reasons, and collapsing them hides the diagnosis.

## Worked example

A hypothetical shop runs 10 prompts across 5 engines, giving 50 pairs.

Named in 12 pairs, so coverage is 24 percent. Of those 12, the description was correct in 7, partially correct in 4, and wrong in 1, giving an accuracy rate of 58 percent. Of 9 citations across those appearances, 2 pointed at the company site and 7 at directories, giving an own-source rate of 22 percent.

The reading: visibility is low but the larger problem is that engines are describing the company from third-party listings rather than from its own pages. The action is not more content, it is entity consistency across the sources engines are actually reading, which is what [AI search visibility](/manufacturing-seo/ai-search-visibility) covers.

## Step 5: set a cadence and hold it

Monthly is enough. Model versions change, indexes update, and a single run is a snapshot rather than a measurement. Manufacturing SEO commercial queries carry AI Overviews in approximately 80 percent of cases, so the surface being measured is present on most of the demand rather than on a fringe of it.

Record the date, the engine versions where visible, and the raw responses. The raw text matters: six months later the useful question is usually what changed in the wording, and a score alone cannot answer it.

## What this measures badly

Stated so the numbers are not over-read.

Assistant responses vary between runs for the same prompt, so a single run carries real noise and small changes between months are not signal. Personalisation and geography affect results even in fresh sessions. And coverage across a ten-prompt panel says nothing about prompts you did not test, which is why the panel should come from real quote requests rather than from imagination.

## Common questions

**Why is the interactive scorer not built?**

Because a form that promises a report and does not deliver one is worse than no tool, which this site learned by shipping exactly that. The method is published now because it works by hand; the automation ships when it works.

**How many prompts is enough?**

Ten to twenty. Fewer produces too much noise per observation, more becomes a job nobody repeats monthly. Consistency across runs matters more than panel size.

**Should competitors be tracked in the same run?**

Yes, and it costs nothing extra. Record which suppliers were named alongside or instead of you. That converts coverage into a share-of-voice figure and identifies which competitors the engines treat as canonical.

**What if the assistant describes the company incorrectly?**

Treat it as the highest-priority finding. An engine repeating a wrong certification, a wrong process, or a wrong location is actively disqualifying the company, and the fix is correcting the third-party sources it is reading rather than editing your own site alone.

**Does this replace rank tracking?**

No. It measures a different surface. Buyers use both, and the two answer different questions about how a supplier is found, as set out in [how industrial buyers search and decide](/buyers).
---
Source: https://manufacturingseo.ai/ai-search/measuring-ai-visibility/
Last reviewed: 2026-08-07
Author: ManufacturingSEO.ai
