AI Visibility

AI Brand Visibility Score

One number for how visible your brand is across AI assistants, and a full account of how it is calculated.

Get My Score

A visibility score compresses hundreds of model responses into one number so you can see movement. That compression is only useful if you know what was thrown away. This page gives the actual factors and weights behind ours, and is honest about what the number cannot tell you.

Last updated

The short version

  • Six factors: mention count 25%, position 20%, sentiment 20%, source quality 15%, context match 10%, competitor gap 10%.
  • Position and sentiment together outweigh raw mention count, deliberately.
  • Per-platform results are multiplied before combining, so a mention on a high-reach assistant counts for more than one on a small one.
  • Scores are only comparable within one version of the scoring engine. When we change the engine, we stop reporting the delta rather than pretending it is your trend.

The Six Factors, and Why Those Weights

Every factor answers a question that raw mention counting cannot. The weights are a judgement about what actually changes a buying decision, and they are worth arguing with, which is why we publish them.

  • Mention count (25%): across the prompt set, how many answers named you at all. The floor of visibility.
  • Position (20%): where you appeared. The first name in an answer absorbs most of the attention; the sixth is nearly invisible.
  • Sentiment (20%): recommended, listed flatly, or named as the option with a caveat. A negative mention can cost more than an absence.
  • Source quality (15%): whether the model leaned on authoritative pages when naming you, which predicts whether the mention will persist.
  • Context match (10%): whether you were named for what you actually sell. Being recommended for the wrong use case produces traffic that does not convert.
  • Competitor gap (10%): your visibility relative to the brands sharing those answers, because visibility is positional, not absolute.

Why Platforms Are Weighted Before Combining

A flat average across assistants would let a strong result on a small platform mask a weak one on a large platform. Since the whole point is estimating commercial exposure, each platform is multiplied by a factor reflecting reach before the scores are combined, with ChatGPT highest at 1.2 and smaller assistants below 1.0. This has a consequence people sometimes dislike: doing well on a niche platform moves the number less than they expect. That is the honest outcome, because it moves your revenue less too.

What the Score Cannot Tell You

Being straight about the limits is what makes the number usable. It is an estimate from a sample of prompts, not a census of everything a model might say. Models are non-deterministic, so there is genuine variance run to run. A score cannot tell you revenue impact, because it does not know your conversion rate or margin. And it is not comparable across companies in different categories, because a crowded category with fifty credible players produces lower scores than a category with four, for everyone in it. The comparison that means something is you against yourself over time, and you against the specific competitors sharing your answers.

When the Ruler Moves

This is the failure mode most scoring products have and do not disclose. If we improve the engine, scores move for reasons that have nothing to do with your brand, and reporting that as your trend would be actively misleading. So we stamp an engine version onto every stored result and withhold the comparison across a version boundary, telling you plainly that this score is not comparable to the last one and the next will be. A visibility trend is only a measurement of you if the ruler stayed the same.

How to Actually Use It

Treat the number as a dashboard light rather than a target. The useful work always lives in the breakdown: which factor is dragging, and which prompts you lost. A score of 40 built from many neutral, late-position mentions needs completely different work from a score of 40 built from a handful of first-position mentions in a narrow category. Optimising the aggregate directly is how you end up gaming your own metric.

Frequently Asked Questions

What is a good AI visibility score?

It depends heavily on category size, so absolute thresholds mislead. In practice, under 20 means you are effectively absent, 20 to 50 means you appear inconsistently, and above 50 means you are a default answer in at least some question types. Your position relative to the competitors in your own answers matters more than the number.

Why did my score change when I did nothing?

Model non-determinism, retrieval results shifting, and competitors publishing all move it. That is why single-point comparisons are weak and a trend across several checks is what you should read.

Can I compare my score to a company in another industry?

Not usefully. A category with fifty credible vendors distributes mentions across all of them and depresses everyone. Compare against your own history and against the brands appearing alongside you.

Do you weight all platforms equally?

No. Each platform is multiplied by a reach factor before the scores are combined, with ChatGPT the highest. Equal weighting would make a strong showing on a small assistant look like commercial exposure it is not.

What happens to my score when you change the engine?

The new score is stored with an engine version, and we suppress the comparison to any score computed under an older version rather than reporting our own change as your trend. Your next check restores the trend line.