Skip to content

Averages & Denominators

Every card on the panel is an average, and no two of them average the same set of probes. The differences aren’t incidental — each exclusion prevents a specific way the number could lie.

The denominators

MetricNumeratorDenominator
AI ScoreSum of AI ScoresAll scorable probes in scope
SentimentSum of sentiment scoresScorable probes that mention the brand
Mention RateDiscovery probes that mention the brandScorable discovery probes
PositionSum of position scoresAll scorable probes in scope
Final ScoreDerived from the AI Score average, see Final Score

“In scope” means whatever you’re looking at: the whole run by default, or a single region, a single engine, or the global-engines group when filtered.

Scorable: the one exclusion that applies everywhere

A probe is scorable when the analysis model returned a usable judgement. When analysis fails after retries, the probe is marked not analyzed and removed from every metric — score averages and mention counts, numerator and denominator alike.

denominator = total probes − probes whose analysis failed The panel reports the excluded count, so the gap between "probes executed" and the averaged count is always visible.

The reasoning is worth being explicit about:

  • A failed analysis means not measured, not “performed badly”. Scoring it as 0 would drag the average down in proportion to a third party’s error rate.
  • Excluding it from score averages but keeping it in the mention-rate denominator would be equivalent to recording it as “not mentioned” — the same error, wearing a disguise.
  • There is no keyword fallback that guesses mention status from the answer text. String matching can’t separate your brand from a same-named company, from the same word used ordinarily, or from a hit inside an echoed search URL, and all three failures overstate visibility.

The probe’s answer text and citations are still stored. The probe detail page shows the full answer, in place of the scores, and a note explaining that it is excluded.

Sentiment: mentioned probes only

Sentiment is a property of how you were talked about. An answer that never brings your brand up has no tone toward it. Counting those as 0 would mean a brand nobody mentions reads as a brand everybody dislikes, and the fix for one problem would look like the other.

So the sentiment average is taken over answers that mention you — in both buckets, since branded prompts are graded for tone too.

The panel shows sentiment rescaled to 0–100 (raw ÷ 15 × 100).

Mention Rate: discovery probes only

Branded prompts name your brand in the question, so they guarantee a mention. Including them would push Mention Rate toward 100% as you add comparison prompts — measuring your prompt library rather than the market.

The panel therefore reports:

  • Mention Rate — discovery probes only, as a percentage.
  • A second line, + x/y in brand-named prompts, so branded coverage is still visible without contaminating the headline.

Per-engine and per-region breakdowns

Both use the same formulas over a filtered subset:

BreakdownSubset
By engineThat engine’s probes, all regions
By regionProbes from geo-adjustable engines only, in that market
Global enginesProbes from engines with no location signal

Cross-region comparison is restricted to geo-adjustable engines because only those run the same prompts through the same engines in every market. Mixing in an engine that runs once in the primary market would give each region a different denominator while the table implied they were comparable.

Region averaging is region-weighted, not probe-weighted

The cross-region average — the baseline each region’s vs Avg delta is measured against — is the arithmetic mean of the regions’ scores, not of all probes pooled:

cross-region average = mean(region₁ score, region₂ score, …) Regions with no probes at all are left out entirely rather than contributing a zero.

With equal probe counts the two approaches agree. They diverge when a region loses a probe to an engine timeout: probe-weighting would quietly shrink that region’s influence on the baseline, whereas “cross-region performance” is a question about regions, not probes.

Rounding

All displayed averages are rounded to one decimal place at the end of the calculation. Intermediate values are kept at full precision, so a number never drifts by accumulating rounding steps.

Reading the gap between counts

Three counts appear on the panel and they are not the same number:

CountMeans
Probes executedHow many prompt × engine × region calls the run made
Probes averagedExecuted minus failed analyses
MentionedScorable probes where the model found your brand

If executed and averaged diverge noticeably, a provider was having a bad day — the run is still valid, just measured on fewer samples. Re-running the same selection costs allowance, so it is usually only worth it when the gap is large.

See these numbers for your own brand The free plan probes every engine we support — no card required.

Start free