Methodology
How we measure whether AI engines see your brand
Most AI-visibility scores are opaque. Ours is not: every number in a Neoval report is computed from findings you can read in the same report, with the rule stated here. If we could not observe something, the report says so and the score excludes it. This page is the contract behind every audit and every monitoring run.
1. Buyer prompts
Questions in the buyer's words
We measure what a buyer would ask, not what a marketer would type.
- Prompts are phrased the way buyers phrase them: "best [category] for [situation]", "alternatives to [leader]", "[category] that [differentiator]", "how do I [job the product does]".
- They are written in the language and market of the audited site. A French site is asked in French; asking in English would test a different set of candidates.
- The brand is never named in the prompt. A prompt containing the brand measures recognition, not visibility.
- A free audit uses five prompts; the deep audit twenty; Starter tracks ten and Growth and Expert thirty, fixed across runs so the trend is comparable.
2. Engines and sessions
Fresh sessions, dated, one engine at a time
An answer is only evidence if it was produced the way a stranger would get it.
- Each prompt runs in its own new session with no memory, no previous conversation and no shared context with the audit itself.
- Engines covered depend on what was bought: ChatGPT on the free audit and on Starter; ChatGPT, Claude, Gemini, Perplexity and Google AI Overview on the deep audit, Growth and Expert. Google search presence is part of every audit. An engine a tier does not cover is left out of the tables and named as not included, never shown as unreachable.
- A deep audit repeats every prompt on three separate days, so a citation that holds can be told from a single lucky sample.
- Every observation records the query, the engine, the date, whether the brand was cited (and at which position), which brands were cited instead, and the sources the engine relied on.
- An engine that could not be queried in a run is recorded as not observed for every prompt. It is never omitted and never estimated.
3. What counts as a citation
Named as an option, for that question
The bar is deliberately strict, so that a change in the number means something.
- Cited: the engine names the brand or links its domain as one of the options in its answer to that prompt.
- Not cited: the engine answered and the brand was absent, including cases where the brand appears only as a negative example.
- Not observed: the engine could not be queried, or refused to answer. Excluded from every rate.
- A mention in a source page the engine cites, without the engine naming the brand in its answer, does not count. It is noted as a source signal, separately.
4. Scores
The rubric, in full
Both scores are deterministic functions of the report's own findings, on a 0 to 100 scale.
| Component | Points | Rule |
|---|---|---|
| Technical controls | 40 | 40 × (ok + 0.5 × warning) / number of controls checked |
| On-page problems | 30 | 30 minus 6 per critical, 3 per high, 1 per medium, 0.5 per low; floor 0 |
| Indexation and coverage | 15 | 5 each: sitemap consistent with the crawl; canonicals on inspected pages; indexed count within 25% of the sitemap |
| Search presence | 15 | 15 × share of audited queries where the site ranks in the top 10; only observed positions count |
| Component | Points | Rule |
|---|---|---|
| AI crawler access | 25 | 25 × share of GPTBot, ClaudeBot, PerplexityBot and Google-Extended allowed by robots.txt (unspecified = allowed) |
| Citation rate | 45 | 45 × (prompts where the brand is cited / prompts actually observed); rows marked not observed are excluded, never counted as zero |
| Citability signals | 20 | 5 each: Organization structured data; a one-sentence positioning statement on the home page; at least one comparison or reference page; visible dates and authorship on content pages |
| Competitor share | 10 | 10 × (1 minus the share of observed prompts where a competitor is cited and the brand is not) |
Exclusion, not zero.A component with no evidence is excluded and the remaining weights are renormalised to 100. The same rule applies within a component: an unverifiable sub-check shrinks the component's weight rather than failing it. Anything without a machine-readable status counts as unknown, never as favourable.
Reportability. Renormalisation is only valid while measured components keep at least half of a score's weight. For GEO that means the citation rate must have been measured on at least one engine; otherwise the GEO score is shown as n/a, never as a perfect score built from the easy components.
Rounding. Half up, to an integer. The verdict paragraph of every report lists which components were scored.
5. What we refuse to guess
Known limits
A method is defined as much by what it leaves out.
- Answers are personalised and non-deterministic. One reading is a sample; a fixed prompt set observed on a cadence is the measurement. Read trends over weeks, not days.
- Keyword volumes and difficulty are estimates from observed result pages unless a search-data source is connected; the report says which.
- We do not infer why an engine chose a source. We record what it cited and let the findings speak.
- The audit reads public pages only, treats website text as untrusted content, and stores nothing from the audited site beyond the report itself.
The rubric is versioned (report-html-v1). Changing a weight means a new version, so runs of the same domain stay comparable. Questions about a specific result: seo-geo@neoval.net. Try it on your own domain from the home page.