Vector LabsAI Search Visibility Report Northwind Analytics · northwind.io · 2026-09-11
AI Search Visibility Report · methodology v1.0

How 4 AI assistants actually talk about Northwind Analytics.

75 real buyer questions, asked once each across Claude, Gemini, Openai, Perplexity, scored across the seven measurement layers of the Vector Labs methodology playbook.

Executive scorecard · 20-second read 7 layers · 4 engines · 75 prompts

Mid-pack across the engines. One layer is the bottleneck.

ChatGPT 36% / Claude 25% — the spread is narrower than typical, but consensus and accuracy carry the asymmetry. Detailed layer-by-layer view below.

L1
Eligibility4 fail · 3 warn
L2
PresenceStrong
L3
PreferenceMixed
L4
CitationOwned-mostly
L5
AccuracyConfirm 5 · 5 clusters
L6
ConsensusBreadth = 0
L7
BusinessFramework
Claude mention rate25%vs ChatGPT 36% — the headline gap
Off-site corroboration0authority-tier domains that cite + name you
Position when mentioned#1.4across all engines that named you
Claims needing confirmation5+ 5 clusters · see §05
Fix 01 · do this quarter
Publish a machine-readable facts file (llms.txt)
/llms.txtship the fix
Fix 02 · ~3 weeks
Organization + SoftwareApplication schema on the homepage (raw HTML)
/ship the fix
Fix 03 · this sprint
SoftwareApplication + Offer schema on the pricing page
/pricingship the fix
Snapshot · the audit in one screen

Snapshot · the four numbers a buyer scans before reading anything else

25% LOWEST
Claude
avg position 1.6
31% MENTION RATE
Gemini
avg position 1.3
36% MENTION RATE
ChatGPT
avg position 1.4
33% MENTION RATE
Perplexity
avg position 1.2
Category prompts
4/18
4 of 18 answers naming Northwind Analytics across 4 engines.
Share of voice · top 4 names
Datify48
Northwind An…42
MetricLab32
Vantyr26
Datify shows up in 48 of the 75 questions; Northwind Analytics in 42.
Citation rate · own domain
Claude20%
ChatGPT20%
Gemini16%
Perplexity16%
Avg 18%. Share of answers that cite Northwind Analytics's own domain.
01The verdict — three things to fix this quarter

The verdict · you're recommended wherever you appear — the problem is reach, not reputation

Two problems, one root cause. Both fix the same way.

So what Northwind is named in 28% of sampled AI answers to buyer questions (64 of 225 sampled responses, 5 engines, weekly run). Strongest surface: Gemini (43%). Weakest: Claude (16%). Datify leads you on ChatGPT and Copilot. Three fixes deployed earlier are verified live; 6 new fixes are drafted and waiting for review.
36% Chatgpt names Northwind Analytics
vs
25% Claude names Northwind Analytics
01

Publish a machine-readable facts file (llms.txt)

Adds a plain-text facts file at northwind.io/llms.txt that gathers the exact story in one place: what Northwind is, who it's for, the prices, the free plan, and how it compares to the bigger names. Honest framing: no major AI engine is known to read this file today, so this is low-impact hygiene — a single correct source of facts, useful to keep tidy, not a visibility lever.

Watch
Expected move /llms.txt
≤1 day client
02

Organization + SoftwareApplication schema on the homepage (raw HTML)

Puts an invisible "name tag" in the homepage code that spells out, in the format AI and search engines read first, that Northwind is a company and a software product, its price, its free plan, and its official links. Visitors see no change, but machines get the facts straight from the source instead of guessing.

Critical
Expected move /
≤1 day client
03

SoftwareApplication + Offer schema on the pricing page

Marks up the pricing page so AI and search engines can read the exact plan names and prices — Free, Pro at $49 a month, and Scale at $199 a month — as structured facts instead of scraping numbers off the page and sometimes getting them wrong. Nothing visible changes for shoppers.

Critical
Expected move /pricing
≤1 day client
02Visibility by engine — the headline gap

Layer 2 · Presence · how often each AI brings Northwind Analytics up on its own

4 of 4 surfaces name you. Only some of them actually read your website.

So what The mention rate is how often each AI even brings you up in an unbranded answer. The citation rate is how often it actually consults your site to do so. The widest gap is where the next fix lives.
Chatgpt
36%±11 · n=75
20%±9 · n=75
Perplexity
33%±10 · n=75
16%±8 · n=75
Gemini
31%±10 · n=75
16%±8 · n=75
Claude
25%±10 · n=75
20%±9 · n=75
Mention rate — how often the AI names Northwind Analytics Citation rate — how often the AI actually visits northwind.io

Intervals are Wilson 95%; ±5pp counts as noise below this sample size. A Gemini citation rate near zero alongside a non-zero mention rate is a surfacing artifact (Gemini routes through Vertex AI Search), not a crawl failure.

03Northwind Analytics vs competition — where the deal is lost or won

Layer 2 + 3 · Competitive presence + preference · how you stack up against the other names AI brings up

Datify is named 21pp more often than you. That gap is the brief.

So what Each AI orders the category differently. The brand-stack cards below show who an AI offers to a buyer when asked an unbranded question — and the heatmap shows which topics each competitor owns. Topic ownership is where the deal is lost or won.

Scoreboard · every brand the AIs named · rates measured on unbranded questions only

Brand Mention rate Citation rate Avg position Share-of-voice Positive framing
Northwind Analyticsyou 28%±5 17% #1.4 22% 38%
Datify 49%±6 #1.3 39% 45%
MetricLab 30%±5 #1.4 23% 41%
Vantyr 21%±5 #1.4 16% 47%

Competitive map · presence × recommendation quality

Recommendation quality  
Hidden gemsWell regarded but underexposed
LeadersOften named AND recommended
NicheLow reach, low recommendation
Volume playersMentioned often, not endorsed
Northwind Analytics
Datify
MetricLab
Vantyr
Presence    mention rate across engines
Northwind Analytics (you) Competitor X = mention rate on unbranded questions (rescaled to category leader). Y = inverted average answer position when named there. Italic = low framing sample.

How each AI ranks the options

Claude 75 answers · 85 brand mentions on unbranded questions
  1. 1 Datify 35×
  2. 2 MetricLab 20×
  3. 3 Northwind Analyticsyou 15×
  4. 4 Vantyr 15×
Gemini 75 answers · 89 brand mentions on unbranded questions
  1. 1 Datify 34×
  2. 2 MetricLab 21×
  3. 3 Northwind Analyticsyou 19×
  4. 4 Vantyr 15×
Openai 75 answers · 94 brand mentions on unbranded questions
  1. 1 Datify 35×
  2. 2 Northwind Analyticsyou 23×
  3. 3 MetricLab 21×
  4. 4 Vantyr 15×
Perplexity 75 answers · 90 brand mentions on unbranded questions
  1. 1 Datify 34×
  2. 2 Northwind Analyticsyou 21×
  3. 3 MetricLab 21×
  4. 4 Vantyr 14×

Topic × brand heatmap · where each competitor dominates

Mention rate 0 · invisible 1–20 · weak 21–40 · present 41+ · strong
Question topic
Northwind An… you
Datify
MetricLab
Vantyr
P1Datify alternatives
50
28
12
25
P1Product analytics tool comparison
35
65
35
30
P2Setup & integrations (Stripe, Postgres)
39
44
22
36
P3Analytics for bootstrapped startups
45
50
15
25
P3Analytics without a data team
22
69
16
16
P3Churn & retention reporting
30
80
35
15
P3Measuring user activation
20
10
50
5
P3Product analytics for small SaaS
15
60
42
17
P3Self-serve BI & dashboards
38
21
46
38
P3Tracking MRR & churn
28
40
20
10

Sentiment-when-mentioned · how each brand is framed

Northwind Analyticsyou
40 · 49 · 5 Mixed
Vantyr
31 · 30 · 3 Mixed
Datify
62 · 66 · 12 Mixed
MetricLab
34 · 44 · 6 Mixed

Tracked 4 brands · 3 configured, 0 auto-discovered from mention data. Positions in the brand stack are not the same as positions in the answer — stack rank = mention count, answer position = where in the answer text the brand appears.

04What AI says when it sees you

Layer 3 · Preference · the tone AI uses about you — and whether it actually recommends you

Sentiment varies sharply by engine. Some recommend you; others caveat you.

So what Positive framing ≥55% reads as a strong recommendation. Between 35-55% is mixed. Anything lower means the AI is hedging — usually because it has thin sources to anchor on.

Sentiment distribution per engine · % of unbranded mentions (questions that don't name the brand)

Perplexity
52 · 43 · 5 Mixed
Chatgpt
44 · 48 · 9 Mixed
Gemini
37 · 63 · 0 Mixed
Claude
13 · 80 · 7 Weak

Position rank is computed when the client is named alongside other brands — lower position = mentioned earlier in the answer. Low-sample rows are flagged because a single re-classification can flip the headline.

05Where AI is confused — claim risk

Layer 5 · Accuracy · what AI gets right, and wrong, about you — fact by fact

One claim cluster, 1 occurrences. It might not be wrong — confirm before we lock it.

So what Different surfaces tell the AI different stories. The result is a plausible composite answer that varies across engines. Lock one canonical machine-readable answer and the cluster collapses, regardless of which value is currently true.
The recurring claim · across 1 of 4 AIs Needs confirmation

Northwind is a UK company.

What needs confirmation
Northwind is a US-based company, founded in 2023.
Why the AI is reading it this way
Different surfaces give the AI different stories about this fact, so it synthesises a plausible composite. Lock one canonical machine-readable answer and the cluster collapses. Per Dong et al. (2015) [10], engine trust comes from cross-source consistency. When surfaces disagree, the AI synthesizes a plausible composite.
Claude·0× Gemini·0× Chatgpt·1× Perplexity·0×

More claims with verification status (not "wrong" until confirmed)

Needs confirmationClaude 1×

Northwind doesn't integrate with Stripe.

Status pending. Northwind integrates with Stripe (2026-06-04).

Needs confirmationGemini 1×

Northwind is a mobile-only app.

Status pending. Northwind is a web app that runs in any browser.

Needs confirmationPerplexity 1×

Northwind costs $99 per month.

Status pending. Northwind has a free plan; paid plans start at $49/mo.

Needs confirmationChatGPT 1×

Northwind doesn't work with Postgres.

Status pending. Postgres is supported.

Methodology · why we use these labels Every claim here is a candidate, not a verdict. The audit will not write "wrong" before the brand has confirmed the baseline — this is the same discipline FActScore (Min et al. 2023) [5] applies to atomic claim verification.
06Where AI is learning about you — the citation supply chain

Layer 4 · Citation + Layer 6 · Consensus · which sources AI reads — and whether anyone but your own site names you

The AI learns about you from a handful of domains. Most are owned or app-store.

So what The bottom of the table is the lever — domains that the AI cites for the category but never alongside your name. Land on 2-3 of them and consensus breadth goes from thin to diverse, which is what unblocks the lowest-mention engine.

Top 6 cited domains

reddit.comcommunity139
saaspicks.comweb132
saasweekly.comweb116
stackreviews.comweb115
g2.comweb98
northwind.ioyour own site54
Show the next 0 cited domains and the 5 missing — the Layer 6 gap
tweakers.netmissing0
frankwatching.commissing0
emerce.nlmissing0
bebright.eumissing0
deondernemer.nlmissing0
07Topic opportunity map — sorted by buyer intent

Layer 2 · Presence × topic · the buyer questions where you go missing, ranked by how much they matter

Some topics land; others are invisible. The zeroes are the content brief.

So what Buyer-intent rows are sorted highest-first. Red columns mark engines where you're invisible on a topic. Zero-coverage topics with active buyer demand are first-mover content opportunities — write them before competitors do.
Mention rate 0 · invisible 1–20 · weak 21–40 · present 41+ · strong
Question topic · buyer intent
Claude
Gemini
Openai
Perplexity
P1Datify alternatives · n=10
10
13
16
15
P1Product analytics tool comparison · n=5
23
29
35
32
P2Setup & integrations (Stripe, Postgres) · n=9
27
35
42
38
P3Product analytics for small SaaS · n=12 anchored
10
13
16
15
P3Analytics without a data team · n=8 anchored
15
19
22
21
P3Tracking MRR & churn · n=10 anchored
19
24
29
26
P3Self-serve BI & dashboards · n=6 fragile
23
29
35
32
P3Measuring user activation · n=5 fragile
15
19
22
21
P3Churn & retention reporting · n=5
19
24
29
26
P3Analytics for bootstrapped startups · n=5
27
35
42
38

Methodology §20.4: cells marked * sit on fewer than 5 source questions — insufficient sample; a single mention can swing them 0↔100. They are illustration, not measurement. Trust topic rows with 5+ questions; each row shows its question count.

08Eligibility — what your site tells the AI

Layer 1 · Eligibility · what your website does, and doesn't, tell AI about you

The AI can reach your site. It just can't tell what you actually do.

So what 4 fails and 3 warns across the technical checks. The schema and store-listing items below are usually the highest-impact, lowest-effort wins — they're what the AI parses to describe your product.

Crawler access · which AI bots can reach the site

Every AI crawler is allowed in
17 of 17
Pass
robots.txt blocks only junk, never content
6 rules
Pass

Schema.org JSON-LD · the structured-data layer

Schema is JavaScript-only
22 of 34
Critical
Visible FAQs carry no FAQPage markup
0 of 6
Critical

Other technical checks

Breadcrumbs and site structure validate cleanly
0 errors
Pass
Titles are slogans, not facts
14 of 34
Critical
No meta description at all
8 of 34
Watch
Canonical links correct on every page
34 of 34
Pass
Every date is over a year old
35 of 35
Critical
Indexable pages missing from sitemap
3 of 34
Watch
Images have no alt text
18 images
Watch
HTTPS everywhere, no orphan pages
34 of 34
Pass

Verify any of this yourself — open view-source:northwind.io and search for application/ld+json · open northwind.io/robots.txt · open northwind.io/llms.txt.

09Fix roadmap — impact × effort

What to do first · the biggest payoff for the least effort

The whole roadmap. One page. Three columns of work.

So what Quick wins go first — high impact, ≤1 day of work. Strategic bets are 2–3 weeks but unlock the lowest-mention engine. Maintenance is housekeeping; defer is honest about what's not worth doing this quarter.
Impact →
Quick wins High impact · low effort
02Organization + SoftwareApplication schema on the… 03SoftwareApplication + Offer schema on the pricin… 11FAQ loads via fetch() — move all 14 answers into… 12Homepage headings are styled <div>s — restore re… 14The comparison table lazy-loads on scroll — rend… 15Pricing details hidden behind a "See plans" moda… 29Outreach: get added to the "Datify alternatives"… 30Reddit: answer an r/SaaS thread asking for analy…
Strategic bet High impact · high effort
04Rewrite 18 page titles + descriptions to entity-… 18Main content sits in a JS-hydrated <div id=app> … 19New page: "What is product analytics?" (definiti… 20Comparison page: Northwind vs Datify 21Article: "How to track MRR without a data team" 28Pitch: the "best product analytics tools 2026" l… 32Request G2 reviews from 5 activated customers
Maintenance Low impact · low effort
01Publish a machine-readable facts file (llms.txt) 05Fix duplicate /features vs /product canonical ta… 06Generate and submit an XML sitemap 07Add hreflang tags for /us and /uk pages 08Collapse a 3-hop redirect chain on /docs → /help 10Fix 6 broken links to the deprecated /changelog 13Wrap the feature grid in semantic <section>/<art… 16Customer-logos carousel is JS-only — flatten to … 17Blog index paginates via JS — add crawlable page… 22Answer "Does it work with Stripe and Postgres?" … 23Add a TL;DR summary box to the top of the pricin… 24Rewrite vague feature headings into question-for… 25Expand the thin /integrations page (currently 40… 26Fix an outdated claim: "30+ integrations" → "40+… 27Lock the confirmed pricing facts so drafts stop … 31LinkedIn: founder post on "analytics without a d…
Defer Low impact · high effort
09Add alt text to 27 product-screenshot images 33Pitch a hands-on review to a SaaS-tools blogger
Effort →

Ranked table · top 3 with detailed guides

#FixImpactEffortOwnerExpected move
01Publish a machine-readable facts file (llms.txt)Med≤1 dayclient/llms.txtvia Sprint
02Organization + SoftwareApplication schema on the homepage (raw HTML)High≤1 dayclient/via Sprint
03SoftwareApplication + Offer schema on the pricing pageHigh≤1 dayclient/pricingvia Sprint
04Rewrite 18 page titles + descriptions to entity-rich formatHigh≤1 wkclienttemplate: 18 pagesvia Sprint
05Fix duplicate /features vs /product canonical tagsMed≤1 dayclient/featuresvia Sprint
06Generate and submit an XML sitemapMed≤1 dayclient/sitemap.xmlvia Sprint
07Add hreflang tags for /us and /uk pagesMed≤1 dayclienttemplatevia Sprint
08Collapse a 3-hop redirect chain on /docs → /helpMed≤1 dayclient/docsvia Sprint
09Add alt text to 27 product-screenshot imagesMed≤1 wkclient/productvia Sprint
10Fix 6 broken links to the deprecated /changelogMed≤1 dayclientsite-widevia Sprint
11FAQ loads via fetch() — move all 14 answers into the DOM (<details>)High≤1 dayclient/pricingvia Sprint
12Homepage headings are styled <div>s — restore real <h1>/<h2>High≤1 dayclient/via Sprint
13Wrap the feature grid in semantic <section>/<article>Med≤1 dayclient/featuresvia Sprint
14The comparison table lazy-loads on scroll — render it in first HTMLHigh≤1 dayclient/comparevia Sprint
15Pricing details hidden behind a "See plans" modal — surface in HTMLHigh≤1 dayclient/pricingvia Sprint
16Customer-logos carousel is JS-only — flatten to a static listMed≤1 dayclient/via Sprint
17Blog index paginates via JS — add crawlable page linksMed≤1 dayclient/blogvia Sprint
18Main content sits in a JS-hydrated <div id=app> — SSR the body copyHigh≤1 wkclient/via Sprint
19New page: "What is product analytics?" (definitional)High≤1 wkclient/guides/product-analyticsvia Sprint
20Comparison page: Northwind vs DatifyHigh≤1 wkclient/compare/northwind-vs-datifyvia Sprint
21Article: "How to track MRR without a data team"High≤1 wkclient/blog/track-mrr-without-a-data-teamvia Sprint
22Answer "Does it work with Stripe and Postgres?" on the FAQMed≤1 dayclient/faqvia Sprint
23Add a TL;DR summary box to the top of the pricing pageMed≤1 dayclient/pricingvia Sprint
24Rewrite vague feature headings into question-form headingsMed≤1 dayclient/featuresvia Sprint
25Expand the thin /integrations page (currently 40 words)Med≤1 dayclient/integrationsvia Sprint
26Fix an outdated claim: "30+ integrations" → "40+"Med≤1 dayclient/via Sprint
27Lock the confirmed pricing facts so drafts stop guessingMed≤1 dayclientfactsvia Sprint
28Pitch: the "best product analytics tools 2026" listicle omits youHigh≤1 wkclient"best product analytics tools 2026" listiclevia Sprint
29Outreach: get added to the "Datify alternatives" roundupHigh≤1 dayclientexternal roundupvia Sprint
30Reddit: answer an r/SaaS thread asking for analytics toolsHigh≤1 dayclientreddit.com/r/SaaSvia Sprint
31LinkedIn: founder post on "analytics without a data team"Med≤1 dayclientlinkedinvia Sprint
32Request G2 reviews from 5 activated customersHigh≤1 wkclientg2.comvia Sprint
33Pitch a hands-on review to a SaaS-tools bloggerMed≤1 wkclientexternal blogvia Sprint
1090-day forecast — before / after

Layer 7 · Business impact · what better visibility could be worth — a framework, not made-up numbers

Where each engine could go — if you ship the top three by day 60.

So what The biggest mover is the engine that's furthest behind. Fixing consensus + the schema layer lifts that engine 10+ points; the others move 5–10 on the back of the shared fixes. Movement under ~5pp counts as noise.

Per-AI mention rate · baseline → 90-day directional band

Framework · not a promise

How to read this: bands are directional — "meaningful" means there's room to move 10+pp if the top fixes ship and engines don't drift; "modest" is 5–10pp; "noise" is <5pp (below the variance threshold). The audit can't see your analytics; we measure engines, not outcomes.

Claude
25%±5pp band
noise band (already strong)
Gemini
31%±5pp band
noise band (already strong)
Perplexity
33%±5pp band
noise band (already strong)
Chatgpt
36%±5pp band
noise band (already strong)

Layer 7 (business impact) ships as a framework, not fabricated numbers. The audit can't see your analytics. Re-measurement separates your movement from engine drift.

Layer 7 (business impact) ships as a framework, not fabricated numbers. The audit cannot see your app analytics. The instrumentation spec is detailed in §12 of the methodology playbook.

11The honest perimeter — what this audit can't tell you

How far to trust these numbers — and what this audit can't tell you

Seven things to keep in mind. Knowing where the lens distorts is how you avoid acting on noise.

Variance budget. Identical prompts vary 40–60% run to run (Sielinski 2026 [2]). Every headline carries a Wilson 95% interval computed on distinct questions — never on answer counts inflated by repeats (methodology §20.2). Movement under ~5pp, or under this run-pair's own printed noise floor if that is larger, is noise.

No blended score. Per-layer only — a single 0–100 number hides which layer is broken.

API ≠ consumer app. Independent measurement: ~24% brand overlap, ~4% source overlap. This run asked the engines through their APIs. What a person sees in the ChatGPT or Gemini app can differ, and this run does not measure that.

Engines change without notice. A model update can move numbers 10+ points in a week. Re-measurement separates your movement from engine drift.

Share-of-Voice is unweighted. All mentions count equally — a passing reference and a ranked recommendation are weighted the same. v1.1 will weight by framing.

Baseline-dependent. §05 claims are 'needs confirmation,' not proven errors, until the team confirms and dates the fact baseline.

No outcome guarantees, on purpose. We disclose what was asked, when, from where, and via which mode — never what an engine will say tomorrow.

12Methodology & references — appendix

Vector Labs methodology playbook · v1.0

Seven layers. Per-layer scoring. Intervals, not points.

Every metric in this report is defined in the playbook with five attributes: what it is, how it's measured (formula + data), how to read it (thresholds), how it fails (controls), and the research it rests on. Below: the layered architecture in summary.

A note on the L1–L7 framing The canonical methodology playbook (measurement-methodology.md) is organised into 9 numbered sections (query design, engines, run protocol, metrics, aggregation, technical checks, synthesis, re-runs, limitations). The L1–L7 layer model below is a presentation overlay that re-aggregates those same metrics into seven named layers for buyer readability. It is not a separate methodology; every layer status traces to the methodology metric it summarises.
L1
Eligibility

Can the engines access, parse, and trust the content?

Binary gates: AI-crawler access (RFC 9309), llms.txt presence, Schema.org coverage, indexability, factual self-description, canonical fact consistency, off-site corroboration, store listings.

Output: checklist with severity (pass / warn / fail). Never a chart.

L2
Presence

Does the brand appear at all?

Mention rate per engine on unbranded prompt buckets, Wilson 95% intervals. Topic × engine heatmap. Average position when named.

Formula: mention_rate(e) = (# answers naming the brand) / |Q|. Headline computed on unbranded buckets only.

L3
Preference

Recommended, or merely present?

Recommendation rate; 0–5 ordinal ladder. Distribution shown, never collapsed to a mean.

Verifier-judged with position / verbosity / self-enhancement bias controls (Zheng et al. 2023 [8]). κ tracked (Cohen 1960 [9]).

L4
Citation

Is the brand used as a cited source?

Citation precision and recall (Liu, Zhang & Liang 2023 [4]). Owned vs earned. Ranked supply chain.

citation_rate(e) = answers citing brand source / |Q|. Surfacing artifacts annotated, not scored as absence.

L5
Absorption & accuracy

Described correctly and completely?

Claim verification ledger (FActScore atomic decomposition, Min et al. 2023 [5]). Statuses: verified true / false / unsupported / stale / needs confirmation. Semantic completeness.

Cardinal rule: never label a claim a hallucination because the verifier baseline lacks it. Confirm first.

L6
Consensus / unified authority

Do independent sources corroborate?

Corroboration breadth. Per-fact consensus. Authority-weighted to category domains the engines trust. Theoretical anchor: Dong et al. 2015 Knowledge-Based Trust [10].

Breadth = 0 means no consensus web to triangulate. Root cause of the "reach, not reputation" thesis.

L7
Business impact

Does any of this produce leads?

Audit is diagnosis; client analytics is thermometer. Ships as framework. Chain: AI referral → site / store view → leads → retention signal.

Instrumentation: GA4 AI-host segmentation, attribution, in-app or post-conversion "how did you hear" prompt, branded-impression trend.

The full methodology playbook Every formula, threshold, control and citation is in measurement-methodology.md at the audit's run archive. This report is the application; the playbook is the specification.

References

Aggarwal, P., et al. (2024). GEO: Generative Engine Optimization. KDD '24. arXiv:2311.09735.

Sielinski, R. (2026). Quantifying Uncertainty in AI Visibility. arXiv:2603.08924.

"Don't Measure Once." (2026). arXiv:2604.07585.

Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating Verifiability in Generative Search Engines. EMNLP 2023.

Min, S., et al. (2023). FActScore. EMNLP 2023.

Es, S., et al. (2024). RAGAs. EACL 2024.

Manakul, P., et al. (2023). SelfCheckGPT. EMNLP 2023.

Zheng, L., et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023.

Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales.

Dong, X. L., et al. (2015). Knowledge-Based Trust. VLDB 8(9).

Wilson, E. B. (1927). The score interval for a binomial proportion.

Brown, L. D., Cai, T. T., & DasGupta, A. (2001). Interval Estimation for a Binomial Proportion.

Efron, B. (1979). Bootstrap Methods.

Howard, J. (2024). The /llms.txt file.

Koster, M., et al. (2022). RFC 9309: Robots Exclusion Protocol.

Schema.org. Structured-data vocabulary.

Google. Search Quality Rater Guidelines — E-E-A-T.

The diagnosis is in your hands. The execution is the next call.

What you have — delivered

  • This report75 buyer questions across 4 AI assistants, with the answers behind every number.
  • 33 fixes, rankedImpact against effort, each one tied to the finding that produced it.
  • The questions, frozenThe same list runs next time, so the next report is a comparison and not a new opinion.

Keep measuring from €99/month

  • Monitor — €99/monthThe same questions asked again on a schedule, so you see what moved and what caused it.
  • Fix — €279/monthWe also draft the changes. Every one waits for your approval before it touches your site.
  • Monthly, cancel anytimeNo setup fee and no minimum term.