How 4 AI assistants actually talk about Northwind Analytics.
75 real buyer questions, asked once each across Claude, Gemini, Openai, Perplexity, scored across the seven measurement layers of the Vector Labs methodology playbook.
Mid-pack across the engines. One layer is the bottleneck.
ChatGPT 36% / Claude 25% — the spread is narrower than typical, but consensus and accuracy carry the asymmetry. Detailed layer-by-layer view below.
Snapshot · the four numbers a buyer scans before reading anything else
The verdict · you're recommended wherever you appear — the problem is reach, not reputation
Two problems, one root cause. Both fix the same way.
Publish a machine-readable facts file (llms.txt)
Adds a plain-text facts file at northwind.io/llms.txt that gathers the exact story in one place: what Northwind is, who it's for, the prices, the free plan, and how it compares to the bigger names. Honest framing: no major AI engine is known to read this file today, so this is low-impact hygiene — a single correct source of facts, useful to keep tidy, not a visibility lever.
Organization + SoftwareApplication schema on the homepage (raw HTML)
Puts an invisible "name tag" in the homepage code that spells out, in the format AI and search engines read first, that Northwind is a company and a software product, its price, its free plan, and its official links. Visitors see no change, but machines get the facts straight from the source instead of guessing.
SoftwareApplication + Offer schema on the pricing page
Marks up the pricing page so AI and search engines can read the exact plan names and prices — Free, Pro at $49 a month, and Scale at $199 a month — as structured facts instead of scraping numbers off the page and sometimes getting them wrong. Nothing visible changes for shoppers.
Layer 2 · Presence · how often each AI brings Northwind Analytics up on its own
4 of 4 surfaces name you. Only some of them actually read your website.
Intervals are Wilson 95%; ±5pp counts as noise below this sample size. A Gemini citation rate near zero alongside a non-zero mention rate is a surfacing artifact (Gemini routes through Vertex AI Search), not a crawl failure.
Layer 2 + 3 · Competitive presence + preference · how you stack up against the other names AI brings up
Datify is named 21pp more often than you. That gap is the brief.
Scoreboard · every brand the AIs named · rates measured on unbranded questions only
| Brand | Mention rate | Citation rate | Avg position | Share-of-voice | Positive framing |
|---|---|---|---|---|---|
| Northwind Analyticsyou | 28%±5 | 17% | #1.4 | 22% | 38% |
| Datify | 49%±6 | — | #1.3 | 39% | 45% |
| MetricLab | 30%±5 | — | #1.4 | 23% | 41% |
| Vantyr | 21%±5 | — | #1.4 | 16% | 47% |
Competitive map · presence × recommendation quality
How each AI ranks the options
- 1 Datify 35×
- 2 MetricLab 20×
- 3 Northwind Analyticsyou 15×
- 4 Vantyr 15×
- 1 Datify 34×
- 2 MetricLab 21×
- 3 Northwind Analyticsyou 19×
- 4 Vantyr 15×
- 1 Datify 35×
- 2 Northwind Analyticsyou 23×
- 3 MetricLab 21×
- 4 Vantyr 15×
- 1 Datify 34×
- 2 Northwind Analyticsyou 21×
- 3 MetricLab 21×
- 4 Vantyr 14×
Topic × brand heatmap · where each competitor dominates
Sentiment-when-mentioned · how each brand is framed
Tracked 4 brands · 3 configured, 0 auto-discovered from mention data. Positions in the brand stack are not the same as positions in the answer — stack rank = mention count, answer position = where in the answer text the brand appears.
Layer 3 · Preference · the tone AI uses about you — and whether it actually recommends you
Sentiment varies sharply by engine. Some recommend you; others caveat you.
Sentiment distribution per engine · % of unbranded mentions (questions that don't name the brand)
Position rank is computed when the client is named alongside other brands — lower position = mentioned earlier in the answer. Low-sample rows are flagged because a single re-classification can flip the headline.
Layer 5 · Accuracy · what AI gets right, and wrong, about you — fact by fact
One claim cluster, 1 occurrences. It might not be wrong — confirm before we lock it.
Northwind is a UK company.
More claims with verification status (not "wrong" until confirmed)
Northwind doesn't integrate with Stripe.
Status pending. Northwind integrates with Stripe (2026-06-04).
Northwind is a mobile-only app.
Status pending. Northwind is a web app that runs in any browser.
Northwind costs $99 per month.
Status pending. Northwind has a free plan; paid plans start at $49/mo.
Northwind doesn't work with Postgres.
Status pending. Postgres is supported.
Layer 4 · Citation + Layer 6 · Consensus · which sources AI reads — and whether anyone but your own site names you
The AI learns about you from a handful of domains. Most are owned or app-store.
Top 6 cited domains
Show the next 0 cited domains and the 5 missing — the Layer 6 gap
Layer 2 · Presence × topic · the buyer questions where you go missing, ranked by how much they matter
Some topics land; others are invisible. The zeroes are the content brief.
Methodology §20.4: cells marked * sit on fewer than 5 source questions — insufficient sample; a single mention can swing them 0↔100. They are illustration, not measurement. Trust topic rows with 5+ questions; each row shows its question count.
Layer 1 · Eligibility · what your website does, and doesn't, tell AI about you
The AI can reach your site. It just can't tell what you actually do.
Crawler access · which AI bots can reach the site
Schema.org JSON-LD · the structured-data layer
Other technical checks
Verify any of this yourself — open view-source:northwind.io and search for application/ld+json · open northwind.io/robots.txt · open northwind.io/llms.txt.
What to do first · the biggest payoff for the least effort
The whole roadmap. One page. Three columns of work.
Ranked table · top 3 with detailed guides
| # | Fix | Impact | Effort | Owner | Expected move | |
|---|---|---|---|---|---|---|
| 01 | Publish a machine-readable facts file (llms.txt) | Med | ≤1 day | client | /llms.txt | via Sprint |
| 02 | Organization + SoftwareApplication schema on the homepage (raw HTML) | High | ≤1 day | client | / | via Sprint |
| 03 | SoftwareApplication + Offer schema on the pricing page | High | ≤1 day | client | /pricing | via Sprint |
| 04 | Rewrite 18 page titles + descriptions to entity-rich format | High | ≤1 wk | client | template: 18 pages | via Sprint |
| 05 | Fix duplicate /features vs /product canonical tags | Med | ≤1 day | client | /features | via Sprint |
| 06 | Generate and submit an XML sitemap | Med | ≤1 day | client | /sitemap.xml | via Sprint |
| 07 | Add hreflang tags for /us and /uk pages | Med | ≤1 day | client | template | via Sprint |
| 08 | Collapse a 3-hop redirect chain on /docs → /help | Med | ≤1 day | client | /docs | via Sprint |
| 09 | Add alt text to 27 product-screenshot images | Med | ≤1 wk | client | /product | via Sprint |
| 10 | Fix 6 broken links to the deprecated /changelog | Med | ≤1 day | client | site-wide | via Sprint |
| 11 | FAQ loads via fetch() — move all 14 answers into the DOM (<details>) | High | ≤1 day | client | /pricing | via Sprint |
| 12 | Homepage headings are styled <div>s — restore real <h1>/<h2> | High | ≤1 day | client | / | via Sprint |
| 13 | Wrap the feature grid in semantic <section>/<article> | Med | ≤1 day | client | /features | via Sprint |
| 14 | The comparison table lazy-loads on scroll — render it in first HTML | High | ≤1 day | client | /compare | via Sprint |
| 15 | Pricing details hidden behind a "See plans" modal — surface in HTML | High | ≤1 day | client | /pricing | via Sprint |
| 16 | Customer-logos carousel is JS-only — flatten to a static list | Med | ≤1 day | client | / | via Sprint |
| 17 | Blog index paginates via JS — add crawlable page links | Med | ≤1 day | client | /blog | via Sprint |
| 18 | Main content sits in a JS-hydrated <div id=app> — SSR the body copy | High | ≤1 wk | client | / | via Sprint |
| 19 | New page: "What is product analytics?" (definitional) | High | ≤1 wk | client | /guides/product-analytics | via Sprint |
| 20 | Comparison page: Northwind vs Datify | High | ≤1 wk | client | /compare/northwind-vs-datify | via Sprint |
| 21 | Article: "How to track MRR without a data team" | High | ≤1 wk | client | /blog/track-mrr-without-a-data-team | via Sprint |
| 22 | Answer "Does it work with Stripe and Postgres?" on the FAQ | Med | ≤1 day | client | /faq | via Sprint |
| 23 | Add a TL;DR summary box to the top of the pricing page | Med | ≤1 day | client | /pricing | via Sprint |
| 24 | Rewrite vague feature headings into question-form headings | Med | ≤1 day | client | /features | via Sprint |
| 25 | Expand the thin /integrations page (currently 40 words) | Med | ≤1 day | client | /integrations | via Sprint |
| 26 | Fix an outdated claim: "30+ integrations" → "40+" | Med | ≤1 day | client | / | via Sprint |
| 27 | Lock the confirmed pricing facts so drafts stop guessing | Med | ≤1 day | client | facts | via Sprint |
| 28 | Pitch: the "best product analytics tools 2026" listicle omits you | High | ≤1 wk | client | "best product analytics tools 2026" listicle | via Sprint |
| 29 | Outreach: get added to the "Datify alternatives" roundup | High | ≤1 day | client | external roundup | via Sprint |
| 30 | Reddit: answer an r/SaaS thread asking for analytics tools | High | ≤1 day | client | reddit.com/r/SaaS | via Sprint |
| 31 | LinkedIn: founder post on "analytics without a data team" | Med | ≤1 day | client | via Sprint | |
| 32 | Request G2 reviews from 5 activated customers | High | ≤1 wk | client | g2.com | via Sprint |
| 33 | Pitch a hands-on review to a SaaS-tools blogger | Med | ≤1 wk | client | external blog | via Sprint |
Layer 7 · Business impact · what better visibility could be worth — a framework, not made-up numbers
Where each engine could go — if you ship the top three by day 60.
Per-AI mention rate · baseline → 90-day directional band
Framework · not a promiseHow to read this: bands are directional — "meaningful" means there's room to move 10+pp if the top fixes ship and engines don't drift; "modest" is 5–10pp; "noise" is <5pp (below the variance threshold). The audit can't see your analytics; we measure engines, not outcomes.
Layer 7 (business impact) ships as a framework, not fabricated numbers. The audit can't see your analytics. Re-measurement separates your movement from engine drift.
Layer 7 (business impact) ships as a framework, not fabricated numbers. The audit cannot see your app analytics. The instrumentation spec is detailed in §12 of the methodology playbook.
How far to trust these numbers — and what this audit can't tell you
Seven things to keep in mind. Knowing where the lens distorts is how you avoid acting on noise.
Variance budget. Identical prompts vary 40–60% run to run (Sielinski 2026 [2]). Every headline carries a Wilson 95% interval computed on distinct questions — never on answer counts inflated by repeats (methodology §20.2). Movement under ~5pp, or under this run-pair's own printed noise floor if that is larger, is noise.
No blended score. Per-layer only — a single 0–100 number hides which layer is broken.
API ≠ consumer app. Independent measurement: ~24% brand overlap, ~4% source overlap. This run asked the engines through their APIs. What a person sees in the ChatGPT or Gemini app can differ, and this run does not measure that.
Engines change without notice. A model update can move numbers 10+ points in a week. Re-measurement separates your movement from engine drift.
Share-of-Voice is unweighted. All mentions count equally — a passing reference and a ranked recommendation are weighted the same. v1.1 will weight by framing.
Baseline-dependent. §05 claims are 'needs confirmation,' not proven errors, until the team confirms and dates the fact baseline.
No outcome guarantees, on purpose. We disclose what was asked, when, from where, and via which mode — never what an engine will say tomorrow.
Vector Labs methodology playbook · v1.0
Seven layers. Per-layer scoring. Intervals, not points.
Every metric in this report is defined in the playbook with five attributes: what it is, how it's measured (formula + data), how to read it (thresholds), how it fails (controls), and the research it rests on. Below: the layered architecture in summary.
measurement-methodology.md) is organised into 9 numbered sections (query design, engines, run protocol, metrics, aggregation, technical checks, synthesis, re-runs, limitations). The L1–L7 layer model below is a presentation overlay that re-aggregates those same metrics into seven named layers for buyer readability. It is not a separate methodology; every layer status traces to the methodology metric it summarises.
Can the engines access, parse, and trust the content?
Binary gates: AI-crawler access (RFC 9309), llms.txt presence, Schema.org coverage, indexability, factual self-description, canonical fact consistency, off-site corroboration, store listings.
Output: checklist with severity (pass / warn / fail). Never a chart.
Does the brand appear at all?
Mention rate per engine on unbranded prompt buckets, Wilson 95% intervals. Topic × engine heatmap. Average position when named.
Formula: mention_rate(e) = (# answers naming the brand) / |Q|. Headline computed on unbranded buckets only.
Is the brand used as a cited source?
Citation precision and recall (Liu, Zhang & Liang 2023 [4]). Owned vs earned. Ranked supply chain.
citation_rate(e) = answers citing brand source / |Q|. Surfacing artifacts annotated, not scored as absence.
Described correctly and completely?
Claim verification ledger (FActScore atomic decomposition, Min et al. 2023 [5]). Statuses: verified true / false / unsupported / stale / needs confirmation. Semantic completeness.
Cardinal rule: never label a claim a hallucination because the verifier baseline lacks it. Confirm first.
Do independent sources corroborate?
Corroboration breadth. Per-fact consensus. Authority-weighted to category domains the engines trust. Theoretical anchor: Dong et al. 2015 Knowledge-Based Trust [10].
Breadth = 0 means no consensus web to triangulate. Root cause of the "reach, not reputation" thesis.
Does any of this produce leads?
Audit is diagnosis; client analytics is thermometer. Ships as framework. Chain: AI referral → site / store view → leads → retention signal.
Instrumentation: GA4 AI-host segmentation, attribution, in-app or post-conversion "how did you hear" prompt, branded-impression trend.
measurement-methodology.md at the audit's run archive. This report is the application; the playbook is the specification.
References
Aggarwal, P., et al. (2024). GEO: Generative Engine Optimization. KDD '24. arXiv:2311.09735.
Sielinski, R. (2026). Quantifying Uncertainty in AI Visibility. arXiv:2603.08924.
"Don't Measure Once." (2026). arXiv:2604.07585.
Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating Verifiability in Generative Search Engines. EMNLP 2023.
Min, S., et al. (2023). FActScore. EMNLP 2023.
Es, S., et al. (2024). RAGAs. EACL 2024.
Manakul, P., et al. (2023). SelfCheckGPT. EMNLP 2023.
Zheng, L., et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023.
Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales.
Dong, X. L., et al. (2015). Knowledge-Based Trust. VLDB 8(9).
Wilson, E. B. (1927). The score interval for a binomial proportion.
Brown, L. D., Cai, T. T., & DasGupta, A. (2001). Interval Estimation for a Binomial Proportion.
Efron, B. (1979). Bootstrap Methods.
Howard, J. (2024). The /llms.txt file.
Koster, M., et al. (2022). RFC 9309: Robots Exclusion Protocol.
Schema.org. Structured-data vocabulary.
Google. Search Quality Rater Guidelines — E-E-A-T.
The diagnosis is in your hands. The execution is the next call.
What you have — delivered
- This report75 buyer questions across 4 AI assistants, with the answers behind every number.
- 33 fixes, rankedImpact against effort, each one tied to the finding that produced it.
- The questions, frozenThe same list runs next time, so the next report is a comparison and not a new opinion.
Keep measuring from €99/month
- Monitor — €99/monthThe same questions asked again on a schedule, so you see what moved and what caused it.
- Fix — €279/monthWe also draft the changes. Every one waits for your approval before it touches your site.
- Monthly, cancel anytimeNo setup fee and no minimum term.