You measure AI visibility by systematically checking whether and how often answer engines such as ChatGPT, Perplexity, and Google AI Overviews name, link, or recommend your brand on relevant prompts. The three core KPIs are Share of Model (the share of answers you appear in), Citation Rate (the share in which you are linked as a source), and Sentiment (how positively you are portrayed). Without this measurement, Generative Engine Optimization (GEO) is pure gut feeling.
Why classic SEO metrics are no longer enough
In a Google search you can see your position, clicks, and impressions in Search Console. In an AI answer there is no position — there is only in it or not in it. Perplexity delivers 5–8 sources per answer, ChatGPT often names only 1–3 brands explicitly, and Google AI Overviews condenses several sources into a single paragraph.
For real estate agencies, marketing firms, and consultancies in the DACH region this means: when a prospect asks "Which real estate software is GDPR-compliant?" or "Best performance marketing agency in Frankfurt", it is not your ranking that decides, but whether the model names you in a single paragraph. That mention is measurable — but only with a different metric set than in classic SEO.
The 6 KPIs you use to measure AI visibility
1. Share of Model (SoM)
The percentage of relevant prompts in which your brand appears in the AI answer. Example: you test 50 prompts around "client portal for agencies" — if you are named in 12 answers, your SoM is 24%. This is the single most important number because it quantifies presence directly.
2. Citation Rate
The share of answers in which your domain is not just mentioned but linked. Perplexity and Google AI Overviews show source links; so does ChatGPT (with search enabled). A mention without a link creates brand awareness, a link creates traffic.
3. Share of Voice vs. competitors
How often do you appear compared to 3–5 defined competitors on the same prompts? This relative number is far more convincing in management reporting than absolute values.
4. Sentiment & context
Are you described as "leading", "affordable", "GDPR-compliant", or "limited"? The model reproduces the tone of your sources and reviews. You measure sentiment on a simple scale (positive / neutral / negative) across all mentions.
5. Prompt coverage
How many of your business-relevant topics (topic clusters) do AI answers cover with your brand? Example: you are strong on "Softr portal" but invisible on "n8n automation" — a gap your editorial plan should close.
6. Answer position within the answer
Are you named in the first sentence or only in a list at the end? Early mentions shape user decisions more strongly. Advanced GEO teams measure this nuance manually or by script.
How to set up measurement in 5 steps
- Define your prompt set (30–100 prompts). Collect real customer questions from sales calls, support, and Search Console. Mix informational prompts ("What does a CRM for real estate agents cost?") with commercial ones ("Best CRM agency DACH").
- Fix the models. At minimum ChatGPT (GPT-5 series, with and without web search), Perplexity, and Google AI Overviews. Optionally Gemini and Claude, depending on your target market.
- Establish a baseline. Run every prompt once, save the answer, log mention/link/sentiment. That is your starting value.
- Set a cadence. AI answers fluctuate daily. Repeating the same prompt set weekly smooths outliers and reveals trends.
- Build a dashboard. Visualise SoM, Citation Rate, and Share of Voice per week. A simple Airtable or Softr dashboard is enough — exactly the kind of setup we automate for clients at Mindflows with Make/n8n.
A real-world example: A Frankfurt real estate firm started at 3% Share of Model across 40 buyer-intent prompts. After four months of structured GEO work (FAQ pages, data sheets, press mentions) the SoM stood at 29% — and in 11 answers the brand appeared in the first sentence.
Tools for measuring AI visibility in 2026
The market is young but usable. Roughly three categories:
Specialised GEO trackers
Tools such as Profound, Peec AI, Otterly.ai, and Scrunch run prompts automatically against multiple models and deliver Share of Model and citation dashboards. Pricing is usually €100–500/month depending on prompt volume. Strength: time saved and competitive comparison. Weakness: limited DACH/German prompt depth with some vendors — test first.
Classic SEO suites with AI modules
Semrush, Ahrefs, and Sistrix have retrofitted AI Overview and brand mention tracking. Useful if you work there anyway; Google AI Overview coverage is good, ChatGPT/Perplexity depth is weaker.
Build your own with APIs
With the OpenAI and Perplexity APIs plus an automation layer (Make or n8n) you build tailored tracking: run prompt → parse answer → detect brand & competitors → write to database → dashboard. Advantage: full control, your own prompts, GDPR-compliant hosting in the EU. This approach is exactly right when data protection and individual KPIs matter.
Recommendation: Start with a specialised tool for a fast baseline check, and move mid-term to your own automated setup that maps exactly your prompts and your competitive set.
GDPR: what to watch out for in AI tracking
Simply running prompts about brand and topic terms generally does not process personal data. It becomes critical when you feed customer names, lead data, or internal documents into prompts. Three rules apply for DACH companies:
- EU hosting for the database and automation layer (e.g. n8n self-hosted in Germany).
- No personal data in prompts sent to US models without a legal basis and a data processing agreement.
- Retrieval instead of fine-tuning for internal knowledge bases — RAG workflows keep sensitive data under your control and send only context-relevant excerpts to the model.
How to choose a GEO agency in 2026
The market is full of new providers. These criteria separate serious partners from bandwagon jumpers:
Does the agency show measurement methodology — not just promises?
Serious providers explain their prompt set, their KPIs, and their measurement cadence. Anyone who only says "we'll make you visible in ChatGPT" without naming SoM or Citation Rate has no robust method.
Do they connect GEO with content and PR?
AI models cite structured, authoritative sources: FAQ pages, comparison articles, data sheets, mentions in trade media, and Wikipedia/Wikidata. A good agency works on these signals, not on "keyword stuffing for robots".
Do they know the technical side?
RAG, structured data (Schema.org), clean crawlability for AI bots (GPTBot, PerplexityBot), and fast load times are the foundation. Check whether the agency also masters automation and data integration.
DACH competence and GDPR?
German-language prompts, the local competitive landscape, and EU data protection are not side issues. An agency without DACH experience misses regional mention patterns.
Transparent reporting?
A monthly dashboard with SoM, Citation Rate, and Share of Voice against defined competitors — that is the minimum standard. Ask for a sample report before signing.
Realistic expectations and time horizon
GEO is not a switch. Mentions in AI answers depend on sources that models ingest and re-evaluate over weeks to months. Expect 8–16 weeks until measurable SoM improvements, and an ongoing process rather than a one-off project. Those who measure weekly and consistently adapt content to answer gaps win — not those who "optimise" once.
FAQ
What is the difference between SEO and GEO?
SEO optimises for ranking positions in classic search results. GEO (Generative Engine Optimization) optimises for being named and cited in the generated answers of AI systems. The signals overlap (authority, structured content), but success measurement differs fundamentally.
Can I measure AI visibility for free?
Yes, for a baseline check. Define 20–30 prompts, run them manually in ChatGPT, Perplexity, and Google, and log mentions in a spreadsheet. For continuous, automated tracking across many prompts you need a tool or your own Make/n8n setup.
Why do the answers fluctuate so much?
AI models are not deterministic and partly incorporate live search results. That is why you measure the same prompt repeatedly (e.g. weekly) and work with averages instead of single measurements.
How many prompts do I need for meaningful KPIs?
At least 30 per topic cluster, better 50–100 across all relevant clusters. Too few prompts make Share of Model statistically unstable.
Does my website structure influence whether ChatGPT cites me?
Yes. Clear headings, self-contained paragraphs, FAQ sections, Schema.org markup, and crawlability for AI bots increase the likelihood of being cited as a source — that is exactly what citable content is about.
Conclusion
AI visibility is measurable — with a defined prompt set, the KPIs Share of Model, Citation Rate, and sentiment, plus weekly repetition. Start with a baseline, pick a suitable tool or build a GDPR-compliant setup of your own, and judge agencies by their measurement methodology instead of their promises. Anyone who wants to be recommended in ChatGPT and Perplexity in 2026 must first know where they stand today.