AI visibility measures how often and how prominently your brand appears in the answers of generative engines (ChatGPT, Perplexity, Google AI Overviews, Gemini). It is measured by querying a fixed prompt set repeatedly and evaluating how often your company is mentioned (citation share), linked and described positively. For real estate brokers, law firms and agencies, this measurement is increasingly replacing pure ranking tracking, because a growing share of searches never lands on a classic SERP.
Why AI visibility becomes a must-have KPI in 2026
Generative answer engines answer more and more questions directly, without a click to a website. For professional service providers, this means: when a prospect asks ChatGPT "Which law firm in Frankfurt specialises in IT law?" or Perplexity "Best real estate broker for commercial properties in Munich", the AI's answer decides who gets the first contact — not position 3 on Google.
Classic SEO tracking (position, impressions, CTR) does not capture this. You need a dedicated measurement discipline that answers three questions:
- Are we mentioned at all? (Presence)
- How often compared to the competition? (Citation Share)
- How are we described — and is it accurate? (Sentiment & factual accuracy)
Step 1: Build a robust prompt set
The foundation of every measurement is a reproducible set of prompts that reflects real customer questions. Rule of thumb: 30–60 prompts per company, split into four categories.
Category A – Category prompts (without brand name)
Here you check whether the AI suggests you on its own.
- "Which real estate brokers in Frankfurt specialise in luxury properties?"
- "Recommend a marketing agency in the DACH region for Google Ads."
- "Which law firm advises startups on GDPR and data protection?"
Category B – Problem/use-case prompts
- "How do I find a broker who also supports the financing?"
- "Who can help me build a CRM automation with Make or n8n?"
Category C – Brand prompts (with your name)
Here you check factual accuracy and sentiment.
- "What does [your company] do?"
- "Is [your company] reputable / which industries is it suited for?"
Category D – Comparison prompts
- "[Your company] vs. [competitor] — who is better for commercial real estate?"
Important: Phrase every prompt the way a real customer would type it (including location and niche). Freeze the wording — only a stable set delivers comparable time series.
Step 2: Define the right KPIs
Measure per prompt and per engine. These six KPIs are the standard we track for clients at Mindflows:
- Presence Rate: Share of prompts in which your brand is mentioned at all. Example: mentioned in 18 of 40 category prompts = 45%.
- Citation Share (Share of Voice): Your mentions divided by all brand mentions in the same prompt set. The most meaningful competitive KPI.
- Average Position in Answer: Are you named as the first, third or last option? Positions 1–2 carry significantly higher click value.
- Link Citation Rate: Share of answers that actually link to your domain (especially relevant in Perplexity and AI Overviews).
- Sentiment Score: positive / neutral / negative, per mention. Scale e.g. −1 to +1.
- Factual accuracy: Are services, location and specialisation correct? Hallucinated false statements are an acute reputational risk.
Example target values for a mid-sized DACH law firm
| KPI | Starting value | 6-month target |
|---|---|---|
| Presence Rate (category prompts) | 20% | 50% |
| Citation Share | 8% | 20% |
| Link Citation Rate (Perplexity) | 10% | 35% |
| Factual accuracy | 70% | 95% |
These values are guidelines — calibrate them to the market size and competitive density of your niche.
Step 3: Measure across engines and regions
Every engine works differently, so measure at least four sources separately:
- ChatGPT (with and without web search): GPT partly uses training knowledge, partly live search. The two modes deliver different results.
- Perplexity: strongly source-based, good link citations — ideal for measuring the impact of content.
- Google AI Overviews: closely tied to classic search and Google Business Profile.
- Gemini / Copilot: growing reach in the DACH region.
Also take region and language into account: ask in German and English, and — if possible — from German IP ranges. Local mentions ("broker in Frankfurt") depend heavily on location signals.
Step 4: Tools for measurement (2026)
Three approaches, depending on budget and maturity:
1. Specialised GEO monitoring tools. Providers such as Peec AI, Otterly.ai, Profound, AthenaHQ or Scrunch track prompt sets automatically across multiple engines and deliver citation share dashboards. Entry pricing usually starts at around €100–500 / month. Ideal for agencies measuring for multiple clients.
2. Your own monitoring with automation. For GDPR control and full data sovereignty, we at Mindflows build setups with Make or n8n: a workflow calls the OpenAI, Perplexity and Gemini APIs with your prompt set (e.g. weekly), an LLM step parses the answers for brand mentions and sentiment, and the results land in Airtable/Baserow. A Softr portal turns this into a client dashboard. Costs: primarily API usage (often < €50 / month) plus setup.
3. Manual sample tracking. To get started, a spreadsheet is enough: 30 prompts, queried manually in each engine every month, KPIs entered. Effort approx. 2–3 hours/month — good for getting a feel before you invest.
Important for DACH: Check the data processing when using APIs. Use EU data residency options, conclude data processing agreements (DPAs) and do not enter any personal client data into prompts.
Step 5: From measurement to improvement (GEO)
Measuring is half the battle. These levers demonstrably improve the values:
- Citable content: Create pages that clearly answer a question in 2–3 sentences (just like this guide). LLMs prefer to cite self-explanatory passages.
- Structured data: Organization, LocalBusiness, FAQ and Service schema help engines extract facts correctly.
- Consistent entity signals: Company name, location and services identical across website, Google Business Profile, LinkedIn and industry directories.
- Third-party mentions: Mentions in trade media, directories and reviews increase the models' trust more than self-descriptions.
- Fact correction: For hallucinated false statements, provide a clear, well-structured reference page ("About us", services) that live search can access.
After every measure, run the same prompt set again — that is the only way to prove impact.
Case example: a broker's office in 90 days
A Frankfurt real estate brokerage started with a presence rate of 15% and 0% link citations in Perplexity. Measures: 12 citable FAQ pages on the buying process, financing and neighbourhoods, LocalBusiness schema, a consistent Google Business Profile. After 90 days: presence rate 42%, link citation rate in Perplexity 28%, factual accuracy up from 65% to 90%. Measurement was carried out throughout with a frozen 35-prompt set via an n8n workflow.
Common mistakes
- Changing prompts: destroys comparability. Freeze the set, version any changes.
- Measuring only one engine: ChatGPT results ≠ Perplexity ≠ AI Overviews.
- Ignoring sentiment: Being mentioned is worthless if the description is wrong or negative.
- One-off measurement: AI answers fluctuate; measure regularly and across multiple runs per prompt.
FAQ
How often should I measure AI visibility?
Monthly as a minimum, weekly during active campaigns. Because LLM answers fluctuate, we recommend at least 3 runs per prompt at each measurement point to form an average.
What is the single most important KPI?
Citation share (share of voice) in the category prompts without brand names. It shows whether the AI recommends you on its own over the competition — the best indicator of real generative visibility.
Do I need an expensive tool or is a spreadsheet enough?
To get started, a spreadsheet with 30 prompts is enough. With multiple clients or a weekly frequency, automation (Make/n8n) or a specialised GEO tool pays off, saving time and producing clean time series.
Is this GDPR-compliant?
Yes, if you do not enter personal data into prompts, use EU data residency and conclude data processing agreements with your API providers. When you run it yourself via n8n, you retain full data sovereignty.
How quickly do the values improve after GEO measures?
First effects in live-search engines (Perplexity, AI Overviews) often within 4–8 weeks. Answers based on training knowledge (ChatGPT without web) react more slowly — here, long-term entity and mention signals count.
Conclusion
AI visibility is measurable in 2026 — with a fixed prompt set, clear KPIs (presence rate, citation share, sentiment, factual accuracy) and the right tool chain across ChatGPT, Perplexity and AI Overviews. For brokers, law firms and agencies in the DACH region, this is the foundation for steering GEO investments instead of continuing to rely on the shrinking classic SERP.