By Richard George
•
August 20, 2026
I sat in a client review earlier this year where two AI visibility dashboards, both paid for by the same organisation, reported the brand's share of voice in ChatGPT eleven points apart. Nobody in the room could explain why. The head of digital asked which number she should put in the board pack. I told her neither, and that the honest answer was a range with a confidence note attached. That went down about as well as you would expect. Thirty years in this industry and I have never seen a measurement category grow this fast on foundations this thin. The IAB now counts more than twenty companies selling AI visibility measurement, each with its own methodology, each capable of producing a different answer for the same brand. Only sixteen percent of brands systematically track AI visibility at all. Meanwhile Fractl data reported by Digiday puts roughly twenty four percent of search and content budgets into AI visibility work. We are spending like the measurement is solved. It is not. What still holds up, for the forceable Let me set the table before I clear it. Branded search volume is now one of the most commercially useful signals a brand has, because it is where AI influence surfaces. Similarweb work and our own at WPP Media found that roughly fifty six percent of AI influenced traffic arrives as branded search rather than an AI referral click, days after the conversation happened. Impression share in Google Search Console still matters, arguably more than clicks do. Paid search conversion and quality score still matter. Crawlability, structured data and site speed still matter, because a page a model cannot read is a page it cannot cite. What has stopped working is the click as the primary unit of account. Ahrefs measured a fifty eight percent reduction in position one click through rate when an AI Overview appears, and Pew's browsing data found users clicked eight percent of the time with an AI Overview present against fifteen percent without. If clicks are still your headline KPI, and Kantar research suggests they are for around seventy eight percent of brands, you are measuring the shadow rather than the object. What actually changed, and why it matters now Three things happened in quick succession this year that should reset how CMOs think about this. First, Google gave us a real, if partial, measurement surface. On 3 June 2026 Google launched dedicated Search Generative AI performance reports in Search Console , rolling out to UK site owners first under CMA pressure. For the first time you can separate visibility inside AI Overviews and AI Mode from ordinary organic. The catch is significant: impressions only. No clicks, no click through rate, no queries. It answers "did I appear" and refuses to answer "what was it worth". Second, the scale question stopped being theoretical. Semrush analysed 126 million US AI search prompts between January and April 2026 across ChatGPT, Gemini, AI Mode and AI Overviews, covering 1,200 brands in 22 industries. Only 36 of those brands held a top 100 position on every platform in every month. The same study found ChatGPT citing an average of fifteen sources per response against Gemini's three. Per platform measurement is not a nice to have. A blended visibility score across engines that behave that differently is an average of incompatible things. Third, and this is the one that should worry vendors, Rand Fishkin ran 2,961 controlled tests across ChatGPT, Claude and Google AI with 600 volunteers . Ask an AI for brand recommendations a hundred times and you get the same list fewer than one time in a hundred, and the same list in the same order fewer than one time in a thousand. His conclusion, which I share, is that visibility percentage across many prompts run many times is defensible, and any tool selling you a ranking position in AI is selling you nothing. The frameworks I would build on, and where I would push back The IAB's Measuring Visibility in the AI Era is the most useful thing published on this in eighteen months. Its four P's give us a causal hierarchy rather than a metrics soup: Presence, whether you appear at all Prominence, where and how visibly Portrayal, in what context and with what accuracy Persuasion, whether any of it moves anything Its real contribution is the split between directional and decision grade data, and a floor that should embarrass a lot of vendor decks: fewer than fifty queries in a measurement programme is classed as exploratory, not even directional. Two honest caveats. Caroline Giegerich told AdExchanger the IAB deliberately avoided calling this a standard, because standards need stability and the market has none. And the framework measures single query, single response visibility. That is not how people use these tools. Real decisions unfold over four, six, ten turns of a conversation, and a brand can be named confidently at turn one and gone by turn five. That gap is why I use the PLUS framework with clients alongside the IAB vocabulary. Four factors the IAB does not fully price in. Personalisation: responses adapt to account history, so universal rankings are a fiction and you should audit by persona instead. Location: outputs are hyper localised, so measure local citation accuracy across your top territories. Uniqueness: models are probabilistic, so batch test the same prompt fifty times and report mention probability, not a snapshot. Stateful ness: context carries, so track brand decay rate across a thread. I want a brand holding above fifty percent mention visibility at prompt four and above forty percent by prompt six. That number tells me more about commercial resilience than any single answer share ever will. Traditional metrics and their AI era equivalents