Issue 04 8 August 2026

What changed in AI search this fortnight

Issue four. Every two weeks I read what changed across ChatGPT, Perplexity, Google AI Overviews, and the models behind them, then turn it into what it means for brands that want to show up when a customer asks an AI for a recommendation. Last issue was about measurement honesty, that a one-shot visibility score is noise. This fortnight the research answered the next question. If the score is noisy, where do the citations actually come from, and who controls the access to earn them. Here is what moved.

TL;DR

Your own website is a small slice of where AI cites you. Muck Rack analyzed 25 million links inside AI answers and found earned media accounts for 84 percent of citations. A separate Trustpilot and Seer study found brands with active review profiles cited roughly 75 times more often than brands without one. ChatGPT leaned on community platforms and pulled about 15 sources per answer, Gemini about 3. Treat the exact figures as single studies on single datasets and measure your own category before betting a budget on them. The direction is the strongest evidence yet: the work that earns citations happens mostly off your own domain.

The access to earn those citations is being repriced. The EU fined Google 890 million euros under the Digital Markets Act, large publishers including Reuters, Politico and The Economist are openly weighing whether to block Google's crawler, and Cloudflare shipped a Monetization Gateway that lets a site charge any caller, including AI agents, for any page, dataset, or API. A page an agent cannot reach, or cannot afford to fetch, cannot be cited. Crawl access is turning into a priced, permissioned asset.

A correction worth acting on this week. Google stopped showing FAQ rich results on 7 May 2026 and removed the feature documentation on 15 June. Missing FAQPage schema is no longer a search-readiness defect, and Google says no special schema is needed for AI Overviews or AI Mode. If a proposal promises you Google FAQ rich results from markup, it is selling something that no longer exists. Keep FAQ markup only where it truly describes visible questions and serves a clear non-Google use.

One cheap audit item with real teeth. In a robots.txt file the most specific user-agent section wins completely. If you add a section for GPTBot, ClaudeBot, or PerplexityBot, your shared disallow rules stop applying to that bot unless you copy them into every section. A real store got its spam URLs indexed exactly this way. Duplicate the rules per bot, or the block silently does nothing.

Google punishes thin content, not AI content. Ahrefs looked at a million top-ranking pages. Fully AI-written pages can still rank, but a higher AI share lined up with lower indexation and lower positions. The read is that heavy, unedited AI output tends to be low quality, and low quality is what gets hit. Quality is the gate, and that has not changed.

And the paid path into AI answers matured. ChatGPT Ads added conversion bidding, geo exclusions, budget pacing and a bulk API, and OpenAI is testing chatbot-native ads that open a conversation with a business agent instead of sending a click to your site. The organic and paid line inside AI search is now a live commercial question rather than a future one.

Vendor moves

OpenAI had the busiest fortnight. ChatGPT Ads grew into a real channel: conversion-optimized cost-per-click campaigns, average daily budgets, automatic budget pacing, geographic exclusions, mobile measurement through AppsFlyer and Adjust, a bulk API, and refreshed product feed ads with pricing and star ratings. On top of that, OpenAI cut GPT-5.6 Luna pricing by about 80 percent and Terra by 20 percent, added a Fast mode, shipped an official Terraform provider and GPT Transcribe, and, per Search Engine Land, is testing chatbot-native ads that start a conversation with an AI business agent instead of linking out to a site. The pattern to hold on to: paid presence inside AI answers is becoming a normal channel with normal campaign tooling, and the organic versus paid boundary in AI search is now a live commercial question.

Google's fortnight was defined by pressure from outside. The European Commission issued its first Digital Markets Act fine against Google, 890 million euros, for favoring its own services in Search. In the same window the Wall Street Journal reported that large publishers, including USA Today, Politico, Reuters and The Economist, are seriously weighing whether to block Google's crawler. Set against that, Alphabet reported Q2 Search revenue of 63.27 billion dollars, up 17 percent, while its reassurances about clicks stayed unverifiable. The framing worth keeping is Search Engine Journal's: Pichai says queries are at an all-time high, Nick Fox repeats billions of clicks weekly from AI features, and none of the claims comes with a definition, a breakdown, or a published method.

On product, Google shipped the Gemini 3.6 Flash family, updated its review-snippet guidance to exclude fake or undisclosed incentivized reviews from markup, made Search Console platform properties global so brands can track how their social and video posts perform in Search, Discover and News, and started rolling out an opt-out control for AI search features as Top Stories begin appearing inside AI Overviews. Google also said that spammed on-site search-box pages may be treated as a site quality problem.

Cloudflare ran an Agents Week built around an agent-native web. The headline item for this work is the Monetization Gateway, now on a waitlist, which lets a site charge any caller for any resource behind Cloudflare, a page, a dataset, an API, or an MCP tool, settled in stablecoins over the x402 payment protocol. Alongside it, Cloudflare shipped finer bot controls that split Search, Agent and Training crawlers so an owner can allow one and refuse another, plus an Attribution Business Insights dashboard that reports crawler appetite. Its PACT protocol, agreed with three rival browsers, is not live yet per Search Engine Journal, and even once live it solves only half of the agentic access problem. This is last year's Content Independence Day with a payment rail attached.

A quieter legal item props the whole tooling layer up. A federal court dismissed Google's DMCA lawsuit against SerpApi, ruling that bypassing anti-bot measures to scrape public search results does not violate copyright, because a plain results page is not a copyrighted work. The legal ground under scraping-based AI visibility monitoring is firmer than it was a month ago, while the technical ground, behavioral bot detection at the edge, keeps getting harder. The fight moved from the courtroom to the edge. Perplexity, separately, added role-based access controls, an API credential vault, and Opus 5, and Anthropic listed Claude Opus 5 as its next flagship tier.

How LLMs read the web

The structural finding this fortnight is where citations come from. Muck Rack's analysis of 25 million links inside AI answers put earned media at 84 percent, and a separate Trustpilot and Seer study found brands with active review profiles cited roughly 75 times more often than brands without. Which sources get cited depends heavily on your category, so this is a prior to verify per brand, not a universal fact. For a car buyer's question the engines tend to cite ADAC, AutoScout24 and carwow rather than the manufacturer's own site. For your category it will be a different set of sources. The point that carries across categories is that the engines cite third-party and community sources by design, and your own domain is a minority of the mix. Measure which domains get cited for your questions first, then decide where to spend.

The FAQ correction closes a gap between content and markup. FAQPage remains a Schema.org vocabulary type, but Google's FAQ rich-result feature is gone as of 7 May 2026, with the documentation removed on 15 June. A page can publish useful visible questions and answers without FAQ schema, and Google's own AI-search guidance says there are no special Schema.org types needed for AI Overviews or AI Mode. If you keep FAQ markup, Google's general structured-data policy still applies: the markup must be a true representation of content a reader can see. QAPage is not a substitute, Google reserves it for a single-question page where users submit answers.

Crawl access is the story to act on. Read the Monetization Gateway and the split Search, Agent and Training bot controls together, and the message to a brand is direct: a page an agent cannot reach, or cannot afford to fetch, cannot be cited. For any site that serves a JavaScript-rendered shell to crawlers, this compounds an existing gap. First the page has to render server-side so there is real text to read, then it has to stay reachable. Crawler access and rendering belong together as one standing check, not two separate ones.

John Mueller flagged a robots.txt mechanic that will bite exactly the sites trying to manage AI crawlers. The most specific user-agent section wins, and it wins completely. A site with a Googlebot section and a general section applies only the Googlebot section for Googlebot, and the general disallows silently stop existing for it. Anyone adding per-bot sections for GPTBot, ClaudeBot or PerplexityBot has to duplicate the shared rules into every section, or the rules quietly do nothing for those bots.

An independent audit put numbers on the readiness gap. Search Engine Land measured 71 businesses and found the average one resolved only 15.6 percent of what AI systems need to trust and retrieve it, with 17 percent showing no AI-retrievable presence at all. The named failure modes are the familiar ones: trust signals buried on secondary pages, client-side JavaScript sites that return no extractable text, and dead or fragmented domains. It reads like an independent replication, at small-business scale, of the rendering problems we see on large consumer sites.

On llms.txt, nothing changed, and that is the finding. SE Ranking analyzed 300,000 domains and found no statistical link between having the file and being cited, with only one of the 50 most-cited domains carrying it. Google says Search and its AI surfaces ignore it. It is cheap, harmless hygiene, not a visibility lever, and it should never be sold as one. Worth noting the other way: OpenAI now publishes an llms.txt with markdown versions of its own doc pages, a signal about how a vendor wants its own content read even while Google shrugs.

Ahrefs added hard numbers to the AI-content question. Across a million top-ranking pages, indexation fell from 49.3 percent for pages with the least AI text to 40.4 percent for pages that are more than 80 percent AI, and pages under half AI content held 82.2 percent of top-three rankings. Fully AI-written pages can still rank. The read is that Google targets low quality, and heavy AI use tends to correlate with lower quality, not that AI authorship is penalized on its own.

Community signal

The counterweight to every scale-your-GEO-content pitch came from Lily Ray. Her research found roughly 40 companies that scaled self-promotional listicles and saw 30 to 50 percent organic visibility drops within weeks, domain-wide, even when the content sat in a single subfolder. Her team's GEO recovery work is now officially starting. Scaled, self-promoting pages are a documented domain-wide risk in both classic search and AI surfaces. Original data and genuine third-party validation are the durable assets.

Several accounts spent the fortnight circulating citation-share statistics, and they are worth learning to recognize. One post claims Reddit is the most cited source at roughly 40 percent across engines. Another repeats Perplexity citing sources about 97 percent of the time against ChatGPT at 16 percent. Simon Leung's post claims 85 percent of AI citations come from third-party sources and only 16 percent of brands track their share of a model. Aleyda Solis summarized a study putting Gemini's same-language citation share in non-English markets near 77 percent against roughly 52 percent for ChatGPT. No primary study backs these exact figures, they trace to single social posts, so keep them out of your slides. The direction matches what we see in client scans. The size of the numbers does not survive checking. If a percentage like this is about to move a budget, measure it for your own category first.

Among the people worth following: Aleyda Solis published an episode on setting AI search strategy and goals, Glenn Gabe reported Top Stories and Preferred Sources now surfacing inside AI Overviews, and Mike King and Barry Schwartz both pushed Similarweb's number that AI Overviews now appear on close to half of US searches. Treat the reach figure with care: Similarweb puts it at 43 percent, other trackers put it far lower, and Similarweb's own methodology has been questioned. The direction, that AI Overviews are now common on informational queries, is not in doubt.

What it means for your brand

Fund the off-site work. If earned media is where most citations in those studies came from, third-party presence belongs in the budget next to your on-site fixes. Reviews, mentions, and being genuinely useful on the sources your category's engines actually cite. Start by measuring which domains get cited for your own questions, then go where they are rather than guessing.

Treat crawler access and rendering as one job. Confirm AI crawlers can both reach and render your important pages. If you sit behind Cloudflare or a similar provider heading toward paid access, know which bot tracks you allow. If your pages are JavaScript-rendered, an AI crawler may see an empty shell, so server-side rendering comes before any citation work pays off.

Stop treating missing FAQ schema as a gap. Do not add FAQPage to chase Google rich results, that feature is gone. Write visible, question-led sections only where research shows real demand, and treat schema as an optional, justified second step. If anyone has sold you FAQ markup as an AI Overviews lever, that recommendation needs correcting.

Fix the robots.txt trap in ten minutes. If you have per-bot sections for AI crawlers, confirm each one repeats your shared disallow rules. If it does not, the block is silently off for those bots. Cheap to check, embarrassing to miss.

Thin content is the real risk, and AI-assisted content on its own is safe. The Ahrefs data backs the line we already give: edit heavily, keep quality high, and AI drafting is fine. Shipping raw model output at scale is what drags rankings down, in both classic search and AI surfaces.

Watch the paid path, do not buy it yet. ChatGPT ads and chatbot-native ads are maturing into a real route into AI answers. For now, earning your way in on trusted third-party sources is cheaper and holds up longer. Keep the paid route on the watchlist for the day earning it fast enough stops being realistic.

Get the next issue in your inbox, every two weeks.

Get the report →

Sources

The references behind this issue.