You marked a page as unavailable in Search Console, expecting it to vanish from Google’s results. Then you see your content cited in an AI Overview, and the assumption crumbles. Turning off visibility in classic search does not automatically remove a page from generative answers. This disconnect is critical for understanding how indexing controls behave in the new AI landscape. We need to look past the familiar blue links to see where your content really ends up.
What Google Extended actually covers
When we talk about Google Extended, we are not referring to a new type of search result, but rather the collective suite of publisher tools that dictate how your content is treated across Google’s entire ecosystem. This term often causes confusion because it sounds like a feature, when it is actually a category of control. Specifically, it encompasses the directives you manage within Search Console and the on-page signals embedded in your HTML.
The scope of publisher controls
To understand the baseline, we need to look at two distinct layers. The first layer lives in Search Console, where you can request URL removals or mark a site as unavailable. The second layer resides on the page itself, primarily through the noindex meta tag or rules defined in robots.txt. These indexing controls serve a singular purpose: they instruct Google’s crawling and ranking systems on how to handle a specific page within the traditional blue-link environment.
For years, this has been the final word on visibility. If you block a page, it disappears from standard search results. However, the landscape has shifted. The llms.txt file now enters the conversation as a newer, specific instruction set designed for Large Language Models. While noindex tells search engines to ignore a page for ranking, llms.txt acts as a roadmap for AI agents, guiding them on how to understand your site’s structure and content priorities.
It is crucial to recognize that these two systems operate with different levels of enforcement. Standard indexing directives are hard blocks backed by Google’s entire search infrastructure. The llms.txt protocol, by contrast, is more suggestion-based. It signals intent to AI tools, but its adherence is not as mechanically absolute as the traditional indexing rules that have governed web visibility for decades. This distinction is where the real complexity lies when trying to manage your brand’s presence in an era where search and synthesis are beginning to diverge.
How AI Overviews select sources
The logic behind source selection in AI Overviews differs fundamentally from traditional ranking. Instead of prioritizing the pages that already hold the top positions in the blue-link results, the system is engineered to direct traffic toward a broader diversity of websites. Data from Google I/O 2024 highlights this shift, noting that users visit a greater diversity of sites for complex questions when these generative answers are present. This means the algorithm weighs topical authority and the uniqueness of data heavily, often selecting sources that might not rank in the top ten for a given keyword.
The incentive of higher engagement
One specific factor driving this diverse selection is the “higher click-through” effect observed in real-world usage. Links embedded within an AI Overview frequently receive more clicks than the same page would if it appeared as a standard organic listing. This higher engagement rate incentivizes the system to maintain a wide range of sources to sustain traffic flow. If the algorithm were to only cite the same few high-ranking domains, it would likely see a drop in overall user engagement. Therefore, the inclusion of niche or specialized sites that offer unique perspectives is not just a byproduct of the technology, but a strategic component to maximize the utility of the summary for the user.
Beyond the ten blue links
The mechanism also contrasts with the static nature of the traditional search results page. In the classic model, position is the primary determinant of visibility. In the generative model, there is a distinct “research” phase where the AI reads and synthesizes content from multiple pages before generating the answer. During this process, a page that ranks lower in traditional results can still serve as a key source for the Overview if its content provides specific, authoritative, or unique data points that other high-ranking pages lack. This shifts the value of content from its rank to its depth and specificity, meaning that indexing controls must be considered in the context of this new, dynamic retrieval process rather than just the static list of search results.
Does noindexing stop AI citations?
The impact of a noindex directive on AI Overviews is more nuanced than on traditional search. While a noindex tag or a “site unavailable” setting in Search Console reliably prevents a page from appearing in standard Google Search results, it does not automatically erase a page from the data streams that feed generative answers. The key lies in understanding the distinction between crawling and indexing. When you apply a noindex tag, you are telling Google’s indexing systems not to include the page in the search index used for display. However, Google’s AI models may still crawl and ingest data from that page for other purposes, such as real-time research to synthesize answers.
This distinction is particularly relevant because AI Overviews rely on a different selection mechanism than traditional search. As discussed, these overviews prioritize topical authority and data diversity over high-ranking status. A page that ranks low—or is completely noindexed—can still be a key source if its content is unique or authoritative for the AI’s research phase. Consequently, simply removing a page from the search index does not guarantee it will be excluded from generative answers.
The current state of enforcement is still evolving. AI Overviews are a relatively new surface, and the way these controls are applied across different Google properties is not yet fully standardized. However, the general consensus among webmasters and experts is that noindexing remains a strong signal. It should, in principle, reduce visibility across Google’s ecosystem, including generated answers. Yet, because the AI models are continuously trained on vast amounts of data, there is no guarantee of immediate or complete exclusion from all AI surfaces. For now, consider noindexing as a request rather than a hard technical block when it comes to AI citations.
The llms.txt signal vs. search directives
The llms.txt file is a plain-text document placed at a site’s root that acts as a roadmap for large language models. It outlines a website’s structure, highlights key content priorities, and suggests which pages are most valuable for an AI agent to process. Unlike traditional metadata, this file is designed specifically to help LLMs navigate complex sites without getting lost in navigation menus or footer links.
However, there is a significant difference in enforcement power between this protocol and standard indexing controls. A noindex directive carries the weight of Google’s entire search infrastructure; it is a hard technical block that prevents a page from appearing in search results. In contrast, llms.txt is a suggestion-based protocol. While many modern AI tools respect these instructions, it is not a definitive barrier in the same way robots.txt or meta tags function for traditional search engines. Treating it as a primary exclusion method can lead to false security.
A practical approach
For best results, treat these two mechanisms as complementary rather than interchangeable. A brand might use llms.txt to guide AI on which products are currently active, helping generative answers stay fresh and relevant. Meanwhile, strictly sensitive or outdated pages should be managed through Search Console controls and noindex tags to ensure they are excluded from Google’s broader ecosystem. This dual-layer strategy addresses both the nuance of AI interpretation and the rigidity of search engine indexing.
Indexing controls remain the primary mechanism for managing visibility within the traditional blue-link ecosystem, acting as the baseline for how Google processes and displays web content. However, AI Overviews function as a distinct layer that prioritizes data depth and source diversity over simple ranking position. This separation creates a scenario where a page might be suppressed in standard search results yet still serve as a valuable citation in a generated answer. As these two surfaces evolve, monitoring them as a single entity becomes less effective. The synchronization between organic rankings and generative answer visibility is no longer perfect, necessitating a separate watch on how your content is interpreted and cited within AI-specific contexts.
