Why 80% of AI Training Data Is English

Published on August 19, 2026

Roughly 80% of the training data behind large language models is in English, according to a 2024 Common Crawl analysis. This statistic shifts the conversation about multilingual AI search from algorithmic design to a fundamental infrastructure gap. When a user queries an engine like Perplexity AI in Korean or Chinese, the system faces a fragmented digital landscape. It must navigate a web where local search engines such as Baidu and Naver sit behind distinct operational barriers that Google does not. Source selection in this context is less about what the model can understand and more about what it can actually reach. The result is a bias toward globally accessible, English-centric sources, leaving high-value local data outside the model’s primary citation radius.

The Baidu and Naver Barrier: Why local engines fragment the web

Google holds the majority share in many regions, but the web is not uniform. In specific markets, local engines dominate the user experience and the underlying data infrastructure. For multilingual AI search to function effectively, it must navigate these distinct ecosystems. However, these local systems often operate behind significant operational barriers that global AI engines like Perplexity AI struggle to penetrate directly.

The core issue is access. To obtain high-quality, proprietary data from the dominant local search engine, developers and marketers often face strict entry requirements. Tools such as Baidu Tuiguang in China, Naver SearchAd in South Korea, and Yandex Wordstat in Russia are not open to everyone. Accessing these platforms typically requires a registered local business entity. Furthermore, network access is often restricted to local IP addresses, creating a digital border that prevents direct, real-time data retrieval from outside the region.

This fragmentation creates a coverage gap for AI systems. Because registering with these local platforms can take weeks, AI engines cannot easily establish a direct, unified feed from the dominant local search provider. As a result, when a user queries an AI about a local topic, the system often lacks access to the most authoritative, up-to-date local source. Instead, it must rely on publicly available web crawls or global indexing, which may miss critical nuances or the most current information available on the local platform.

Market Dominant Engine Key Access Requirement
China Baidu Local business entity and local network access
South Korea Naver Local business entity and local network access
Russia Yandex Local business entity and local network access

This structural barrier means that non-English sources from these regions are less accessible to AI training and retrieval pipelines, impacting the depth and accuracy of answers in those specific languages.

How Perplexity AI navigates the non-English source landscape

When users ask how Perplexity AI selects sources, the answer is less about algorithmic preference and more about operational reach. The engine is designed to prioritize high-authority, accessible, and structured content that can be reliably retrieved and verified. In English-speaking markets, this usually means pulling from well-indexed sites like major news outlets, academic repositories, or government databases that are open to public crawls.

The situation shifts significantly when we look at non-English sources in fragmented digital landscapes. In regions where local search engines like Baidu or Naver dominate, the engine cannot access their proprietary, entity-gated data. Because these local platforms require specific local business entities and network access to use their advanced tools, they do not provide a unified, direct feed to global AI systems. As a result, Perplexity and similar engines are forced to rely on global indexing—often dominated by Google’s infrastructure—or public web crawls that may not reflect the depth of local knowledge.

This creates a distinct coverage vs. accuracy trade-off. A query in Japanese might successfully pull from a well-maintained global index, where the content is accessible and structured. However, a query in Korean might struggle if the most relevant, up-to-the-minute local news is locked behind Naver’s proprietary ecosystem or requires entity verification to access. The engine is left choosing between a potentially outdated but accessible global source and a locally accurate but inaccessible one.

Ultimately, the “choice” of a source is often the choice of the most accessible version of a fact, not the most authoritative local version. For multilingual AI search, this means that visibility in AI answers is heavily dependent on whether your content is part of the public web’s open structure. If a source is hidden behind a local firewall or requires specific regional access, it effectively remains invisible to the AI, regardless of its local reputation or authority.

The structural gap in multilingual AI language processing

The core issue in AI language processing is not just grammar or vocabulary, but the underlying data distribution. LLM training corpora are roughly 80% English by token count, a fact rooted in the 2024 Common Crawl analysis. This creates a significant language discount for non-English sources, meaning the model has inherently more experience and nuance with English text than any other language.

This imbalance directly impacts citation density. When a user asks a question in Spanish or German, the AI has fewer high-quality references to draw upon compared to an English query. The result is a thinner net of available sources, which forces the system to rely on whatever is accessible rather than what is most authoritative. This gap is a structural feature of how these models learn, not a temporary oversight.

Another critical factor is nuance loss during on-the-fly translation. When an AI translates content from a local language into its working mode, it can strip away cultural context and specific phrasing that gives the original source its value. Consequently, a native-language source with clean structure often becomes more valuable to the AI than a translated English equivalent, because it preserves the semantic integrity the model needs to verify accuracy.

Ultimately, the challenge for multilingual AI search is not just linguistic translation, but the infrastructure reality of a fragmented web. The non-Google web is harder to index consistently, and this fragmentation means that even when the AI understands the language, it may not be able to reach the best sources. This creates a persistent gap between what the model can comprehend and what it can actually access and cite.

Frequently asked questions on non-English AI search

Why is my brand invisible in non-English AI search?

It is common for business leaders to ask why their local website, which ranks well on regional platforms, does not appear in AI-generated answers. The core issue is access. AI engines often lack direct access to proprietary data from local search giants. Consequently, they rely on global training data that is heavily biased toward English. If your content is not part of that accessible global index, the system simply cannot see it, regardless of its local relevance.

Does Perplexity AI use Baidu or Naver for sourcing?

Generally, no. Perplexity AI does not use local dominant engines like Baidu or Naver as primary, direct sources. This is due to the technical and legal barriers, such as the need for local business entities and specific network access, which prevent direct integration. Instead, the system pulls from the public web, which means the “best” answer is often the most accessible one, not necessarily the most authoritative local source.

How can I make my non-English content more citable?

To increase the likelihood of being cited by AI, you need to provide a clear citation hook. This means using explicit answer formats that directly address common queries. You should also implement local structured data, specifically schema markup with the inLanguage property, to help the engine understand the linguistic context. Positioning your brand as an authoritative source within the public web ensures that the AI has a reliable, easy-to-extract option when it generates a response.

The path forward for non-English content

As AI search becomes a primary referral source, the gap between English and non-English content quality will likely widen rather than narrow. The infrastructure underlying multilingual AI search is currently imbalanced, with 80% of LLM training data in English, which creates a persistent structural disadvantage for non-English sources. Brands cannot force engines to prioritize local languages, but they can reduce friction for the crawler.

The practical path forward is to focus on the infrastructure of non-English content. This means prioritizing clean structure, explicit schema markup with correct inLanguage tags, and clear, direct answers. By making content the easiest source to extract and verify within a fragmented landscape, brands increase their likelihood of being cited by systems like Perplexity AI. The goal is not to compete with the English data bias, but to become the most reliable anchor in the local digital environment.

AEO/GEO

Want to learn more?

Contact us for direct consultation and support.

Contact us

Related Articles

Audit Risk in Localized AI Answers: Currency Rate Errors
Multilingual & international aeo

Audit Risk in Localized AI Answers: Currency Rate Errors

A Berlin-based customer receives a quote from your AI assistant. The display shows a clean Euro price, but the backend log records the transaction at a...

Read article
Why AI local pricing leaks revenue via silent currency conversion
Multilingual & international aeo

Why AI local pricing leaks revenue via silent currency conversion

An AI assistant quotes a customer in Tokyo a final cost of 15,000 yen. The transaction processes smoothly, the customer is satisfied, and the system logs a...

Read article
19 LLMs, 3,991 Figures: Why Regional Media Dominates AI
Multilingual & international aeo

19 LLMs, 3,991 Figures: Why Regional Media Dominates AI

A 2026 study published in npj Artificial Intelligence reveals that the political leanings of large language models align closely with the geopolitical...

Read article
One Prompt, Six Markets: Where Buyer Personas Break Down
Multilingual & international aeo

One Prompt, Six Markets: Where Buyer Personas Break Down

Most teams build a detailed buyer persona prompt once, run it through their AI stack, and then reuse it for every global campaign. The result is content...

Read article
Why Korea's 'Core Market' Status Exposes Flaws in Buyer Prompts
Multilingual & international aeo

Why Korea's 'Core Market' Status Exposes Flaws in Buyer Prompts

Between 2020 and 2024, iHerb did not simply translate its interface into Korean to grow in Asia. It re-engineered its buyer prompts to respect South Korea’s...

Read article
3 reasons machine-translated pages lose AI citations
Multilingual & international aeo

3 reasons machine-translated pages lose AI citations

Many global teams assume the language barrier is the only hurdle to international visibility. They translate their best-performing English content, publish...

Read article