Ask ChatGPT Which human languages do you work best in? and it will output a list that should surprise no one: English takes the lead, followed by Spanish, Portuguese, French, and German all in the top five.

Now ask Gemini which human languages it can work in, and it returns a tiered list, with, unsurprisingly, the same languages at the top: English, Spanish, French, German, and Portuguese. Interestingly, Gemini included a 'Note on Fluency' at the end of its output, warning that its highly precise, nuanced, and speedy responses were limited to major global languages. As for lower-resource languages, one could expect grammar and idiomatic expressions to be, occasionally, less precise.

Low-resource languages are human languages with limited amounts of computer-readable data available. This can happen in a variety of ways: there can be few speakers of that specific language, or there can be a considerable amount of people who speak the language but not much digitized data in it. In other cases, as researchers at Stanford point out, a language may have plenty of digital data and speakers, but lack the resources or public awareness needed to actually have models implemented in it. It is also worth mentioning that roughly half the world's languages are solely transmitted orally, meaning that they have no native or traditional writing system, much less suitable digitized data.
According to Ethnologue, there are 7,170 languages in use today worldwide. If we consider only languages with established writing systems (which is problematic in and of itself), we are left with about 3,800 human languages. Only about 20 languages have enough training data online to create natural language processing systems. And even within those twenty, differences exist, as shown via the exchanges with ChatGPT and Gemini.
As Statista showed, in October 2025 English dominated online content, being used by nearly half of all websites worldwide. Spanish ranked second, accounting for around 6 percent of global web content, followed closely by German at 5.9. Not only is the fall-off here tremendous, notice also that low-resource languages do not even get a share of the web-content pie, despite the fact that speakers of low-resource languages amount to around 1.2 billion people worldwide.

When people talk about AI's data problem, they are usually referring to the fact that the data used to train these models is low-quality, unethically sourced, or simply insufficient. But there is, arguably, a more serious and insidious problem when it comes to the training data used by LLMs: a language problem.
LLMs perform substantially better in high-resource languages (particularly English) compared to low-resource languages, due to severe imbalances in training data. As Peppin et al. pointed out in a recent paper, "This 'language gap' has far-reaching implications which ultimately leave certain language communities around the globe marginalized. AI models present both limited language support and biases are introduced that reflect Western-centric viewpoints, undermining other cultural perspectives." A good example of this is an LLM advising a user in a rural, non-Western region on managing a disease. It might suggest solutions that rely on Western healthcare systems, specific insurance models, or non-existent infrastructure, completely failing to adapt to local cultural practices or actual available resources.
Leaving underrepresented languages and communities behind doesn't just mean they don't get to use the newest AI tools; businesses and workers in developing regions are totally excluded from AI-driven productivity gains. That means that where others can take advantage of automated customer service tools, guided coding, and rapid data analysis, these communities fall behind in the fast-evolving digital world. Furthermore, "in regions where universal health care remains a challenge, AI-powered diagnostic tools that only function in English create a new layer of health care inequality."
Not only are high-quality outputs limited to high-resource languages, but so are effective harm and bias detection. Yong et al. showed in their 2024 publication that "translating English inputs into low-resource languages increases the chance to bypass GPT-4’s safety filter from <1% to 79%." By translating unsafe inputs into low-resource languages such as Zulu or Scottish Gaelic, they were able to evade the model’s safety measures and elicit harmful responses almost half of the time. Combining different low-resource languages increased the jailbreaking success rate to around 79%. Compare this to the original English inputs, which had less than a 1% success rate in generating harmful outputs.
Though the pace of technological improvements is remarkable, this remains an issue. A 2026 research paper by Marx et al. showed that "Simply translating harmful prompts into low-resource languages no longer effectively bypasses LLM safety guardrails. However, multi-turn conversations that distribute harm intent across conversation turns remain effective." They concluded that poor automated translation quality instead of, as one would hope, stronger safety guardrails, is responsible for the lower jailbreak rates (relative to English) in these low-resource languages.
Because of the very nature of the problem, many of the apparent solutions are also restricted to high-resource languages: for instance, where these benefit from synthetic data and high technical expertise, low-resource languages don't. Peppin et al. said it well: "Data availability is one of the most potent levers of progress." They posit that a variety of sources of data can be beneficial for improving multilingual coverage, while simultaneously acknowledging that one of the most formidable challenges in this field is the quality and quantity of data available. Interestingly, where there seems to be consensus in the field with regards to translated data propagating errors, lacking nuance, and generally being insufficiently high-quality, they found that "it is better to increase coverage of data by including both human, synthetic and translated data rather than solely prioritizing human annotations."
Not only does language low-resourcedness go beyond mere data availability and reflect systemic issues in society, it also affects who engages with technology, and how. How we engage with technology shapes the way we think about problems and, in turn, the way we think about culture and the world. When it comes to the language problem present in LLMs, we shouldn't just worry about cultural erasure, but about cultural homogeneity. Ensuring a significant portion of the world's languages is well-represented in the technology that exists will allow everyone (not just speakers of high- or low-resource languages) to live in a richer, better world.
Sources




