Google says its technologies and products now work in more than 300 languages, used by more than 7 billion people. That progress represents nearly 86% of the world’s population, but it also makes one gap clear: thousands of languages and dialects still have little presence online and in digital tools.
The company argues that translating words is not enough. For artificial intelligence to be truly useful, it must also understand tone, pauses, emotions, language mixing and the cultural particularities of each community.
From translating text to understanding how we speak
For years, speech recognition systems followed a fairly rigid chain: turn audio into text, process that text and then turn it back into audio. The problem is that important details of human communication get lost along the way.
Do you always speak in perfect sentences? In real life, we interrupt one another, laugh, hesitate, use local expressions and mix languages. That is what happens with Spanglish or Hinglish, for example.
Google says it is now working with models capable of processing audio directly, without relying solely on a written transcription. According to the company, this strategy makes it possible to better interpret both what is said and how it is said.
The tools mentioned include:
- Gemini 3.5 Live Translate: real-time spoken translation in 70 languages and more than 2,000 language pairs, with recognition of language changes and emotional cues.
- Gemini 3.5 Transcribe: a speech-to-text model designed to work with noisy audio, specialized terms and cleaner formats. It also powers voice typing features in Gboard.
- Universal Speech Model: a model trained on 12 million hours of audio to transfer knowledge from languages with abundant data to others with fewer digital resources.
The most ambitious goal is the so-called 1,000 Languages Initiative, which seeks to create technology for the world’s most widely spoken languages, including those that have been overlooked by traditional systems.
Data must also represent communities
The web is dominated by a small group of languages. That is why training AI to understand less-represented languages requires more than collecting text from the internet: it means working directly with the people who speak them.
Google highlights three open-data partnerships:
- WAXAL: a speech dataset covering 27 languages in sub-Saharan Africa, spoken by more than 100 million people across more than 26 countries. It includes tonal variations and conversational rhythms that often disappear from conventional databases.
- Project Vaani: an initiative developed in India that has collected more than 30,000 hours of speech in 109 languages, with the participation of more than 155,000 people.
- Amplify Initiative: a collaboration with more than 1,600 local experts and 20 universities across four continents to collect multimodal data with cultural context.
The company also introduced Language Explorer, an interactive tool for visualizing LinguaMeta, an open repository containing information about more than 7,000 spoken, written and signed languages.
The importance of these projects goes beyond research. The data can help create tools for farmers, teachers, healthcare workers and other community members who need to communicate in their own language.
AI without an internet connection and for basic phones
A digital tool is not truly accessible if it only works with a fast connection and a modern phone. Google points out that more than 3 billion people still lack reliable internet access.
To address this limitation, the company developed TranslateGemma, a family of lightweight translation models based on Gemini and trained in 55 languages. Their design allows them to run on the device, reducing dependence on the cloud.
But even an efficient model can remain out of reach for people who use basic phones. That is why Google is supporting Viamo in developing Ask Viamo Anything, a voice assistant that brings Gemini features to conventional phones through interactive voice response systems.
The pilot program in Rwanda has reportedly used Gemini to answer more than 2 million questions. The idea is simple but powerful: access an AI tool without needing an app, an advanced screen or a permanent connection.
Accessibility also includes sign languages
Language is not limited to regional dialects. Millions of people communicate using sign languages, and many voice tools still struggle to understand nonstandard forms of speech.
Google is working on Sign Language-to-Text, or SL2T, a system trained with more than 50 sign languages. The technology powers sign-to-text features in Gboard and Live Transcribe on Pixel 11, beginning with American Sign Language and its translation into English.
The company presents this development as a first step toward improving the accessibility of digital products for the approximately 70 million people who rely on sign languages to communicate.
Pronouncing names correctly is also a way to respect culture
Some details seem minor until they go wrong. A navigation app that mispronounces the name of a city or street can cause confusion, but it can also erase part of that place’s identity.
In New Zealand, Google worked with Māori language experts to improve the pronunciation of geographic names in Google Maps. The company says that incorporating authentic pronunciations into its text-to-speech models helps better reflect the history and heritage of communities.
What good is an AI that translates many words if it does not understand how the names of your own region are pronounced? Linguistic accuracy also has a cultural dimension.
Broader AI, but still a work in progress
Google says its language technologies are already integrated into nine platforms, including Search, Android, Chrome, YouTube and Google Play, reaching more than 5 billion people.
However, reaching more users is not enough. The real challenge is creating systems that understand context, respect cultural differences and allow each person to express themselves without having to adapt completely to the technology.
The promise of AI for everyone is not measured only by the number of available languages. It also depends on who participates in creating the data, which voices are heard and whether the tools work under the real conditions of each community. That is the difference between translating a language and helping a person be understood on their own terms.
Original source
https://blog.google/innovation-and-ai/technology/ai/ai-for-every-language
