Eighty-nine percent of the prompts Vietnamese users send to Gemini arrive in Vietnamese.
Thailand sits at 87 percent, Indonesia at 84, and across the six Southeast Asian markets Google measured, close to seven in ten prompts come in something other than English. Google put the numbers out on 14 July in its first regional report on the assistant, beside a line saying its user base here more than doubled in a year.1
The English-first assistant lost that argument without anyone announcing it.
What replaced it is a real regional stack. AI Singapore’s SEA-LION family now runs across more than eleven Southeast Asian languages, with 760,000 downloads and 180,000 API calls a month behind it. In Jakarta, GoTo and Indosat’s Sahabat-AI covers Bahasa Indonesia, Javanese and Sundanese, and sits inside Dira, the voice assistant that takes spoken commands for Gojek rides and GoPay transfers.
Patrick Walujo, GoTo’s chief executive, called it digital sovereignty when the model launched in November 2024.
Read the model cards and the word thins out. SEA-LION’s current lineup is built on Llama, Gemma, Qwen and Apertus, which is to say on Meta, Google, Alibaba and a Swiss consortium. Sahabat-AI’s 8B model is about 50 billion tokens of continued pre-training stacked on a SEA-LION checkpoint that is itself a Llama.
The sovereign part is the last layer.
The sovereign part is the last layer.
Everything under it is borrowed, and more of it comes from Hangzhou each quarter. Alibaba’s Qwen took more than half of all open-source model downloads worldwide by March, close to a billion cumulative, on Interconnects AI figures reported by the South China Morning Post.
Language is the cheapest layer in the stack to own, and the first one a larger lab takes back.
Language is the cheapest layer in the stack to own, and the first one a larger lab takes back. Google’s seventy percent is the demonstration: it bought the gap shut, then pushed Gemini Spark out in the same languages the week the report landed.
Fluency is rented ground for anyone building here.
The instruction data is harder to lift. Sahabat-AI’s Gemma2 model was tuned on roughly 448,000 Indonesian instruction-completion pairs, 96,000 in Javanese and 98,000 in Sundanese, assembled with Gadjah Mada University rather than scraped off the open web. Those sets exist because somebody in Yogyakarta sat down and wrote them.
Dira then renews the corpus with every ride booked and every transfer sent.
A lab in Mountain View or Hangzhou can buy Javanese fluency, and the price is falling. It cannot buy the recording of a Surabaya customer arguing with a driver about a cancelled order, because that conversation only happens inside the company that runs the rides. GoTo owns the queries; the weights it can always re-base.
Which is the reading a Jakarta or Ho Chi Minh City operator should take into the next procurement meeting: treat the model as a rental, and spend the budget on the pipeline that fills it.
Every one of those Vietnamese prompts is a training example, and Google is collecting eighty-nine percent of them.
Footnotes
-
The same report notes the region generated five billion images with Nano Banana and close to a million songs with Lyria 3, which is a fair account of what the local-language fluency is mostly being spent on. ↩