top of page
IMG_5940.jpg

ASEAN-KOREA

Cultural & Creative Sectors Research

Southeast Asia Is Building AI in Its Own Languages (Part 1)

  • Aug 5
  • 9 min read

SEA-LION, and the Problem of a "Language Without Sovereignty"

Introduction: Who Owns the Voice of AI?

Ask a global model like GPT-4 or Llama for farming advice in Javanese, or how to fill out a government form in Cebuano, and what comes back? Usually a stilted, translationese answer, or a reply that switches to English, or something plausible-sounding built on a mistaken cultural premise. This is exactly the void that local derivative models are now stepping in to fill: Sahabat-AI, co-developed by GoTo Group and AI Singapore, is an Indonesian-focused adaptation fine-tuned with 448,000 Indonesian instruction-completion pairs alongside 96,000 Javanese and 98,000 Sundanese pairs, supporting Indonesian, Javanese, Sundanese, and English. The issue isn't merely subpar performance. These global models were never built for Southeast Asia in the first place.


Catalysing AI Innovation forSoutheast Asia: SEA-LION is a family of efficient, open-source, multilingual, multimodal language models designed to understand Southeast Asia’s diverse languages, cultures, and contexts.
Catalysing AI Innovation forSoutheast Asia: SEA-LION is a family of efficient, open-source, multilingual, multimodal language models designed to understand Southeast Asia’s diverse languages, cultures, and contexts.
SEA-LION: Continued Innovation and Collaboration: Together with NVIDIA and our partners, both regional and global, SEA-LION will continue to push the boundaries of what’s possible, ensuring that Southeast Asia remains at the forefront of AI innovation. This partnership represents a commitment to ongoing research, development, and implementation of cutting-edge AI technologies tailored to the unique needs of the region.
SEA-LION: Continued Innovation and Collaboration: Together with NVIDIA and our partners, both regional and global, SEA-LION will continue to push the boundaries of what’s possible, ensuring that Southeast Asia remains at the forefront of AI innovation. This partnership represents a commitment to ongoing research, development, and implementation of cutting-edge AI technologies tailored to the unique needs of the region.

The SEA-LION team names the structural cause of this bias directly. Existing LLMs often display strong biases in cultural values, political beliefs, and social attitudes—because their training data, especially content scraped from the internet, is heavily influenced by Western, Industrialized, Rich, Educated, and Democratic (WIRED) societies. People from non-WIRED societies are less likely to be literate, use the internet, or have their output easily accessible, leading to a significant imbalance in the data.


This is the starting point for Southeast Asia's project of building AI in its own languages. This piece traces that movement across two parts. Part 1 examines the background, purpose, and technical characteristics of its flagship, SEA-LION. Part 2 surveys the language-model landscape across the wider region and takes up the questions this movement poses for the long-term creative economy.


What Is SEA-LION?

The name SEA-LION (Southeast Asian Languages In One Network) carries its design intent within it. It is a family of open-source Large Language Models anchored by the Products Pillar of AI Singapore and forming part of Singapore's National Multi-Modal LLM Programme (NMLP), built to better understand Southeast Asia's diverse contexts, languages, and cultures.


What stands out is that this is not a single company's product but a national project. AI Singapore is a national programme supported by the National Research Foundation, Singapore and hosted by the National University of Singapore. In other words, SEA-LION set out from the beginning with the character of both an industrial policy and a cultural one.


The Principle of Full Openness

The most striking choice in SEA-LION's design is transparency. The project commits to opening every one of these layers: pre-training data, model training code, model weights, fine-tuning data, and evaluation benchmarks. Its latest models are released under a fully open MIT license so they can be used by everyone without restrictions.


This openness is strategy, not merely goodwill. Only by releasing the models as open source can regional developers, enterprises, and governments build derivative models suited to their own uses—and only then can it grow from a single Singaporean model into an ecosystem for all of Southeast Asia.


Major Languages of Southeast Asia by Number of Speakers

Language

Total Speakers (approx.)

Primary Country / Region

Script

Indonesian (Bahasa Indonesia)

~200 million

Indonesia (national lingua franca)

Latin

Vietnamese

~90 million

Vietnam

Latin

Filipino / Tagalog

~87 million

Philippines

Latin

Javanese

~68–82 million

Indonesia (Java)

Latin / Javanese

Thai

~60–69 million

Thailand

Thai

Burmese

~40–43 million

Myanmar

Burmese

Sundanese

~32 million

Indonesia (West Java)

Latin / Sundanese

Lao

~27–30 million

Laos (+ Isan region, Thailand)

Lao

Malay (standard)

~19–33 million

Malaysia, Brunei, Singapore

Latin

Khmer

~18 million

Cambodia

Khmer

Cebuano

~16 million

Philippines (Visayas, Mindanao)

Latin

A note on the "Malay world" figure: Malay as a macrolanguage—counting Indonesian, standard Malay, and closely related varieties together—reaches roughly 290 million speakers. The table lists Indonesian and standard Malay separately, since they function as distinct standardized languages despite being mutually intelligible.

Figures are approximate totals (native + second-language speakers), compiled from Ethnologue-based and language-industry sources; ranges reflect variation across estimates.


Why these numbers matter more than they first appear

Here is the counterintuitive part: a large number of speakers does not mean a large amount of digital data. Javanese is spoken by nearly 70 million people, yet it has no official national status—Indonesians conduct all their formal, written, online business in Indonesian instead. The result is that one of the world's most-spoken languages leaves a surprisingly thin digital footprint: few news sites, few digitized books, little of the clean text an AI needs to learn from. The size of a language on the street and its size on the internet are two very different things—and it is precisely this gap that a project like SEA-LION is built to close.

The second pattern hides in the last column. Four of these languages—Indonesian, Vietnamese, Filipino, and Malay—are written in the Latin alphabet, which means they can quietly borrow the fonts, keyboards, and text tools already built for English. But Thai, Burmese, Khmer, and Lao each carry their own script, descended from ancient Brahmic writing systems, with no spaces between words and complex stacked characters. For an AI model, this is where things break: a system trained mostly on English has to learn not just a new language but an entirely new way of seeing text on a page. When a global model "fails" at Burmese, it is often failing at the script before it ever reaches the meaning.

Put those two facts together—thin digital data and unfamiliar scripts—and you have the real reason Southeast Asia cannot simply wait for global models to catch up. The languages that need AI the most are exactly the ones the mainstream models find hardest to see. Building for them is not a matter of translation; it is a matter of teaching the technology to read the region on its own terms.



Why Build It? The Keyword Is "Sovereignty"

The key to understanding SEA-LION is sovereignty—and here sovereignty carries three layers of meaning.

First, data sovereignty. 

Models like GPT-4 or Meta's Llama excel at English and major European languages but frequently struggle with the low-resource languages of Southeast Asia. Global models also fail to account for local cultural contexts or the region's propensity for code-switching—mixing English with local vernaculars, such as Singlish in Singapore or Manglish in Malaysia. Previous iterations of SEA-LION focused on creating a sovereign capability for the region, ensuring that Southeast Asian data isn't just a footnote in the training of US-based models.

Second, linguistic sovereignty. 

Southeast Asia is a region where more than 600 million people depend on diverse linguistic ecosystems. SEA-LION v4 is trained on over 1 trillion tokens with heavy emphasis on a curated Southeast Asian dataset, making it particularly strong in handling low-resource regional languages, dialects, and cultural contexts where global foundation models often fail.[6] This makes it a crucial enabler for digital equity in a region where over 600 million people rely on diverse linguistic ecosystems.

Third, cultural sovereignty. 

This layer matters most to those working in cultural policy. Across the region, a model like SEA-LION outperforming global giants in local dialects proves that heritage is a data point that cannot be ignored. It ensures that when an AI assists a farmer in Java or a worker in Hanoi, it speaks their language—literally and culturally.

This is not abstract. Thailand offers a clean example of a country using SEA-LION as the skeleton of its own cultural sovereignty and then localizing it. WangchanLion is an instruction-finetuned model based on SEA-LION—a pan-ASEAN pretrained LLM led by AI Singapore—and its finetuning is a collaborative effort between VISTEC and AI Singapore. The project has since deepened into full-scale Thai pretraining: WangchanLION-v3, released in collaboration between AI Singapore, VISTEC, and SCB10X, is an 8-billion-parameter model pre-trained on 47 billion high-quality Thai tokens, with those 47B tokens also released on AI Singapore's HuggingFace. Crucially, it treats data transparency itself as a public good: unlike other Thai pretraining models that focus on open-sourcing weights alone, WangchanLION-v3 openly shares comprehensive details from data-cleaning pipeline experiments to fine-tuning results, precisely because the lack of accessible pretraining corpora poses challenges for reproducibility and further research. This is what "cultural sovereignty as shared infrastructure" looks like in practice—one country's localization effort feeding open datasets back into the regional commons.


How Is It Built? A Technical Evolution

SEA-LION's technical trajectory carries an instructive policy lesson in its own right.


Building From Scratch vs. Building On Top

The early versions were built ambitiously from scratch. The first versions of SEA-LION, released in December 2023, were trained from scratch using SEA-LION-PILE (about 1 trillion tokens). But the team soon changed course. The newer version is based on continued pre-training of good open-source models—version 2 is based on Llama 3—because this approach may be more sustainable over the longer run.

This pivot compresses the real-world dilemma facing any low-resource region. When both data and compute are scarce, building everything from the ground up is ideal but unsustainable. So the strategy becomes: take a mature global model (Llama, Gemma, Qwen) as the foundation, then layer regional data thickly on top. And it is precisely here that the fundamental tension of Part 2 is hidden—if the foundation belongs to someone else, how intact is the "sovereignty" built on top of it?


Generation by Generation

SEA-LION has iterated rapidly. In v4, its first multimodal models extend capabilities beyond text to handle image + text inputs with massive 256K native context windows and specialized regional OCR, while continuing its focus on Southeast Asian languages, culture, and use cases. In v4.5, it uses rapid specialization of state-of-the-art open foundation models via knowledge distillation and model merging, delivering high-capacity reasoning and agentic tool-use capabilities.

The performance is striking. On the SEA-HELM benchmark, SEA-LION v4 achieves a top ranking among models under 200B parameters across Burmese, Filipino, Indonesian, Malay, Tamil, Thai, and Vietnamese tasks, and globally places #5 out of 55 models tested. What's notable is that the model not only outperforms open-source peers like Llama 3, Qwen 3, and Gemma 3 but also holds its own against proprietary models.


The multimodal turn is more than a spec sheet—its practical payoff shows up in the messy documents of everyday regional commerce. Take fintech: PetInsureX, a Thai team, built a privacy-first pet-insurance solution for Southeast Asia combining a SEA-LION multilingual assistant with automated OCR claims processing, manipulated-injury photo detection, explainable fraud scoring, and faster payouts—simplifying claims while providing multilingual explanations for customers. Reading a crumpled local-language receipt or claim form is exactly the kind of task where region-specific OCR and language understanding must work together, and where a global model's blind spots become operational failures.


Where Is It Headed? Smaller, and Into Industry

SEA-LION's roadmap points not toward "bigger" but toward "more precise, and closer." The team taps local partners from industry and government to create "offspring models" for lower-resource languages of Southeast Asian countries. The future of SEA-LION is moving toward smaller models and domain-specific fine-tuning to ensure higher impact in AI deployments.

Behind this lies the logic of field deployment. As embodied AI and edge AI (deployed directly from an endpoint device rather than a centralized cloud) become frontier sectors, the team is looking to build more small language models that can run on local devices. There is a clear recognition that this is not only about language: "It's not just the linguistic element, but juxtaposing it against the industry element," citing how healthcare in Southeast Asia looks different from how it looks globally.


Healthcare is where this juxtaposition turns concrete—and where localization can be a matter of safety. SIMIS, a Singapore AI-medical platform, integrated SEA-LION with custom vision models to recognize medical devices and provide real-time, step-by-step guidance in eight regional languages, reducing user errors and demonstrating how culturally aware AI can improve healthcare outcomes and safety at scale. In a low-resource clinic, a model that speaks the local language and can see the device in front of the health worker is not a convenience feature; it is the difference between correct and fatal use. This is the clearest answer to why "smaller, closer, domain-specific" is the destination, not "bigger."

SEA-LION is not simply "Southeast Asia's ChatGPT." It is a national and regional project built on the proposition that linguistic representation is itself a question of sovereignty in the digital age. Its fully open license, its continued-pretraining strategy, and its shift toward small, domain-specialized models are all pragmatic choices a low-resource region has made in order to have its own voice amid resource constraints.

But SEA-LION alone cannot speak for all of Southeast Asia. Indonesia goes its own way, Thailand another, Vietnam yet another—each pushing its own models and strategies. And beneath all of this movement lies that fundamental tension noted above: sovereignty built on someone else's foundation.

Part 2 maps this region-wide landscape country by country and takes up three questions this movement poses for the creative economy. Whose data is it? How does heritage become an asset? And what does the word "sovereignty" conceal?


bottom of page