Southeast Asia Is Building AI in Its Own Languages (Part 2)
- 5 days ago
- 10 min read
Eleven Voices, and Three Questions for the Creative Economy
Introduction
In Part 1, through SEA-LION, we examined the problem of a "language without sovereignty" and the background to Southeast Asia's project of building AI in its own languages. But SEA-LION is only one axis of this story. Part 2 surveys the region's language-model landscape—who is building what, and with which strategy—and takes up three fundamental questions this movement poses for the long-term creative economy.
I. The Regional Landscape: One Current, Many Channels
Southeast Asia's language models divide broadly into two tiers: pan-ASEAN models and individual national models.
Models Aimed at the Whole Region
Model | Developer | Characteristics |
SEA-LION | AI Singapore (Singapore national project) | 11+ SEA languages, fully open MIT, multimodal & agentic extensions |
SeaLLMs | Alibaba DAMO Academy | Mix of commercial and open, for SEA-language users |
Sailor / Sailor 2 | SEA AI Lab + academia | Qwen-based, open source |
These three are frequently benchmarked against one another (on SEA-HELM and others) and effectively compete to become the regional standard.
Strategies That Diverge by Country
This is where it gets interesting. ASEAN does not move as a single voice. Rather, each country takes a distinctly different path suited to its own conditions.
Singapore, beyond SEA-LION, also has MERaLiON. MERaLiON is one of Singapore's two national LLMs (the other being SEA-LION), downloaded more than 90,000 times since its December 2024 launch—signaling strong demand for AI that understands local nuance and context.[1] MERaLiON is particularly strong in speech/audio and empathetic response.
Indonesia has chosen the path of data sovereignty. Rather than framing itself merely as a host for servers, Jakarta has doubled down on digital sovereignty. NVIDIA's investment in Solo—through Indosat's Sahabat-AI initiative—is designed not to warehouse compute but to build a large language model in Bahasa Indonesia, trained on local context and national data priorities. This signals an ambition to refine intelligence, not just rent infrastructure.

Thailand leads with technology diplomacy. The prime minister has engaged in direct lobbying with global technology leaders including NVIDIA's Jensen Huang and Apple's Tim Cook, offering a green utility tariff that guarantees long-term, low-cost renewable electricity for green data centers. At the model layer, Typhoon (developed by SCB 10X) is the flagship, and like SEA-LION, it was made available for free use and further development by Thais. Its creative-economy relevance is concrete: Typhoon is built on Mistral-7B with an expanded Thai vocabulary and pretraining tuned for Thai vocabulary, context, and cultural nuance, and its natural home is exactly the work of local content—Thai-language customer-service chatbots and financial-advice assistants that read local phrasing far more fluently than a global model. This is the answer to "why bother with a local model": register, idiom, and politeness that a global model flattens are precisely where a domestic model earns its keep.
The same logic extends across the region. Malaysia's MaLLaM was trained from the ground up on Malaysian data, which in principle makes it better suited to the country's multilingual, code-switching reality—the "Manglish" texture of everyday Malaysian speech—that is central to authentic local advertising, dialogue, and script writing.
Vietnam and Malaysia belong among those making the most technically ambitious choices. SEA-LION, PhoGPT (Vietnam), and MaLLaM (Malaysia) remain anomalies as several of a small set of models in Southeast Asia pretrained from scratch. While most regional models fine-tune existing foundations, these exceptionally built from the ground up.
Malaysia is also aggressive on infrastructure. Malaysia has focused on fiscal incentives and infrastructure to attract AI capital.
What the Landscape Tells Us
Two structural features emerge from this picture.
First, the diversification of base models. From 2020 to 2022 (pre-ChatGPT), all the models in Southeast Asia relied on U.S. base models. In 2023, diversification began—toward local initiatives, France's Mistral, and international collaborative efforts. In 2024, the number of regional LLMs doubled, and Chinese models (specifically Qwen) accounted for a quarter of all new models.
Second, the difficulty of integration. Member states pursue markedly different AI strategies—Malaysia on fiscal incentives and infrastructure, Indonesia on data sovereignty and local-language models, Thailand on technology diplomacy and energy policy. Each is rational at the national level, yet collectively they illustrate why a single, binding AI regulatory regime would be difficult to implement in Southeast Asia. The reason is clear: the region is made up of countries with very different political systems, economic capacities, and levels of technological development, so copying a single global model designed for more uniform regions is neither practical nor realistic.
This very diversity and unevenness is a constant of ASEAN cultural and creative-economy policy. And language models are the lens that reveals that constant most sharply.
II. Three Questions for the Creative Economy
Now to the heart of the matter. What questions does this wave of language models pose for Southeast Asia's long-term creative economy?
Question 1 — Whose Data Is It? The Double Edge of "Heritage" as Asset
The most seductive—and most dangerous—proposition is this: heritage is a data point that cannot be ignored. The fact that AI can outperform global giants in local dialects shows that cultural and linguistic heritage can become a competitive asset. Thailand and Vietnam are indeed investing in "local language models" to prevent "algorithmic bias" where AI might unintentionally mirror Western social values over Eastern ones.
The optimistic scenario from a creative-economy standpoint is clear. When a region's folktales, dialects, oral traditions, and archives are "assetized" as AI training data, cultural heritage that was long undervalued acquires new economic value. The door opens for creators and small institutions too—as the ASEAN cultural-policy brief notes, easy access to digital platforms has democratized cultural exchange, allowing smaller organizations and individual practitioners to participate in regional and international dialogues.
But the same brief immediately names the shadow side. Digital transformation can lead to the homogenization of cultural content at the expense of local and indigenous content. More structural problems follow: unfavourable revenue-sharing models for local creators, digital piracy, and intellectual property protection.
Here lies a paradox worth dwelling on. The very process of "assetizing" heritage as data can simultaneously disembed it from its original creators and communities. The moment a Javanese oral tale—or the intricate narrative structures of a Wayang Kulit shadow-puppet performance, or the motifs of a hand-drawn Batik pattern—enters the model's weights to generate new scripts and images, whose asset is it? The institution that scraped and trained on it, or the community of Dalang (master puppeteers) that transmitted the story for a thousand years? Is the individual creator fairly compensated, or is their ancestral craft simply strip-mined for a tech company's feature update? These are questions cultural policy must answer now.
The legal issue of data provenance is already concrete. Even for Qwen-based SEA-LION, data provenance raises legal flags over copyrighted web material in multiple languages.
Question 2 — How Does Heritage Become a Sustainable Asset? The Problem of "Foundation Dependence"
The fundamental tension foreshadowed in Part 1 now surfaces fully. How intact is sovereignty built on someone else's foundation?
The case of Qwen-SEA-LION compresses this dilemma. AI Singapore acknowledges that some observers worry about the geopolitical dependencies inherent in adopting a China-origin foundation like Qwen, yet claims that sovereignty goals outweigh such concerns because open licensing ensures forkable roadmaps. The logic runs: even if the foundation belongs to someone else, because it is open source it can always be taken and modified, so it is not dependence.
This claim is only half right. Even if forking is permitted by license, the biases and worldview already steeped into the base model are inherited intact. Carnegie's report poses exactly this question—building on Qwen raises queries about source bias in its Chinese- and English-language corpora. The risk of escaping Western bias only to switch to another.
Why this matters for the creative economy is that, in cultural-content production, the "worldview of the foundation" governs the "aesthetics and narrative of the output." Whether drafting a screenplay, producing regional tourism content, or generating educational material, the base model's cultural defaults seep in subtly. This is not hypothetical. Studies find that Qwen, trained primarily on Chinese data, shows strong entrenchment in culturally specific views, and that multilingual Qwen models can even slip into unintended Chinese output when prompted in a non-dominant language unless smoothed after the fact. Ask such a model to write a story set at a traditional Southeast Asian wedding, or around a contested historical event, and its foundational weights may quietly default to Sinocentric norms and aesthetics, flattening the local nuance the content was supposed to carry. As one analysis of the region's shift to Chinese base models warns, if Qwen carries perspectives filtered to align with its corporate and political origins, Southeast Asia may have more than only Western bias to be mindful of. True cultural sovereignty is decided not at the license layer but at the data layer—and, ultimately, the architecture layer.
This is why Indonesia's "from scratch" path, and the ground-up pretraining of PhoGPT and MaLLaM, carry real policy significance. Expensive and slow, but the only route out of foundation dependence. Conversely, SEA-LION's pivot to continued pretraining followed the real-world logic of longer-run sustainability. This trade-off between sovereignty and sustainability is the core dilemma of Southeast Asia's creative-economy infrastructure.
Question 3 — Whom Is "Sovereignty" For? The Gap in Access
The last question is the most uncomfortable. Does the grand narrative of "sovereign AI" actually empower the region's creators and small cultural actors, or does it create yet another gap?
The ASEAN responsible-AI report names a painful reality. SMEs experience the highest barriers—cost, lack of local-language tools, and absence of practical guidance—creating a "Missing Middle" in accessible AI solutions. Large firms face their own obstacles: deep organisational and cultural barriers, with legacy systems and hierarchical processes complicating adoption.
This "Missing Middle" is fatal for the creative economy. Most of the cultural and creative industries are held up by freelancers, small studios, independent creators, and regional cultural institutions. Picture a small indie game studio in Manila that wants to generate endless NPC dialogue in Taglish—the Tagalog-English blend its players actually speak—or a solo webtoon artist in Bandung: neither can absorb the steep, per-call API costs of proprietary global models for open-ended creative iteration. If precisely this layer is shut out of frontier AI by lack of resources, "cultural sovereignty" remains a slogan of national-scale infrastructure and never descends into the creative ecosystem on the ground.
Yet the direction is right. For the region to gain the capacity not only to apply AI tools but to shape them in ways that support long-term competitiveness and innovation—this is the crux. SEA-LION's roadmap toward small models that run on local devices and domain specialization—letting that Manila studio run a capable local-language model offline on a standard laptop, no metered API in sight—can be read as an attempt to narrow this gap.
The question is whether this miniaturization and localization actually reaches the hands of the individual creator.
Policy Imagination for Eleven Voices
Southeast Asia's language-model wave is no mere technology trend. It is a vast experiment in who has the right to narrate their own culture, in their own language, in the age of AI.
To summarize, the three questions converge as follows:
The ownership problem: Assetizing heritage is both opportunity and risk of exploitation. Without governance that returns a fair share to creators and communities, "cultural AI" becomes a new form of extraction.
The dependence problem: Openness at the license layer does not guarantee genuine sovereignty. Without self-reliance at the data and architecture layers, one may merely trade Western bias for Chinese bias.
The access problem: If the achievement of national-scale "sovereign AI" fails to reach the "Missing Middle"—the small creators and cultural institutions on the ground—sovereignty remains a slogan.
And a Fourth Horizon: When Regulation Tightens, Cooperation Matters More
There is one more axis on the horizon. As global regulation tightens, the frame of the whole discussion shifts. Europe is the bellwether: from 2026, the EU AI Act requires every provider of a general-purpose AI model to publish a public summary of training datasets, and under the EU Copyright Directive creators can reserve their rights to prevent their work from being used in AI training—developers must check for copyright reservations and exclude or license that content before training. These obligations reach beyond Europe's borders: Article 53 of the EU AI Act imposes enforceable copyright obligations on all general-purpose AI providers regardless of where training occurs, with territorial scope extending beyond the Union.
The intent—protecting creators—is sound. But there is a bitter irony for smaller languages. As one Bruegel analysis warns, tightening opt-out and licensing rules may result in biased training datasets—especially penalizing smaller language and cultural communities, since leading AI models are already biased in favour of large English-language communities and lack cultural diversity. When permissible data shrinks, it is the low-resource languages—already data-starved—that are squeezed hardest.
This is precisely why, for language, cross-border cooperation and cultural exchange become more important, not less. Language data is inherently transnational: a single language spans borders, and the corpora, archives, and consented community datasets that low-resource models need cannot be assembled by any one country alone. As regulation fragments the global data landscape, the ability to pool resources, share benchmarks, and build trusted cross-border data flows—through frameworks like ASEAN's DEFA and interoperable governance—becomes the difference between a language that thrives in the AI age and one that quietly falls out of the training set. Sovereignty and cooperation are not opposites here; for language, they are two sides of the same page.
Eleven-plus languages, and as many cultures. For ASEAN to keep this diversity as an asset rather than surrender it to the pressure of homogenization—and to the narrowing of what may lawfully be learned—will take policy imagination as fine-grained as the technical sovereignty it pursues.
References
Singapore upgrades MERaLiON with more empathetic AI for SEA — Entelechy Asia. https://entelechyasia.com/2025/05/28/singapore-upgrades-meralion-with-more-empathetic-ai-for-sea/
SEA-LION, Multilingual and Multicultural Context Evaluation — Charunthon Limseelo / Medium. https://medium.com/@boatchrnthn/sea-lion-multilingual-and-multicultural-context-evaluation-5cad7295b46e
An Introduction to "Typhoon", the Market's Most Capable Thai LLM — SCB 10X. https://www.scb10x.com/en/blog/typhoon-innovative-thai-language-model
Speaking in Code: Contextualizing Large Language Models in Southeast Asia — Carnegie Endowment. https://carnegieendowment.org/research/2025/01/speaking-in-code-contextualizing-large-language-models-in-southeast-asia?lang=en
ASEAN Pushes Forward on AI Sovereignty — The Southeast Asia Desk. https://www.thesoutheastasiadesk.com/p/aseans-race-for-ai-sovereignty-begins
SEA-LION GitHub repository — AI Singapore. https://github.com/aisingapore/sealion
Building ASEAN's Responsible AI Ecosystem 2025 — EU-ASEAN. https://eu-asean.eu/wp-content/uploads/2026/01/AI-Paper-Jan-2026_050126.pdf
What Singapore's SEA-LION teaches us about the makings of local-language AI — GovInsider. https://govinsider.asia/intl-en/article/what-singapores-sea-lion-teaches-us-about-the-makings-of-local-language-ai
EU AI Act 2026: New Rules for Training Data and Copyright — Scalevise. https://scalevise.com/resources/eu-ai-act-2026-changes/
AI Training Data Copyright 2026: IP Risks & EU/US Rules — AI Governance Desk. https://aigovernancedesk.com/ai-training-data-copyright-rules/
The European Union is still caught in an AI copyright bind — Bruegel, Analysis 32/2025. https://www.bruegel.org/analysis/european-union-still-caught-ai-copyright-bind
Further Reading & Source Links
Regional landscape & national strategies
"The ASEAN AI Dilemma: From Zero-Sum Rivalry to Digital Strategic Autonomy," Modern Diplomacy, Jan 2026: https://moderndiplomacy.eu/2026/01/10/the-asean-ai-dilemma-from-zero-sum-rivalry-to-digital-strategic-autonomy/
"Why AI Regulation in ASEAN Is Harder Than It Looks," KAIRI AI / Medium, Jan 2026: https://medium.com/kairi-ai/why-ai-regulation-in-asean-is-harder-than-it-looks-f192faedae7c
"Advancing Southeast Asia's AI Future Through Sovereign AI Models," FULCRUM (ISEAS), Feb 2026: https://fulcrum.sg/advancing-southeast-asias-ai-future-through-sovereign-ai-models/
"ASEAN Pushes Forward on AI Sovereignty," The Southeast Asia Desk, May 2026: https://www.thesoutheastasiadesk.com/p/aseans-race-for-ai-sovereignty-begins
"Singapore's Multilingual AI Breakthrough With Qwen-SEA-LION," AI CERTs, Mar 2026: https://www.aicerts.ai/news/singapores-multilingual-ai-breakthrough-with-qwen-sea-lion/
Creative economy & cultural policy
ASEAN Socio-Cultural Community — Policy Brief No. 31 (2026), ASCC R&D Platform on Media, Culture and Arts
"Building ASEAN's Responsible AI Ecosystem 2025," EU-ASEAN, Jan 2026:
Regulation & training-data governance
"EU AI Act 2026: New Rules for Training Data and Copyright," Scalevise, Feb 2026
"AI Training Data Copyright 2026: IP Risks & EU/US Rules," AI Governance Desk, Jun 2026
B. Martens, "The European Union is still caught in an AI copyright bind," Bruegel, Analysis 32/2025
Foundational analysis
"Speaking in Code: Contextualizing Large Language Models in Southeast Asia," Carnegie Endowment for International Peace, Jan 2026: https://carnegieendowment.org/research/2025/01/speaking-in-code-contextualizing-large-language-models-in-southeast-asia?lang=en



