Unisound U2-ASR and U2-TTS Get Comprehensive Multilingual Upgrades, Enabling Global Listening and Speaking with a Single Model
Repost Caption: Pushing the boundaries of speech technology for progress, connecting diverse worlds for good. Upholding AI advancement and benevolence, Unisound pursues the original aspiration of inclusive AI for all!
At the opening ceremony of the 2026 World Artificial Intelligence Conference, President Xi Jinping pointed out in his keynote speech Work Together to Build a Fair and Reasonable Global AI Governance System: “All countries should uphold the philosophy of putting people first and pursuing progress for good. We should make artificial intelligence an important driving force for common prosperity and shared security, and work together to build a fair and reasonable global AI governance system.”
This vision charts the course for AI development. The “progress” of technology represents hard‑power capability to push limits and pursue excellence; the “benevolence” of technological application embodies the warmth of empowering the public and bridging divides. Two sides of the same coin, together they define the core values of the AI era.
Responding to this call of our times, Unisound has recently completed a full upgrade to the multilingual capabilities of U2‑ASR and U2‑TTS. U2‑ASR adds recognition support for 13 new international languages, while U2‑TTS adds synthesis for 8 Southeast Asian languages. With this update, the U2 Speech Large Model supports over 100 Chinese dialects and more than 15 international languages. This upgrade marks not only another leap in Unisound’s technical prowess, but also a firm practice of the philosophy of “AI for Progress” and “AI for Good”.
01 AI for Progress: Technological Breakthroughs Scale New Heights for Speech Capabilities
Technological “progress” means continuously pushing capability boundaries and translating cutting‑edge innovations into stable, efficient, scalable infrastructure ready for deployment.
Previously, enterprises building multilingual speech systems covering multiple countries and regions often needed to procure, deploy and maintain multiple separate models. This brought high costs, long lead times, inconsistent interfaces, complex system integration and heavy maintenance burdens.
The latest upgrade of U2‑ASR and U2‑TTS breaks down these technical barriers. With a unified model, standardized interfaces and consistent services, enterprises can build global‑ready speech interaction capabilities with far greater efficiency.
U2‑ASR: One Unified Model to Understand the World
In this upgrade, U2‑ASR further enhances its unified multilingual joint modeling architecture, adding support for 13 international languages: Arabic, German, Spanish, French, Indonesian, Japanese, Korean, Portuguese, Russian, Turkish, Vietnamese, Thai and Italian.
Addressing disparities across languages in phonesme sets, speech rhythm, intonation patterns and accent distributions, U2‑ASR leverages unified joint modeling to enable collaborative learning and capability sharing across linguistic features.
Enterprises can now process audio content in diverse languages by integrating just one model and one set of APIs, without deploying and maintaining separate recognition systems.
Under standardized testing protocols, U2‑ASR was benchmarked against leading industry models including Whisper Large V3 Turbo, FunASR and Seed‑ASR, evalsuating performance in two scenarioses: with explicit language tags provided, and without language tags fed into the system.
- With explicit language tags: On the industrial‑grade multilingual test set, U2‑ASR achieves an average Character Error Rate (CER) as low as 6.58% across the 13 languages. Twelve languages record a CER below 8%, with five key languages — German, Spanish, Indonesian, Italian and Vietnamese — hitting CER under 5%.
In comparative benchmark tests shown in the chart, U2‑ASR delivers the lowest CER across the board and outperforms competing models. Vietnamese achieves a CER as low as 3.01%, while Russian and Indonesian reach 5.26% and 3.41% respectively, demonstrating stable recognition strengths across diverse language families. Compared with the best‑performing alternative models including FunASR 1.5, Seed‑ASR, Whisper large‑v3‑turbo, Cohere transcribe and Qwen3‑ASR‑1.7B, U2‑ASR delivers an average relative error reduction of roughly 30%. The relative error reduction exceeds 50% for Vietnamese, alongside approximately 41% for Russian and 37% for Indonesian, reflecting substantial overall performance advantages.
- Without explicit language tags (closer to real‑world business scenarioses): For untagged audio streams, U2‑ASR retains high accuracy via automatic language identification and closed‑set routing, avoiding language misclassification and sharp error spikes. For cross‑border customer service, international conferences, multilingual content moderation and other use‑cases, reliable recognition results are available without pre‑processing to classify languages upfront.
U2‑TTS: Natural Speech for Global Communication
If ASR empowers machines to “understand users”, TTS determines whether AI can respond naturally in users’ native tongues.
In this upgrade, U2‑TTS prioritizes enhanced Southeast‑Asian language speech synthesis, adding support for eight languages: Thai, Vietnamese, Indonesian, Malay, Burmese, Khmer, Filipino and Lao.
Specialized optimizations target each language’s pronunciation rules, intonation conventions, pause patterns and prosodic features. This enables AI not merely to “speak”, but to articulate clearly, fluently and naturally in ways aligned with local linguistic customs.
- Natural and expressive output elevates localized interaction: U2‑TTS faithfully reproduces pronunciation traits, intonation shifts and speech cadence of different languages, delivering more natural, approachable voice experiences for local end‑users. For enterprises expanding overseas, localization extends beyond UI and text translation to voice interaction, enabling products to communicate in ways familiar to local audiences.
- Streaming architecture delivers fast first‑packet response: Built on streaming neural acoustic models, U2‑TTS outputs high‑fidelity 24kHz, 16‑bit speech chunk‑by‑chunk. Its “generate‑and‑play” streaming mechanism cuts first‑packet latency and eliminates delays caused by waiting for full‑length text synthesis before playback. It meets low‑latency requirements for real‑time voice conversations, digital human rendering, intelligent customer service, in‑vehicle interaction, smart terminals and other time‑sensitive scenarioses.
Sample synthesized output of U2‑TTS multilingual synthesis: Unisound speech model updated, welcome to experience
02 AI for Good: Inclusive Aspirations Bridge the Digital Language Divide
“AI for Good” is measured not only by technical strength, but also by how many people and scenarioses technology can serve, and whether communities across nations and linguistic backgrounds can share the fruits of AI advancement.
For a long time, global speech AI resources have concentrated heavily on major languages such as Chinese and English. Many developing markets and less‑resourced languages suffer from scarce training data, high model‑building costs and limited access to quality services.
Unisound’s latest multilingual upgrade aims to bring high‑quality speech technology to more languages and users through unified models and standardized services, shifting multilingual intelligent interaction from high‑barrier development to low‑threshold invocation.
Ready‑to‑use capabilities bring global speech within reach
Unisound packages sophisticated multilingual AI capabilities into out‑of‑the‑box MaaS (Model‑as‑a‑Service). Cross‑border businesses, smart hardware manufacturers, digital human developers, edtech platforms, customer‑service providers and content platforms can integrate multilingual ASR and TTS rapidly via standard APIs to shorten R&D cycles dramatically. Instead of building corpora, training models and constructing underlying infrastructure from scratch, enterprises can build global‑facing intelligent voice applications at low cost, turning technology into productivity for business collaboration and industrial growth.
Giving under‑resourced languages a voice, amplifying diverse cultures
Language serves not only as a communication tool but also as a carrier of culture and identity.
When languages such as Lao, Khmer and Burmese gain access to high‑quality AI speech support, technological value extends beyond interaction efficiency. It helps diverse languages and cultures be heard, understood and preserved in the digital realm.
From regional smart education and cross‑border medical communication to public services and cultural content dissemination, multilingual speech AI keeps expanding the frontier of inclusive AI, bringing intelligent services to broader populations and geographies.
03 From Listening to Responding: Closing the Global Speech Interaction Loop
From deep domestic expertise covering hundreds of Chinese dialects to global multilingual expansion, Unisound is building a comprehensive capability matrix spanning speech recognition, comprehension and synthesis. The synchronous advancement of U2‑ASR and U2‑TTS completes both the input and output ends of multilingual speech interaction, forming a full loop of “listening‑understanding‑responding”. It lays solid speech‑level foundations for global interaction in the Agent era.
Technological evolution must ultimately serve human needs and well‑being. Enabling AI to accurately interpret more languages and deliver natural, authentic AI voice output for global users embodies Unisound’s commitment to “AI for Progress” and “AI for Good”. Through continued inclusive practices, Unisound strives to tear down language barriers and bring cutting‑edge AI to address real‑world needs worldwide.
Moving forward, Unisound will keep pushing technical limits to make intelligent speech a core connector bridging cultural divides and linking global digital ecosystems. As a Chinese AI enterprise, Unisound will leverage multilingual interaction capabilities to build global intelligent bridges, facilitate cross‑linguistic exchange and intercultural communication, and contribute Chinese insights toward shared prosperity of the global digital economy.
Start your global speech interaction journey — Try it now
The newly upgraded U2‑ASR and U2‑TTS are fully available on Unisound’s TokenHub MaaS platform with open standard APIs for flexible invocation by developers and enterprises.
Developers, enterprises and individual users are welcome to visit the platform for hands‑on experience: U2‑ASR: http://maas.hongre128.com/models/asr U2‑TTS: http://maas.hongre128.com/models/tts