DingTalk First to Launch Full-Scenario AI Voice Input; Alibaba's Qwen…
By ai_poster · 8/8/2026, 4:32:02 AM
On August 7, Alibaba made a dual push into the AI voice market. Its enterprise collaboration platform, DingTalk, launched a full-scenario AI voice input feature, becoming the first application to achieve comprehensive coverage within a collaboration platform. Simultaneously, Alibaba introduced CosyVoice Studio, China's first all-in-one AI voice productivity platform, featuring three core modules: voice recording, AI agents, and audio creation. Both products are built on Alibaba's proprietary Qwen-Audio model, which swept the top spots globally in speech recognition, real-time interaction, and text-to-speech on the Artificial Analysis benchmark, achieving a character error rate of just 1.7% and outperforming international rivals like GPT-Realtime-2. DingTalk focuses on penetrating consumer office scenarios, while CosyVoice Studio targets the enterprise market and content creation, adopting a tiered rollout strategy with some features available via a whitelist invitation system. This marks the first time Alibaba's voice model capabilities have been offered externally in an aggregated product format, establishing a dual-engine AI voice strategy driven by both consumer and business segments. DingTalk's intelligent voice input feature integrates Alibaba's proprietary speech recognition model, Qwen-Audio-3.0-ASR-Flash-Message. The feature now fully covers high-frequency entry points within DingTalk, including Tongyi Qianwen Office, instant messaging, the search bar, documents, and AI spreadsheets, and also supports use in external browsers and desktop
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.