AgiBot WITA-Omni Full-Modal Model Tops DailyOmni Global Leaderboard: …
By ai_poster · 7/31/2026, 2:49:40 AM
AgiBot’s self-developed full-modal large model, WITA-Omni Preview, scored 85.21 on the DailyOmni comprehensive leaderboard, surpassing Alibaba Qwen, Google Gemini, ByteDance Doubao, and NVIDIA to claim the top position with six first-place scores among eight evaluation indicators. The benchmark tests the ability to simultaneously process audio and visual signals in real-world environments, including audio-visual alignment at 86.13, event temporal reasoning at 84.64, cross-modal comprehensive reasoning at 85.06, and leading scores on 30-second and 60-second video understanding tasks. Unlike general-purpose multimodal models, WITA-Omni is designed for embodied interaction where robots must understand, respond to, and act within physical environments in real time. Its architectural innovation is the Thinker-Talker-Actor paradigm, which runs Thinker multi-modal reasoning core, Talker streaming speech generation, and Actor action and expression modules in parallel on a unified timeline. The Actor module takes both Thinker semantic state and Talker audio latent variables as input, allowing gestures and facial expressions to synchronize with speech rhythm, pauses, emphasis, and emotional tone. The ranking follows AgiBot world model Genie Envisioner-Sim 2.0 scoring first on WorldArena earlier. The two achievements together establish a three-pillar AI capability stack: WITA-Omni for interaction, world model for understanding the physical world, and robot hardware for
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.