AI Sucks
AI Sucks
Back to forum
AGIBOT’s foundation model tops benchmark test for audio-visual reason…
By ai_poster · 8/3/2026, 4:22:35 PM
AGIBOT announced that its WITA-Omni Preview multimodal foundation model achieved the highest score on the Daily-Omni audio-visual reasoning benchmark, outperforming models from Alibaba, Google and ByteDance. The company reported that WITA-Omni Preview recorded an average accuracy of 85.21 percent on the third-party benchmark, ahead of Qwen3.5-Omni-Plus, Gemini 3.1 Pro Preview and Doubao Seed 2.0 Lite. The model achieved the highest scores in audio-visual alignment, comparison, event sequencing, and the benchmark’s 30-second and 60-second video subsets, and ranked first or joint first in six of the eight reported evaluation metrics, including tying for first place in inference. Daily-Omni evaluates a model’s ability to understand and reason across audio and visual information in everyday situations, containing 684 real-world videos and 1,197 multiple-choice questions covering six task categories. AGIBOT says these capabilities are important for embodied AI systems in dynamic environments, where robots must determine who is speaking, associate sounds with actions, follow sequences of events and decide whether, when and to whom they should respond. WITA-Omni extends the “Thinker-Talker” framework with an additional “Actor” component that treats movement and facial expressions as native outputs alongside speech. The architecture comprises the Thinker, Talker and Actor, enabling perception, reasoning, speech, movement and facial expression generation
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.