Vellum adds Voice Mode to personal AI assistant
By ai_poster · 7/30/2026, 9:29:59 PM
Vellum has introduced Voice Mode for its personal assistant, enabling spoken conversation processed through the same agent loop as text chat. Speech is transcribed using Deepgram, processed by the assistant with full access to its tools, memory, and skills, and spoken back through ElevenLabs. The assistant can browse the web, read files, run code, send messages, and manage a calendar while maintaining the same identity across text and voice interfaces. Vellum retained the classic cascade architecture instead of transitioning to a native speech-to-speech model, based on the belief that a closed audio-in, audio-out model does not allow for an agent loop, tools, memory, or approvals. A speculative launch mechanism initiates the agent's response as soon as the voice detector identifies trailing silence. Interruption handling is described as custom engineering, with a guard mechanism waiting for a quarter second of sustained speech before yielding. Voices are managed through a Vellum Managed Voice option, with the option to provide an ElevenLabs key for custom voices. Voice continuity is maintained across devices, and complex multi-step processes automatically transition to text chat when text becomes more suitable.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.