Mistral releases a new open source model for speech generation
Mistral Introduces Voxtral TTS
French AI company Mistral has unveiled Voxtral TTS, a new open source text-to-speech model designed for integration into voice applications, AI assistants, and enterprise solutions such as customer support. With this release, Mistral enters the competitive landscape alongside companies like ElevenLabs, Deepgram, and OpenAI.
Key Features and Capabilities
- Language Support: Voxtral TTS can generate speech in nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic.
- Custom Voice Creation: The model allows users to craft a custom voice with less than five seconds of sample audio, accurately capturing unique accents, inflections, and vocal nuances.
- Edge Device Compatibility: Due to its small size, the model can run efficiently on devices like smartwatches, smartphones, and laptops.
- Advanced Multilingual Support: Voxtral TTS can seamlessly switch between supported languages without losing the distinct qualities of the speaker's voice, offering advantages for real-time translation and dubbing.
Performance and Technology
Built for agility, Voxtral TTS boasts a time-to-first-audio (TTFA) of 90 milliseconds with real-time generation capabilities. A 10-second voice sample can be processed in roughly 1.6 seconds, ensuring responsive and interactive experiences.
The underlying architecture is based on the Ministral 3B model, which emphasizes human-like speech over robotic tones. The open source nature and customizable features make it appealing for enterprise adoption.
Strategic Vision
Mistral plans to offer a comprehensive platform capable of handling multimodal streams — text, audio, and images — to enable end-to-end agentic systems for businesses. This broader product strategy follows their earlier releases of transcription models optimized for both batch and real-time processing.
For further details, see the original article on TechCrunch.