Cohere Launches an Open Source Voice Model Specifically for Transcription
Cohere has unveiled its first voice recognition model, Transcribe, designed to deliver fast and accurate speech-to-text transcriptions. With just 2 billion parameters, Transcribe is optimized for consumer-level GPUs, making self-hosting accessible for more users.
Key Features of Transcribe
- Supports 14 languages: English, French, German, Italian, Spanish, Portuguese, Greek, Dutch, Polish, Chinese, Japanese, Korean, Vietnamese, and Arabic.
- Impressive accuracy: Achieves an average word error rate (WER) of 5.42 on Hugging Face’s Open ASR leaderboard, outperforming competitors like Zoom Scribe v1 and IBM Granite 4.0 1B.
- Scalability: Capable of transcribing 525 minutes of audio per minute, making it one of the fastest in its category.
Performance and Availability
Transcribe demonstrated a 61% average win rate over similar models in human evaluations for accuracy and usability, though it performs less robustly with Portuguese, German, and Spanish. Cohere plans to integrate Transcribe into its enterprise agent orchestration platform, North, and offers the model for free via API and on Model Vault, its managed inference service.
Rising Demand for Voice AI
With growing interest in automatic speech recognition, tools like Transcribe are increasingly vital for apps focused on note-taking, dictation, and voice analysis. Cohere’s recent financial performance indicates solid growth, and there’s speculation about a possible IPO in the near future.
For more details, see the original article by Ivan Mehta at TechCrunch.