Mistral Voxtral

Apache 2.0 open-weight models (24B and 3B) downloadable from Hugging Face

What Mistral Voxtral does

  • Apache 2.0 open-weight models (24B and 3B) downloadable from Hugging Face
  • Built-in Q&A/summarization over audio, not just raw transcription
  • Function-calling directly from voice input for agent workflows

Mistral Voxtral — straight answers

What is Mistral Voxtral?

Mistral Voxtral is listed under Speech & Transcription Engines, in the AI Models & Local Execution category on Flocci AI Tools. Apache 2.0 open-weight models (24B and 3B) downloadable from Hugging Face. It is freemium — a usable free tier with paid plans above it, and it lives at mistral.ai.

Is Mistral Voxtral free?

Mistral Voxtral is freemium: there is a free tier you can work in and paid plans above it. It is one of 391 freemium tools in the catalog, all labelled the same way so you are never surprised by a paywall.

What can Mistral Voxtral do?

Mistral Voxtral does 3 things the catalog singles out: Apache 2.0 open-weight models (24b and 3b) downloadable from hugging face; built-in q&a/summarization over audio, not just raw transcription; function-calling directly from voice input for agent workflows.

What is the best free alternative to Mistral Voxtral?

NVIDIA Parakeet is the closest free alternative: it sits in the same Speech & Transcription Engines sub-category and is free. OpenAI Whisper, WhisperX and AssemblyAI also start free. The full list is on the alternatives page.

See the full list →

Mistral Voxtral alternatives

Compare all alternatives →

OpenAI Whisper

Speech & Transcription Engines
free
  • Open-source speech-to-text, run locally free
  • 99+ languages and translation
  • Robust to accents and noise
  • Powers countless apps

Deepgram

Speech & Transcription Engines
freemium
  • Fast, accurate speech-to-text API
  • Real-time streaming and diarization
  • Voice-agent building blocks
  • Free credits to start

AssemblyAI

Speech & Transcription Engines
freemium
  • Speech-to-text with audio intelligence
  • Summaries, topics and sentiment
  • Speaker diarization
  • Free tier for developers

NVIDIA Parakeet

Speech & Transcription Engines
free
  • Tops the Hugging Face Open ASR Leaderboard with 6.05% average WER at 0.6B params
  • CC-BY-4.0 license, free for commercial and non-commercial use
  • RTFx ~3,386 — extremely fast inference relative to accuracy

Speechmatics

Speech & Transcription Engines
freemium
  • Free starting credit with no card required, 56+ language support
  • Model-training discount program cuts costs 33% for data-sharing customers
  • On-premises/VPC deployment for privacy-sensitive enterprise use

Soniox

Speech & Transcription Engines
freemium
  • Real-time speech-to-text priced well under typical competitor rates
  • Combines transcription + translation in one streaming API call
  • Diarization and speech translation bundled at no extra cost

Head-to-head comparisons