sherpa-onnx

Offline speech-to-text, text-to-speech, VAD and speaker diarization without internet

What sherpa-onnx does

  • Offline speech-to-text, text-to-speech, VAD and speaker diarization without internet
  • Runs on Android, iOS, Raspberry Pi, WebAssembly and desktop with many language bindings
  • Apache-2.0

sherpa-onnx — straight answers

What is sherpa-onnx?

sherpa-onnx is listed under Speech & Transcription Engines, in the AI Models & Local Execution category on Flocci AI Tools. Offline speech-to-text, text-to-speech, VAD and speaker diarization without internet. It is free, with no paid plan attached, and it lives at github.com.

Is sherpa-onnx free?

sherpa-onnx is listed as fully free — there is no paid tier attached to it in the catalog. That makes it one of the 399 entries on Flocci AI Tools with no upgrade path built in.

What can sherpa-onnx do?

sherpa-onnx does 3 things the catalog singles out: Offline speech-to-text, text-to-speech, vad and speaker diarization without internet; runs on android, ios, raspberry pi, webassembly and desktop with many language bindings; apache-2.0.

What is the best free alternative to sherpa-onnx?

faster-whisper is the closest free alternative: it sits in the same Speech & Transcription Engines sub-category and is free. FunASR, Moonshine and NVIDIA Canary-Qwen 2.5B also start free. The full list is on the alternatives page.

See the full list →

sherpa-onnx alternatives

Compare all alternatives →

OpenAI Whisper

Speech & Transcription Engines
free
  • Open-source speech-to-text, run locally free
  • 99+ languages and translation
  • Robust to accents and noise
  • Powers countless apps

Deepgram

Speech & Transcription Engines
freemiumLeaving soon
  • Fast, accurate speech-to-text API
  • Real-time streaming and diarization
  • Voice-agent building blocks
  • Free credits to start

AssemblyAI

Speech & Transcription Engines
freemiumLeaving soon
  • Speech-to-text with audio intelligence
  • Summaries, topics and sentiment
  • Speaker diarization
  • Free tier for developers

NVIDIA Parakeet

Speech & Transcription Engines
free
  • Tops the Hugging Face Open ASR Leaderboard with 6.05% average WER at 0.6B params
  • CC-BY-4.0 license, free for commercial and non-commercial use
  • RTFx ~3,386 — extremely fast inference relative to accuracy

Mistral Voxtral

Speech & Transcription Engines
freemium
  • Apache 2.0 open-weight models (24B and 3B) downloadable from Hugging Face
  • Built-in Q&A/summarization over audio, not just raw transcription
  • Function-calling directly from voice input for agent workflows

Speechmatics

Speech & Transcription Engines
freemiumLeaving soon
  • Free starting credit with no card required, 56+ language support
  • Model-training discount program cuts costs 33% for data-sharing customers
  • On-premises/VPC deployment for privacy-sensitive enterprise use

Head-to-head comparisons