upvote
When you are running the model locally and trying to do other tasks at the same time, Whisper Large feels too clunky on normal hardware.

I like Parakeet because it's good enough, relatively light, and fairly fast. I'm using this for things like meeting transcription and dictation. Since I'm sending most of the text to an LLM to clean up afterwards, it works well enough.

I'm hoping for something the size of Parakeet (or smaller) but better quality. It feels like with all of the advances in making smaller models better in the LLM space, someone should be able to come up with a lightweight and better quality speech to text model.

reply
I confirm this as my experience as well. Parakeet did not perform well for me (English with heavy accent). Whisper large is much better. But once we get to model of this size look at qwen3-asr (1.7b parameters). I'm quite happy with how it performs (both transcription quality and speed on rtx5080). By the way the model is also available on openrouter and very/very cheap. Performance is reasonable (<1s response), but some requests are delayed (>10s)
reply
Thank you for this suggestion. I am now trying it in 8-bit quantization, and it seems to work very well, at least as well as Whisper Large. It is a little bit faster on my Mac, too.

The real verdict will be after several days of use with dual language dictation, but so far it looks really good!

reply
It’s really domain dependent I think. If you are doing anything conversational interfacing with less AI-familiar users, latency matters a lot.

If you’re feeding the results into a very smart LLM, it will figure out what you meant (but crucially ONLY if you warn it or tell it to do so, in some cases!). If you’re writing code directly or creating something for public consumption, you can’t tolerate mistakes. If you’re taking notes for yourself you just want it to work cheaply.

If you are ok with the complexity you can run both, and a Meta/Google open model with native audio, and let a smart LLM doctor it up. If performance really matters you can pay for a proprietary model or train one yourself. Until you get to that point, I think it probably doesn't matter much either way. Voice is just too easy to fiddle with

reply
For English I've found Parakeet v2 and Parakeet unified to perform significantly better than v3. For deposition audio in a quiet room and neutral american accents I've hit 96-97% on a number of sessions with v2 and I've seen as low as 83% for heavy accents with overlapping speakers.
reply
Whisper large-v3-turbo is still the SOTA non-cloud model? Damn. I'm wanting to move away from it to something smaller/better but based on this comment it seems like I made the right move by sticking with whiser.cpp this whole time. It's been 2 years since large-v3-turbo was released...
reply
Much to my surprise, this seems to be the case.
reply