upvote
What do streaming implementations do when a bigger context reveals a different interpretation/parse? When I use Whisper in the terminal, I can see it going back and correcting itself. Are corrections off the table for a true streaming transcription?
reply
I am guessing here.

Corrections based on larger context should also be part of the streaming output -- maybe include replacement text for previous chunk/s identified by chunk ID.

reply
[flagged]
reply
My mind is boggled by how many implementations miss this.

Handy has Nemotron Streaming and it works fabulously, FWIW. I’ve vibed a kind-of-working Deepgram API server into it but haven’t gotten around to finishing it. It’s something that should exist IMO!

reply
Yeah! Streaming is crucial for any real time use.
reply
My very first vibe-coded app (~Sonnet 3.6) was a dictation client for personal use.

I don't really understand how streaming would work compared against my normal flows. When I dictate, I set a toggle and then do stream of thought as I poke around between windows. When I'm ready to 'flush', I navigate to some target and give it focus for the text to flow.

Does dictation software now keep sort of unfocused floater previews and come with to-clipboard shortcuts or similar?

reply
This is the main reason I lean on Deepgram over local services.
reply
You can absolutely do high quality, low latency, even multilingual local streaming nowadays. As the commenter above says, Nemotron 3.5 Streaming is awesome. We make heavy use of it in our transcription app.
reply