Corrections based on larger context should also be part of the streaming output -- maybe include replacement text for previous chunk/s identified by chunk ID.
Handy has Nemotron Streaming and it works fabulously, FWIW. I’ve vibed a kind-of-working Deepgram API server into it but haven’t gotten around to finishing it. It’s something that should exist IMO!
I don't really understand how streaming would work compared against my normal flows. When I dictate, I set a toggle and then do stream of thought as I poke around between windows. When I'm ready to 'flush', I navigate to some target and give it focus for the text to flow.
Does dictation software now keep sort of unfocused floater previews and come with to-clipboard shortcuts or similar?