upvote
Sorry about your father. He needs a dictation model, not a general purpose speech-to-text model. They ignore umms and ahhs, change things like “an elephant, no a monkey, went up the tree” to “a monkey went up the tree,” support saying punctuation aloud sometimes, etc.

Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge...

reply
For essentially infinite and fast dictation I use https://github.com/cjpais/Handy on Parakeet streaming (cohere is far better, but slower and has a token output limit so you cant ramble for many minutes). And then just do a cleanup pass with a cheap LLM, it will in my experience, do far better than trying to voice control to go edit a sentence or change words. I just weave instructions into my writing. I understand this requires technical know-how, but for those with it, this is the best solution I have found to long form writing without my hands.
reply
I literally today pushed v0.1 of my dictation app that does both in a single package. You can connect either to a cheap LLM but I've set it up out of the box to use Qwen3.5 9b which does the job well for free, and everything stays local, no telemetry. You can also use Claude/OpenAI/Local providers. https://github.com/lnenad/lipwise
reply
I like the look of this but I may have to do a port if I go ahead with it. Have a primarily Python on Android project that could use STT.
reply
Difference with voice ink?
reply
Open source, multi platform
reply
Does the AI processing work fast or consume less resources on 4-5 years old laptops with 8GB RAM?
reply
For slower machines it's much better to use a cloud model, but Qwen3.5 4B can do the job and is fast enough on older machines.
reply
Neat, thanks!
reply
Another plug for Handy, and wanted to share something cool about it.

You can set it "Push to talk" mode (like a walkie-talkie radio), and when you're done talking and release the button, it can paste the text into any text field.

You can even replicate ChatGPT voice conversation mode, by having Handy as your speech input, and then (I forgot the extension) enabling a speech-to-text model for OpenCode. Surprisingly relaxing flow for certain tasks, like tweaking a website's styles.

reply
Whisper flow does a good job of this too, speaking to transcribe directly into a focused text input
reply
Agreed, record and transcribe however you can and then use a LLM to clear out.

I prompt it to:

"Attached (or underneath) is the transcript of a self recording i've done with tons of rambling and some incorrect words transcriptions, please do a pass clearing out and arranging any typos or possible misunderstandings. Keep original in parenthesis when not sure if it's a misunderstanding. Do not summarize or alter the nature of the content, simply tidy the transcript."

reply
Original in parentheses is very interesting.
reply
Love Handy. It has become a core piece of how I use computers.

In case it's helpful to anyone else using it, at first it felt a bit slow to me, because there was a noticeable pause after I finished a message before it would quickly type it all out. I changed the input method from direct to clipboard and it's way faster now, almost instantaneous.

reply
FluidVoice is working on that I believe w/Fluid-1 model https://github.com/altic-dev/FluidVoice
reply
> And then just do a cleanup pass with a cheap LLM

I am especially interested in this part. Could you please share the prompt you are using to instruct the LLM to clean up the dictation? Thanks in advance!

reply
I do this with a summary prompt + a JSON config. The config keys are the full name of the thing that gets commonly mistranscribed, and the values include a phonetic spelling of the name and a n "entity id" which is just the name of a markdown file that provides information about the entity, which the LLM can use to infer what should have been referenced if it's not clear.

My prompt is pretty simple:

    - Intro of the purpose (faithful, high-quality summaries of meeting transcripts useful for readers who did not attend the meeting)
    - Some basic rules
        - Summarize only what is actually said in the transcript.
        - If a speaker's name isn't clear from the transcript, describe them generically (e.g. "a participant") rather than guessing
        - If the transcript is a fragment or cuts off, say so rather than inventing a resolution
        - Only state a causal or explanatory link between two facts (e.g. why someone is absent) if that link is clearly stated in the transcript
        - Preserve relative time references precisely rather than flattening them. If a speaker says something will happen "in two hours" don't reword it as "today" or "same day"
        - Maintain the original voice and framing to keep original intent
        - Attempt to infer meeting attendance based on who actually spoke or was spoken to
        - Prefer plain, direct sentences over compound or clever phrasing
        - Respond directly with the summary. Do not preface it with phrases like "Here is a summary" or "Based on the transcript".
    - A preferred structure for the summary, with examples
    - An instruction to attempt correction of mis-transcribed words based on the included JSON configuration, but only if the certainty is high. If unsure, leave as-is, but add (spelling?) after the word to flag it.
    - The JSON config as a string.
I invoke this with a skill that can also update the JSON config if I identify new mistranscriptions, so it gets updated regularly.
reply
Technical knowhow? Handy is a gui right? (Also Parakeet is great I use it everyday)
reply
Thanks for sharing Handy!
reply
+1 for handy and then using LLM's for the cleanup pass, though what are your observations on feeling as if sharing that output though?

Because I have seemingly mixed opinions on it, on one hand, I did put the effort but on the other, the output is AI generated so I am unsure about sharing it with others (because they might think its AI generated)

Do you use it for very small edits (removing just the uhhm's?) or for slightly more edits.

The way that I use it sometimes is that while thinking, I will write something which can sometimes make me feel as if a better re-write can better explain my thoughts or rephrasing it as such. For example. I will think about X topic, connect it to Y, then try to add some more points about X again.

I found LLM's to do a really decent job at generating the final outputs as such, but as I said, I am left sometimes feeling a little confused as to sharing it or not because of it being AI generated and the end user not knowing if I put an actual effort into creation of it or not.

Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it.

reply
what do you mean:

"what are your observations on feeling as if sharing that output though?"

and "Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it."

could you rephrase the question

for STT, it's literally you saying it, with a model transcribing, and then another model correcting a little, and you also have control over editing it. I don't think any of the arguments on "etiquette re: sharing AI output" apply here.

reply
> for STT, it's literally you saying it, with a model transcribing, and then another model correcting a little, and you also have control over editing it. I don't think any of the arguments on "etiquette re: sharing AI output" apply here.

Yes, for doing extremely mild edits (like just removing uhh's etc.) this might be true

but I sometimes feel as if writing can allow me to shift paragraphs, so if I am writing para 1, para 2, I can shift back to para 1 and write another sentence in it and edit some parts of para 1 to include that point

but when I am doing STT, although I can move towards the other para, I find myself just speaking in a complete flow and just write in para 2 only.

Thus when I ask AI to write, I would prefer it to move the statements to appropriate paragraphs and in just general, create a more comprehensible viewpoint from all the STT text that I had written.

This does generate AI generated text which can be detected as such. Uploading it on blogs makes me feel as if people might read what they might consider "AI slop" and so the ethics part (as I myself don't wish to read AI slop)

the problem with AI written or edited texts is that I am unsure of how much effort the other person has put in (just a single prompt or a detailed thought was put in), and I feel as if, others feel the same way.

Should one try to show the rough draft as well to try to show that it was an effort which was human generated or that human effort was used, but that means having a proper disclosure that it was AI-generated/AI-assisted, which I feel as if offputs a lot people (including me) as because of the above logic, that there's still friction for the user within testing if real effort was put into place and I am unsure how effective sharing drafts of it could be.

I don't want my blogs to be tainted and treated as AI-slop because I care about them so I am unsure of what to do. I have multiple things that I have written which if I pass through AI can create some meaningful blog piece but as it stands, they are rough drafts and I find myself putting low efforts or being lazy in actually editing them myself as well (and potentially putting in multiple hours) when AI can be used to help create a more polished version just as well and get across my point.

reply
I think you're overthinking it. you clearly care a lot about the output and respecting the reader. you're not the problem lol
reply
[flagged]
reply
[flagged]
reply
I've been using Gemini Desktop App purely for dictation. It's a miracle! For the first time in my life I'm blown away by the quality of my (heavy accent) speech recognition. Just be sure to disable the "speak to window -> reasoning" option to make it purely dictation and stop from writing whole emails for you.
reply
> I don't think the challenge with speech to text was size of the binary

The usecase for small models like this is making on-device STT/TTS more accessable. This is important if your usecase is sensitive to either privacy or latency, but this comes at the cost of quality.

My experience has been that these small TTS models are unexpectedly good if your audio is in distribution (western accents, higher quality audio, common vocabulary), but pretty quickly degrade as you move outside of that. They often dont support more complex features such as diarization, multilingual, or realtime streaming either.

reply
> I don't think the challenge with speech to text was size of the binary

It's a very compelling aspect of the problem.

If you can get a model under certain size thresholds, that means you can eliminate latency domains. For example, if the model is able to fit entirely inside L3, the latency of servicing requests drops by an order of magnitude (or better) compared with a model that resides in L3+DRAM.

reply
It's true that L3 is an order of magnitude lower latency than DRAM, but that's mostly going to translate into a throughput difference rather than a latency difference, when looking at the system as a whole.
reply
There are many different challenges, each requiring their own solution. I, for one, really miss the old Google Assistant on my Android phone. It would very reliably play most songs that I wanted to hear on Spotify. Gemini fails at this almost every time, and is significantly slower. It's actually a difficult problem, as the songs people want to hear are regularly being released, are often associated with uncommon names, or have words in unusual orders, so normal LLM style tools just don't cut it.
reply
This reminds me of how the thing iPhones had pre-Siri (so we're talking pre-2010), which was entirely offline, did a better job than even the most modern thing at "Play [one of the finite set of songs in my library]." I sometimes get absurd matches from bands I've never heard of, when the right answer is something right there in my library.
reply
It's odd that the matching algorithm does not simply prioritize the music already in your library, but I see something similar in other domains where machine learning/information retrieval is used. E.g., in Apple Maps, I might have the map centered over my location and type in a restaurant nearby. Often, Apple Maps will find a restaurant with the same name on the other side of the country. This strikes me as an easy thing to fix (and Apple Maps has had this bug from the beginning), but if somebody knows something abou this, I'd love to know. Maybe it's harder than I imagine.
reply
I know it makes me sound like a lunatic to anyone listening, but I always use the most condescending monotone with Siri in the car, like this. "Directions to Wal-Mart,... in... Beaverton,... Oregon." Doesn't matter if I've been to that Walmart 197 times since Apple Maps 1.0 came out, because if you don't be as explicit as , 1 out of 10 times, it'll decide that you must mean a random Walmart 12 states away. Or like "Wall Plastering Incorporated."

And even when it's getting it right, and if there's only one Walmart in Beaverton, it still to this day needs to ask "One option is Walmart on Expressway Road in Beaverton...." Maybe it's correct in its 0% confidence level there, since it's so bad, but... I don't get how you could design something that bad, even before LLMs existed. I feel like I could do better, even using their Speech-to-text engine, with the processing backend built of pure regexes and if/elses.

reply
It's just useless for me now, the change happened some 5 or so years ago.

"Hey Siri, play [song]"

Leads to, take your pick:

- "You'll need to unlock your iPhone first."

- "I couldn't find [song] on Podcasts" (??????)

- "Playing [a totally different song]"

- "I couldn't find any music by [song, but it thinks it's a band]"

- "Playing music by [song, again it thinks it's a band]"

reply
Yes!! I get all of those. My favorite is that 10% of the time it asks me on what app I want to play the music (Oh, Apple, you're suddenly deeply respectful of competing on an equal playing field?).

And don't forget whatever the current phrasing is for "I'm sorry, my shit's all fucked up" and "My network connectivity had a blip and I'm unwilling to retry" and "Even though I have on-device STT models, and now LLMs too, and an on-device database of your music, which is downloaded, I won't bother without the cloud.

reply
so what's your problem? sell your apple products. I Know... you're from california, so this is heresy.

but once you calm down and stop hyperventilating from my suggestion, you'll see that the only reasonable and pragmatic course of action is to move to android and linux.

reply
Settings -> Accessibility -> Side Button -> under "Press and Hold to Speak" choose "Classic Voice Control"
reply
This is because, in the case of a restricted set of possibilities, voice recognition circa 2000 was actually very very good.

If you can do something with an extremely limited vocab, voice recognition was fine using off the shelf microchips in the 70s, where you wired in a microphone connection and had discrete pins for output actions.

LLMs are basically only useful for utterly free form transcription, but that doesn't actually help you turn that into tasks to perform and parameters for those tasks

The core "problem" in voice recognition is that freeform speech is an abysmal UX paradigm and provides zero discoverability, and LLMs IMO have not improved the situation of actually doing anything with the resulting text.

The other day I tried to prompt Gemini 3 times to tell me what the heck the business with a weird sign I saw was. The first prompt worked with a stale location context and therefore was way off, the second prompt had to reach out to google servers, and came back with recognizing the physical space I was discussing, but told me that I was talking about an event that takes place in the museum next door that I had told the model was next door to the business in question, the third try it still seemed to understand where I was referencing, but insisted I couldn't possibly be talking about anything there.

It took 1 second on google maps to find exactly what I was referring to, which was the business in Google's system located at the exact map location the model had found.

I'm sick and tired of people turning to LLM and "AI" tools to pretend they are better, when the problem is that these companies don't even use existing good solutions because they just don't care.

reply
One of the most common things I did with google assistant was tell it to remind me to do something, and it would reliably create a reminder on my calendar. With Gemini, it is very unreliable. Sometimes it does a google search. Sometimes it just opens a Gemini chat where it parrots back a (sometimes garbled) paraphrase of my request. Etc.

It's also really terrible at recognizing names of my contacts, probably because those names are not represented in the training data.

reply
deleted
reply
I hope you will find a solid solution for your father.

I was researching STT for people with speech disorders two years ago and essentially everything was boiling down to three problems at the end of the day - data scarcity, irregularity of way of speaking and thus constant ambiguity in translation, and individual differences in speech patterns among patients.

reply
This is a pretty cool one I saw recently: https://youtu.be/_j806JHhCRo?si=tr8ENJo_VlJF-b4y
reply
Have you tried playing music from his childhood for 20 minutes a day before his writing sessions?

In some cases, this may improve function for a few hours. Best regards =3

reply
i mean for something this small, it can be fit into a l3 cache on a cpu and be essentially always on various purposes
reply
Unrelated: I love your username.
reply