upvote
Presumably it's not, but the TTS voice in the video sounds to me more like formant synthesis than diphone - it reminds me of my DECtalk.

The project credits does mention espeak (which is formant based) as well as various other TTS projects, although it sounds like they are only using the pronunciation part of espeak, not the voice synthesis.

https://github.com/moonshine-ai/moonshine#acknowledgements

reply
It certainly sounds similar, but seems more nimble with phonetic pronunciation in the demo.

Having it run on a pico would be pretty impressive =3

http://cmuflite.org/

https://github.com/festvox/flite

reply
> Having it run on a pico would be pretty impressive

Yes, although relative to the DECTalk DTC01, a Pi Pico is a beast !

Pico : dual core ARM @ 133 MHz, 2MB flash, 264K RAM

DECTalk: 68000 @ 10 MHz + TMS 32010 @ 20 MHz (5 MIPS), 256K ROM, 64K RAM

reply
deleted
reply
deleted
reply