upvote
Speech-to-text models predict the next token of text from the preceding tokens of text and the current tokens of speech.
reply
Thanks, I did some learning and it fell more into place.
reply