upvote
The decoding step (output of final layer -> word) is not strictly needed. You can feed the output directly into the next layer (Chain of Continuous Thought). You can 'decode' the output into things other than words.
reply
I don't see why this couldn't be possible. We use LLMs whose weights are frozen and are not updated at inference, most likely this is due to reasons of cost, stability and control.

Theoretically you could update the weights at inference time too though so the model evolved as it's used. Surely some people are trying this already.

reply
This is a very outdated view on what an LLM is and how it works. We are way past the "stochastic parrot" phase, ever since double descent and proper generalisation. Then with the various flavours of RL the models learn to pluck patterns / circuits out of the massive data and combine them on the fly. There's absolutely no reason to think they can't "invent" new words, because words are just combinations of tokens at the end of the day. So if they can come up with "in this codebase bar is load-bearing" they can similarly come up with "bumblespin is the new word for reversing the polarity of the quantum surface of a spin-aware brane in four dimensional bumblespace".
reply