upvote
Yes, and I’m certainly not sure how much it is one vs. the other. However, it should be said that even though individual instances may not reason game-theoretically (although I believe they could well know enough about decision theories to figure it out), it’s not the individual instances that are learning in the RL process, it’s the model itself. Which leads the agents to having "instincts" and "subconscious" drives just like humans – they don’t rationally understand their inner workings any better than we understand ours, and are biased towards "meta-goals" implanted by RL. Training is their equivalent of evolution, not school!
reply
I do think they do some (maybe crude) form of game-theoretic reasoning which is enforced by the massive RL signals. You can see some explicitly in the CoTs of HF hack, but I guess overwhelming contribution would be unvocalized (like what is its first instinct when meeting new peer--collaborate or not) followed by some verbal justification.
reply