The abilities of LLMs are emergent. You can experiment with LLMs and know as much about their observable behavior as the experts. But there’s no way to “crack open” an LLM and see precisely where each skill or tendency lives; as far as we currently understand, it’s all mashed together.
If stochastic gradient descent isn't taught in whatever CS Theory 101 is now, it really should be these days.
Lot's of neat stuff to learn.