Deep NNs and LLMs should not work based on our theoretical understanding. The fact that they do should give us a pause instead of us flatly denying their unexpected performance.
Tasks like using a CNN to detect digits has been well understood since the 2000s. The explosion in the capability of LLMs is very surprising, sure, but where is the concrete proof that such systems "should not work"?
token != step.
Just you try executing a complex command one word at a time.
They're doing RL on open problems these days, not just next token prediction.
...but many orders of magnitude faster, and in a way that scales horizontally really well, which is quite useful even if the quality isn't quite what a the absolute best humans can do.
One way AIs really really really excel is pulling together a lot of different data sources and reasoning over that data. In the past sure we could collect data and create huge datasets, but the analysis of that data - extracting themes, finding commonality or issues etc - either required extensive human research and analysis at best, or at worst crude regexes or keyword matching.
Now an AI can pour over that data and make its own inferences and decisions and findings that we've simply not been able to do before at this kind of speed or scale just because of time and resources.
And the AI, having done that, can propose new things for e.g. training, i.e. new things that no human has ever done before that the AI is simply repeating. For example it can propose a task that it knows from it's research is hard for it to solve currently, and then we just throw compute and randomness at it to find the "best" solution from many many attempts, then repeat until we hill-climb up to a perfect 1.0 score (... although of course we have to try and avoid cheating/attempts to short-ciruit the eval)
So this could be coding tasks, UI control tasks, protein folding, maths, chemistry etc etc. Anything that is easily and objectively programmatically scored. You can run this in a loop many times, each time you go around the loop the model gets smarter, learns more things from it's research, new areas of loss it can optimise etc etc.
It's harder where there is not a way to objectively score the outcomes (e.g. art, creative writing). Often this uses a fuzzy "judge" model that is trained specifically to give the work a score based on it's appraisal. This works but you can see how we might end up with feedback loops, so often it is paired with humans who provide feedback to provide supervised fine tuning datasets.
Tl:Dr - It's not just "repeating what it's seen". AI is finding new ideas and creating new things millions of times a day, and that is just software engineers asking it to write code or fix bugs, let alone people using it for actual research or whatever.
You think wrong ( https://dl.acm.org/doi/10.1145/3442188.3445922 ) ... except about irritating. Yes, its irritating to people insisting next-token predictors are intelligent.
Citation needed
To me, it is not at all obvious that the "level" of the training set is an upper limit to the capabilities of an LLM.
Sure, the LLM hasn't been exposed to material more advanced than the most capable human domain experts have produced. However, it has seen and learned from a vast amount of information that these domain experts are completely unaware of. Why shouldn't the LLM be able to use that information to produce output that's beyond the capability of a domain expert?