This is where applied business programmers part ways with philosophers of mind and cognitive scientists. However, the lack of an inner model of the world is a formidable limitation, while the reliability of purely statistical-inferential processes will never be adequate for some basic building blocks of modernity.
But arithmetically speaking, is any of that true? Is model progress truly "accelerating"? How can you compare the delta from GPT-2 to GPT-3 with the delta from GPT-5 to GPT-6 with no sense of irony? And, at the risk of being quite blunt, do you know the meaning of the term "exponential"?
On the MMLU benchmark, GPT-2 had an accuracy of 32.4%. GPT-3 improved this to 43.9%.
GPT-4 scored 86.4%. On GPQA Diamond, it scored 31%, vs GPT-5 at 86%. GPT-6 scores 96%.
If we take ARG-AGI-2, it would be 9.9% for GPT-5 vs 95% with GPT-6.
The benchmarks do not corroborate the picture you were drawing about improvements slowing down between GPT major versions.
I remember their capabilities like so:
GPT-3 could produce convincing looking texts, sometimes.
GPT-4 was somewhat smarter and would give more accurate answers. At that time image understanding was released, wasn't it?
Then with GPT-5 we have a reasoning model, another jump in capabilities.
Now, compare GPT-6 Astra with its ability to implement software, work on long horizon tasks, and visual understanding.