But arithmetically speaking, is any of that true? Is model progress truly "accelerating"? How can you compare the delta from GPT-2 to GPT-3 with the delta from GPT-5 to GPT-6 with no sense of irony? And, at the risk of being quite blunt, do you know the meaning of the term "exponential"?
On the MMLU benchmark, GPT-2 had an accuracy of 32.4%. GPT-3 improved this to 43.9%.
GPT-4 scored 86.4%. On GPQA Diamond, it scored 31%, vs GPT-5 at 86%. GPT-6 scores 96%.
If we take ARG-AGI-2, it would be 9.9% for GPT-5 vs 95% with GPT-6.
The benchmarks do not corroborate the picture you were drawing about improvements slowing down between GPT major versions.
I remember their capabilities like so:
GPT-3 could produce convincing looking texts, sometimes.
GPT-4 was somewhat smarter and would give more accurate answers. At that time image understanding was released, wasn't it?
Then with GPT-5 we have a reasoning model, another jump in capabilities.
Now, compare GPT-6 Astra with its ability to implement software, work on long horizon tasks, and visual understanding.