upvote
> 15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.

There must be something really of with those benchmarks. Yes, hallucinations gotten better, but I don't see that the big frontier models got so much better in the last 12-18 Months. They just put out bigger wall of texts and feel smarter. But they still make way too many stupid errors

reply
12 months ago "way too many stupid errors" was constant news. Today, you rarely hear about those anymore.

Sure, the novelty of the errors has worn off a bit and thus the reporting. Nevertheless the quality has improved immensely in this regard.

Also, AI video generation is now so good and accessible that it is very, very regularly used for memes, disinformation and proper (short) movie projects. AI image generation even more so (Mitch McConnell anyone?).

Pretending progress hasn't been mindboggling is insane.

reply
> Today, you rarely hear about those anymore.

Maybe people just got bored of reporting and reading about them.

reply
Maybe it got a lot less and I just got used to it. True.

Still feels too much for me. Breaks my workflow for no reason. Too much overhead for me, if I can't trust the output

reply
Ah yes, memes are clearly the sign of massive progress.
reply
No. We need objectively around 192 to 512gb of very fast memory to be able to run really useful models. I don't see local hardware with these specs coming in 1 to 2 years. There are a big number of initiatives currently taking place to increase ram output. But it will take another 3 years minimum to close the current supply issues. China is fast pacing forward to have its own chip baking factories with small enough nano scales to have fast chips. Will also take a few years.
reply
> 10-15 years? The current rate is closer to 10-15 months.

The leaps between models have gotten smaller and smaller. 2023-2024 models were rocketing up in quality. 2024-2025 I’d say was pretty impressive too. But 2025-2026? Very easy to feel the slowing pace of improvement. I agree 10-15 years is overly conservative but 10-15mo is far too bullish.

reply
The speed of model releases, in my view, is actually getting faster and faster. There were nearly nine months between GPT-3.5 and GPT-4. And now in just over one month, major models already included Claude Fable 5, Claude Sonnet 5, the GPT-5.6 series, Kimi K3, GLM 5.2, Qwen 3.8 Max, Grok 4.5... and the official DeepSeek V4 release is coming soon.

Iteration speed is now measured in days.

reply
Agreed. That said, I think months wins out as “closer to.”
reply