If anything, in some domains frontier models have become worse - claudisms and chatgpt idioms are making them notoriously bad at prose without a considerable amount of prompting and tweaking.
However there are many other domains which have verifiable rewards in the process of learning them, despite their overall impact not being verifiable. For example, the life sciences, an LLM could be given access to data about an organism, and then make predictions about how a drug or gene therapy will affect that organism. In economics, LLM's could create models of behavior, evaluate predictions over time and see how well those predictions match reality.
>claudisms and chatgpt idioms are making them notoriously bad at prose without a considerable amount of prompting and tweaking.
I think this is because a lot of people genuinely like the claudisms, even though a small minority of technical people don't.
Anecdotally, creative writing communities do report drastic differences in quality between models (it is my understanding that Gemma 4 is especially praised) so there is something to it being verifiable but it's a far cry from being something that can be distilled into objective math/code-like evals. One could even say it comes down to vibes
>For example, the life sciences, an LLM could be given access to data about an organism, and then make predictions about how a drug or gene therapy will affect that organism.
Yes, biotech companies are doing this right now but ultimately they are just predictions and you still need cold hard biological data to ground them and iterate on. That's slow and expensive, especially for the juicy fields (human biology, food and crops, clinical trials) and you can't just plug billions of VC capital into the pipeline and hope for RSI. Physical (gotta procure all those labs and their equipment), human (gotta hire and pay specialists to run experiments), biological (gotta wait for organisms to reproduce, drugs to take effect, crops to grow), regulatory (gotta convince agencies that your fancy new drugs are legit) bottlenecks get in the way.
And biology is one of the easier fields where a path to RSI (if not accelerating) is conceivable. Good luck iterating on macroeconomics data where there are no replicates and 'experiments' take literal years if not decades.
>I think this is because a lot of people genuinely like the claudisms, even though a small minority of technical people don't.
My uncharitable take is that the abtruse jargon makes people feel smart for understanding it and gives a sense of belonging (as a closed circle of initiates who understand LLM cant), just like rationalists love to repackage old or unsavoury idea under new nerdy smart sounding names.
For example, they have become more and more unintelligible when you ask for explanations or descriptive text. They assume you see the same context as them and shortcut explanations.
I've had to craft a skill to get them to produce remotely understandable explanations of even mildly complex/non-mainstream subjects.
This may be Curse of Knowledge https://en.wikipedia.org/wiki/Curse_of_knowledge on their part, and it also impacts human experts but still. Becoming better at one thing does not mean you become better at everything else, although I do agree with you that the better they become the more things there are they become good at, but their ability is still quite jagged and maybe increasingly so.
Knowledge can occasionally get in the way of some tasks, such as a master illustrator trying to draw like a child. Or as you've mentioned, an expert trying to explain to a beginner.
Even so, I disagree with you about model progress on communication - maybe there's a little jaggedness between minor model versions, but Opus 5.5 is a much much better communicator than, say, Claude 3 Opus.
LOL. Human (and animal!) intelligence is notoriously jagged. We're all basically idiots except for very narrow areas where we focus.
Smarter just means smarter.
See for example Paul Framton https://www.theguardian.com/world/2013/mar/30/physicist-mode...
This is the classic error in mixing up knowledge / experience and intelligence.
Intelligence is the ability to quickly process, understand, learn, and use information. It does not imply having already been exposed to any given information.
A smart doctor may have terrible computer skills. But this is from lack of experience and interest, not lack of ability to learn how to use computers.
But take that smart doctor, and a relatively dumb person who is equally inexperienced with computers. Now give them a month and incentive to learn computers. The doctor will learn much more; that's intelligence.