upvote
You have committed the classic blunder of confusing your abstraction layers.

"Probabilistic next word prediction" and "humor" sit about as far apart as "modulating airflow with meat flaps" and "humor" do. One is an interface through which an action is performed and the other is a highly abstract capability.

Would you claim that a podcast comedian is fundamentally incapable of being funny because all he ever does is wiggle the air with his throat meat flaps? Probably not.

Absolutely nothing about "probabilistic next word prediction" forbids "making intuitive/orthogonal leaps in context". The interface is expressive enough.

And empirically? The "sense of humor" in LLMs is yet another "a function of model scale" capability. GPT-4.5 was reportedly funnier than both GPT-4o and o1. Fable 5 is reportedly funnier than Opus 4.x. It's one of those ever-elusive "big model smell" signs that are hard to measure with anything other than vibes.

Under the "humor as an opposed social intelligence test" family of hypothesis, what "being funny" reflects is the funny guy's ability to model and predict you and your reactions. For the comedian to be able to make the audience laugh, he must know his audience well, model it accurately enough to be able to spot the "breaking points" of humor, things they'd find unexpected and clever and thus "funny", and then weave those things into the jokes.

Then, a bigger LLM gets better at humor because it has a more accurate model of how humans think of things - including the "ha-ha" gaps. It's a "theory of mind" capability. It's not "special", it's just hard.

reply
Humor requires a sudden orthogonal leap from context. That's what a punchline is.

I think you're saying that you can eventually train models to arrive at that destination by training on existing jokes, effectively encoding these leaps as probabilities.

In that case, the model isn't actually making an intuitive/comedic leap; they're just following new probability chains in attempting to approximate examples they've seen in training.

I'm suggesting that something architecturally different is necessary to create a model which can make intuitive/comedic leaps.

Try to get a frontier model to write a clever, funny joke which hasn't been seen before. Or, try to get it to make an intuitive leap that leads to a novel discovery.

You can use them to guide your own efforts along these lines, as a sounding board. But with current architecture I just don't think either is possible for an LLM to do on its own.

reply
> Try to get a frontier model to write a clever, funny joke which hasn't been seen before.

This is what I refer to when it comes to larger models like Fable 5 being funnier. They are more capable of doing that. They can deliver that "sudden orthogonal leap from context" of yours more reliably.

It's not a "fundamental inability" and never was. If you crank the scale up and a capability appears, "current architecture" was never the problem.

reply