upvote
[delayed]
reply
Have you played Arc 3? It seems like more of a simple optimization problem (think Sokoban) than anything approaching fluid intelligence. Whether a multi hundred billion dollar company would spend time benchmaxxing a highly publicized benchmark that claims to confer AGI is an exercise left to the reader, but I doubt Claude Plays Pokemon is suddenly going to get past Mt. Doom now.
reply
> Claude Plays Pokemon is suddenly going to get past Mt. Doom now.

I miss him... But for reference he did get past Doom and got pretty far in the strength puzzle too before he cut cut off. He was looping and just brute forcing it.

reply
Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores.

I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.

reply
If they were benchmaxxing, surely they would score higher than 30% on ARC-AGI.
reply
Doing a quick search it seems like the average human score is 49%?

I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.

reply
nah they could make educated guess about arc and benchmaxx it too.
reply
deleted
reply
It is pretty clear at this point that current models are good at maths and problems with verifiable rewards. And puzzles are essentially math problems. Still a long way before we can say their "fluid intelligence" is effectively applicable to the real world.
reply
I keep wondering why there aren't more real world tests.

Maybe hook up a bunch of the AIs to a stereo camera and a couple of microphones and give them control over actuators to so they can drive cars. Then lets race them around a somewhat complex course.

When they are good enough at driving on tracks, put them on the road. Maybe see which can drive a truck with 400 cases of Coors from Texarkana, TX to Atlanta, GA and back within 28 hours.

reply
reply
It’s so weird to think how computers have mastered stuff that we used to think took intelligence (like chess, go, mathematics problems) but are doing so poorly at things any idiot can do (like drive a car).
reply
It’s still “only” at 30%, and “fluid intelligence” isn’t very well-defined. The models are getting more capable, but what that means in absolute terms is anyone’s guess, because we don’t have a thorough understanding on what exactly constitutes human intelligence.

I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs.

reply
Doubleplus benchmaxxed
reply