> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
Not related to your post, but a fact I keep mulling over. The fact I don't trust the current crop of LLM's enough and I consider LLM's as a tech will hit a ceiling pretty hard, it doesn't mean parallel improvement curves won't spring up out of other research that will lead to much higher capabilities than currently.
https://chessbench-ai.github.io/#leaderboard
It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.
> About their ELO ratings from their own website:
> A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.
I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..
Please folks at least use your AIs to read stuff before making claims.
AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.
A GM is 2600 they can beat me in under 20 moves...
Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.
Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.
But given how easily I can crush them and how often they want to make illegal moves (btw above bench seems to use a harness that pokea the model until it gives valid moves).
I would rate them around 500-800 big range but at that level it's all about if the model can recall an opening or not. If it plays good first 4-8 moves the person on the end will fumble for certain and they win.
I can play good/best moves till 14-15 moves if I remember the lines and find someone who falls for it.
If you could give them the lines as prompts like the best 20-30 openings then they will be around 700-800.
700 is around the rating for a human who doesn't know the tricks but can do bare minimum calculations and understands the rules thoroughly.
Anyone who casually plays on a regular basis can beat them more often than they lose. As you said if you just know the core openings (and end games, both of which you can get a handle on with modest effort) you will generally win.
Edit: reminder we had computers beating the best players in the world literally decades ago. LLM’s are remarkable tools but the current promises and expectations are ridiculous
In some ways this is reflective of the AI experience at large, sometimes shockingly competent but then also sometimes ludicrously incompetent.
Don't use the word infinite in any CS claims. They can recreate or approximate monte Carlo tree search and it technically is still a correct solution in your framing of the problem so long they defeat you.
You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.
We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing's flight guidance system to do so.
>If they cared to have it perform well in chess games, you'd see a different shape and behavior.
So the things they claim are on the verge of AGI actually aren’t? They need to be trained for specific tasks?
Delusion runs deep in HN circles.
I say that as someone heavily invested in AI startups and projects and as someone working in the field.
I think most people on HN should touch grass and find real human contact. Lmao
Incredible reasoning all around here.
And no an AGI system doesn't need to play chess on a certain level to be disruptive to you and me and whole industries. It only needs to be as good as a person and cheaper.
Just because you define AGI as something it doesn't has to be,doesn't mean i need to touch grass.
This chess comparision is one of the most ignorant and stupid arguments i have heard after the parrot thing
Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count
Why should that matter?
So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.
And the stuff i'm using LLMs daily is just fake?
I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.
It simply means that LLMs are smarter than you, but not smarter than the average person
Because we want to use this as a replacement for humans, and the average human can learn the rules of chess without needing to see the rules explained hundreds of thousands of times in millions of games.
So, yeah, it matters if a model has millions of examples of something in its training set and still cannot follow the rules.
People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.
We all know AI can code, but the question it all stemmed from what if it's AGI or GM level in chess on it's own.
You can't just back pedal from the statement that apparently being able to code a chess engine is the same as being good at chess.
I can write a chess engine that beats Magnus Carlson without AI that alone neither makes me GM level or AGI or any of the other claims the above comments seem to be making?
He definitely needs to touch grass.
Y'all seem to miss the point of this forum. Building and hacking and science and engineering.
I swear there's a whole lot of you who just like to look down instead of up. There's a whole universe up there.
Is code omnipotent, I have been in software all my life and I would hard agree here.
Sure stuff LLMs can do with being good at parts of code reproduction is incredible. And honestly it's the new way to do a lot of things but I have not see an iota of proof that it can scale across the board.
For instance Maths is just code with different symbols and slightly less universally legible concepts.
AI is the best invention at figuring out or walking the search space and directionally doing logically computation over general software adjacent stuff.
But that's it, I am certain a bunch of companies will make a lot of money despite no AGI.
I think people either don't understand AGI or don't understand how real world works.
Until an LLM can bow it's head take responsibility for mistakes made and ensure they aren't repeated again with 100% confidence to the leadership it's inarguably a tool a rather questionable one at that.
> The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.
These comments indicate a complete failure to understand the technology.
I won't respond again.
I don't believe this.
You refer to "subagents", so this is not just an LLM but an LLM with some kind of agentic harness. Any reasonable harness and prompt, given internet access and appropriately prompted to succeed on this task, is more than capable of firing up Lichess or chess.com and relaying moves back to you. The free levels will be enough to beat you.
A frontier model can also likely one shot a chess engine that plays at your level, again if given an environment in which it can do that.
I completely believe the LLM on its own can't play a full game of chess at your level. Though I'd bet that with enough reinforcement learning it is possible to train a pure transformer architecture to do that. We just don't do it because there are other approaches that play chess much better.
As a 1500 elo human I can tell you that a 1500 elo chess engine doesn't play like anything like a 1500 elo human.
GPT-6 almost never suggests an illegal move anymore while even Sol still did so time to time
They are of course "wrong" if you don't read the faint fine print and sensibly interpret them as FIDE or similar ratings.
"Elo is relative to the ChessBench field."
What kind of intelligence is "playing <____> but we don't tell you the rules" supposed to test?
An LLM is the wrong approach for playing chess.
2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.
3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.
The idea that anything other than a breathless blog post about the latest model snapshot is useless is really poisonous to proper debate on AI issues
also I think the operative letter in AGI is the G - and if the G is short for 'variably competent savant-like hyperfocus on certain kinds of software coding and not any other general skill' then its not really G at all, is it?
A bash script can clone and build stockfish, feed in human moves, and reply. By your standard, this bash script would "destroy any human at chess."
Are you interested in assessing the intelligence of the model, or the intelligence of the tools the model can use?
People are good are 1900 or 2100 above and the top ones who spend decades in the field i.e. deep expertise are well in the 2200-2700 range.
A 1100 player is none of these things, they are purely relying on strategic reasoning there is a good chance they cannot name a single opening or articulate clearly why a move was appropriate. 1100 is quite low bar.
Delusional, but then Claude fable also isn’t beating any human at chess, the engine is.
It is also not clear whether the llm adhering to a grammar is necessary for intelligent agents.
Certainly,a harness can easily correct for it.
I'm an expert in my field, read my comments, my gramma is shit.
You might never have tried to program before, so I don't blame it on you.
But most programmers, even experienced ones, see grammar and type errors regularly.
Harnesses do correct things, sure.
Languages allow a certain flexibility in their grammars - you can read a sentence without that adhering it exactly to the grammar.
Games and programming languages (including lean) does not allow this flexibility.
A very intelligent person would likely also reason in terms of probably outcomes before correcting a statement to adhering entirely to the grammar.
Certainly it must be like that, otherwise reviews in math was rendered moot.
Do we blame research mathematicians for not adhering to the grammar?
It can be very interesting and even entertaining to know where models don't do well. I don't find something like chess to be very instructive about anything though, nobody is going to pay for AI to play chess at any significant scale even if it could do it perfectly.
> The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
I have found that not to be the case, and I am not an AI power user or cheerleader by any means. There's probably a bunch of even "simplest" tasks where AI doesn't do well and might never. That doesn't take away from the cases where it works well and is a productivity booster. It doesn't even have to be solving millennium prizes or any other breakthrough creativity or reasearch, it's still very useful in places.
Yes some people can do it but most people can't even if they're unusually intelligent.
You really need to be giving the LLM a board representation.
EDIT: I see that they actually were giving the LLMs a board representation and they still played badly. Fair enough then.
But nobody wants that.
There's more to games than simply winning you know.
Maybe not, but you'd be surprised how little it takes.
A six year old child can learn the rules of chess well enough to be able to play legal moves only in a single day. And they can improve their game at a pace which is almost frightening to behold. I have taught children, and I've witnessed significant improvement materialise in a single game. LLMs have probably thousands of chess books, games, videos, etc in their training data, yet they are unable to even follow the rules.
This is, at the very least, interesting. It illustrates many of the things brains can do, which current ML systems in general, and LLMs in particular, can't.
Time is irrelevant to training; the more relevant comparison is "how many games does a human need to play to get diminishing returns".
The caveat is: It depends on the task.
Are there reams of chess moves that the model can train off of? No.
Are there reams of math papers the model can train off of? Yes.
I think the line of criticism around LLMs sucking at chess makes more sense when you understand what the AI companies are saying about the future trajectory of these models.
The entire recursive self improvement story falls apart once you point out that there is not much "cross domain transfer learning". Meaning that training an LLM to become good at coding, math, etc, will eventually transfer into them being good at other skills that were not explicitly trained for.
Using games like chess which have little economic value is actually a good test for this. What's even more surprising about them sucking at chess is how much information about chess strategy exists in the training data.
This is as false as something can possibly be. There are open databases of millions of chess games spanning hundreds of years.
The only reason LLMs are this bad at chess is because the labs don’t care about chess performance so they’re not going out of their way to train the models for it. The ability they do have is from what chess information happens to be in the training data, plus whatever general reasoning abilities they may be able to apply.
For real??
If the models were actually intelligent, the way that the boosters claim, they wouldn't need to be tuned to play chess in order to be good at it. That's kind of the point of intelligence, that it is generically applicable to whichever task one wishes.
A human being has general intelligence and needs A LOT of training and finetuning to become good in chess.
And there is a relevant and significant difference between the expectation of an AGI and an ASI system.
Your assumptions/intuition about generic human intelligence feels quite incorrect, considering LLMs currently play better than a brand new human player would (presumably without any attempt to fine tune them specific on chess, such as playing thousands of games).
They’ve ingested all the literature on playing chess, a brand new human player has not.
We seem to be moving goalposts to the point that humans don’t even live up to the expectations of the AI critics. The only way you get better at chess is by playing a lot of games and learning from mistakes, that goes for humans or AI agents, not simply by reading about chess.
How can you play without being aware of the rules and how can you learn from your mistakes without knowing they are mistakes? That’s what I said about reading a book of two. It is to kickstart the process. Then mastery is gained over time through practice.
This kickstarting then gradual refinement is how most people learn. And the foundational knowledge stays. Even a basic player knows to not do illegal moves.
It feels like you're trying to say that humans never make illegal moves while learning chess, which doesn't match with my experience. I'm trying to understand your overall point.
You have the first stage, pre-training, which is learning from next token prediction. That's where the model memorises a lot of facts about things and generally gets good at forms of writing. It's like reading a lot of books on programming and reading through a lot of source code. It's learning how to autocomplete code, essentially. Doing that requires a developing a reasonable understanding of code, but it's also learning how to autocomplete bad code as well as good, and won't make it a "good" programmer.
Pre-training uses a method called Cross-Entropy Loss to update the weights of the network.
Then comes post-training. This is where the model is trained against huge sets of example problems, like fixing a bug, adding a new feature based on a spec, etc. They are set the task and try to complete it inside a training environment. Once they're done, their complete solution is evaluated (either by humans, or by some separate evaluation model that was developed based on human feedback) and they are updated based on whether the solution was good or not.
Post-training uses a different method called Proximal policy optimization to update the weights of the network.
So these really are very different forms of learning, and mainstream LLMs are not post-trained to be good at chess. They could be. You could easily create a reinforcement learning environment that evaluated and improved their ability to play and win at chess. The result would be a very strong chess playing AI, something we know is possible because the strongest chess playing programs we have are neural network based, but it is not a priority for AI companies.
Coding is a matter of translating the natural language description of a problem to the code specification while keeping the semantics fixed (and imputing the unspecified semantics as necessary). It is not considerably more difficult than translating between two dissimilar natural languages. Chess isn't a matter of language translation, but a compute heavy game of finding the best move out of many possibilities with wide variation in the quality of each move. Chess takes directed practice and reinforcement whereas language translation does not.
People think that if one mention exists in the training set, then the LLM is perfect at it.
> "current frontier models need laborious oversight and guardrails on even the simplest tasks"
I feel this statement is extreme. I can't personally reconcile it with any of the projects we're regularly seeing get delivered largely by LLMs now.
What are you thoughts? Like, what's your position here? Even if you sincerely believe frontier models need laborious oversight on even the simplest of tasks, do you think that accurately captures and reflects the current state and progress of frontier LLMs?
Don't get me wrong, there's lots of things LLMs can't do well, but the idea that they're basically not helpful for even the simplest of tasks seems... disingenuous?