upvote
> current frontier models

> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

reply
I tested both myself and a weak bot against Astra xhigh, https://lichess.org/study/27lCQqDa. It's still pretty bad at chess, though it takes longer to devolve into illegal moves.
reply
So give me a falsifiable point in time, a model you would claim succeeds at the task. One does not get to hotfix-patch updater out of the pressures of reality. Today is the day.
reply
deleted
reply
The gap in capabilities is mostly quantitative and not qualitative.
reply
Is it? I am on the fence on this, but it does seem like there are some qualitative improvements between the models.

Not related to your post, but a fact I keep mulling over. The fact I don't trust the current crop of LLM's enough and I consider LLM's as a tech will hit a ceiling pretty hard, it doesn't mean parallel improvement curves won't spring up out of other research that will lead to much higher capabilities than currently.

reply
They still need supervision though
reply
deleted
reply
deleted
reply
The actual current frontier plays somewhere around GM level.

https://chessbench-ai.github.io/#leaderboard

It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.

reply
I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims.

> About their ELO ratings from their own website:

> A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.

I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..

Please folks at least use your AIs to read stuff before making claims.

AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.

A GM is 2600 they can beat me in under 20 moves...

Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.

Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.

reply
HN is no different than Reddit, or any social media for that matter, in that commenters pretend to read articles.
reply
Back in 2001, our social medium was Slashdot and no one ever pretended to read the article. No one read the article either. It was slashdotted most of the time anyways.
reply
that is if it even a human commenter at all
reply
State-sponsored psyop meta comments aside, the models obviously continue to get better, but there is still a lot of 'guard railing' required to keep even the latest models completely on-task. The chess example is interesting because it's clearly a well-studied and established domain so the rules, strategies, and whatever else is in the training data should make yield excellent results; but clearly there is some behavior in these systems that's difficult to engineer out.
reply
Exactly, especially when you touch “forbidden things”, like questioning why rust IS not the best system programming language, you will be punished so hard by “expert”s.
reply
HackerNews is Gell-Mann amnesia that refreshes on every comment on every thread.
reply
What levels are they actually at in your experience?
reply
Sub 1300 that's my rating in the singular official tournament I participated at.

But given how easily I can crush them and how often they want to make illegal moves (btw above bench seems to use a harness that pokea the model until it gives valid moves).

I would rate them around 500-800 big range but at that level it's all about if the model can recall an opening or not. If it plays good first 4-8 moves the person on the end will fumble for certain and they win.

I can play good/best moves till 14-15 moves if I remember the lines and find someone who falls for it.

If you could give them the lines as prompts like the best 20-30 openings then they will be around 700-800.

700 is around the rating for a human who doesn't know the tricks but can do bare minimum calculations and understands the rules thoroughly.

reply
As someone who used to compete for years and plays currently as a hobbyist, you’re absolutely correct. LLM’s are terrible at chess and if anyone wants to sober up their view on AI, try it yourself.

Anyone who casually plays on a regular basis can beat them more often than they lose. As you said if you just know the core openings (and end games, both of which you can get a handle on with modest effort) you will generally win.

Edit: reminder we had computers beating the best players in the world literally decades ago. LLM’s are remarkable tools but the current promises and expectations are ridiculous

reply
So you can see an actual game on that website, and the play seems pretty decent to me for a while (~1700 lichess = 1300 elo) until move 28 when black throws away their queen for absolutely no reason in an incomprehensible blunder.

In some ways this is reflective of the AI experience at large, sometimes shockingly competent but then also sometimes ludicrously incompetent.

reply
> even if I give them literal infinite time and all the subagents and internet access..

Don't use the word infinite in any CS claims. They can recreate or approximate monte Carlo tree search and it technically is still a correct solution in your framing of the problem so long they defeat you.

reply
The AI can write a chess bot program that will beat you.

You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.

We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing's flight guidance system to do so.

reply
It's not quite the same, but the in-flight chess game provided by Delta was known to be absurdly hard: https://news.ycombinator.com/item?id=46593395
reply
I believe I remember reading it was based on Glaurung's code (which eventually evolved into what we now know as the juggernaut Stockfish).
reply
I can write a chess bot program that will beat you. Does that mean I’m good at chess?

>If they cared to have it perform well in chess games, you'd see a different shape and behavior.

So the things they claim are on the verge of AGI actually aren’t? They need to be trained for specific tasks?

reply
They’ll never be AGI simply because the definition will be constantly updated to be some steps ahead of them.
reply
I'm pretty sure "competent at chess without external aids" has been on the standard AGI checklist since before personal computers were a thing. How can you claim an intelligence is general if it can't make sense of such a highly constrained board game? This is solidly table stakes.
reply
So AGI needs to be trained on something to work well on it. Lovely reasoning we have right here.

Delusion runs deep in HN circles.

I say that as someone heavily invested in AI startups and projects and as someone working in the field.

I think most people on HN should touch grass and find real human contact. Lmao

Incredible reasoning all around here.

reply
An AGI doesn't stand for 'perfect intelligence' it stands for artificial general intelligence.

And no an AGI system doesn't need to play chess on a certain level to be disruptive to you and me and whole industries. It only needs to be as good as a person and cheaper.

Just because you define AGI as something it doesn't has to be,doesn't mean i need to touch grass.

This chess comparision is one of the most ignorant and stupid arguments i have heard after the parrot thing

reply
AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI

Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count

reply
>LLM can’t beat an avg chess player.

Why should that matter?

reply
If something has general intelligence it should be able to read the rules of a game and follow them. Therefore an artificial general intelligence (AGI) should be able to do this.

So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.

reply
So we humans are not a general intelligence then?

And the stuff i'm using LLMs daily is just fake?

I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.

reply
> And the stuff i'm using LLMs daily is just fake?

It simply means that LLMs are smarter than you, but not smarter than the average person

reply
I would bet a lot of money that Astra can follow the rules of chess (perhaps if repeated within the context window). Also, this is a different argument than what I responded to.
reply
I can write you a benchmark to prove it even with a heavy handed system prompt Astra will make an illegal move during the course of the games first few moves are generally ok since it's just throwing out learned moves.
reply
I'd genuinely like to see the results of that.
reply
I would definitely take you up on that.
reply
> Why should that matter?

Because we want to use this as a replacement for humans, and the average human can learn the rules of chess without needing to see the rules explained hundreds of thousands of times in millions of games.

So, yeah, it matters if a model has millions of examples of something in its training set and still cannot follow the rules.

reply
We're not talking about learning the rules of chess here, but playing a competent game from just being shown the rules. Why is it so hard for people to keep track of the thread of discussion?
reply
But we are. The models can't even follow the rules: they try illegal moves all the time.
reply
The fact that LLMs can play chess at any level is a strong indication we are in AGI.
reply
Can they if they frequently make illegal moves?
reply
No it isn't. Computers could play chess long before LLMs, better than LLMs can in fact. That didn't make them AGI.
reply
I'm stating that certain folks are trying to use the software-generating product as an AGI/ASI and then complaining when it doesn't play chess very well.

People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.

reply
It's not even a "software-generating product". It's only half of it. Most of the heavy lifting is done by absolutely not-AI compilers, analyzers and the like. If not for these programs, written well before AI boom, them LLMs would be no better at programming than they are are at pure LLM based calculations or writing.
reply
Then why respond at all for the sake of responding?

We all know AI can code, but the question it all stemmed from what if it's AGI or GM level in chess on it's own.

You can't just back pedal from the statement that apparently being able to code a chess engine is the same as being good at chess.

I can write a chess engine that beats Magnus Carlson without AI that alone neither makes me GM level or AGI or any of the other claims the above comments seem to be making?

reply
He keeps posting with a particular type of tone.

He definitely needs to touch grass.

reply
Try to embrace hacker ethos and stop hating.

Y'all seem to miss the point of this forum. Building and hacking and science and engineering.

I swear there's a whole lot of you who just like to look down instead of up. There's a whole universe up there.

reply
I agree w/ this perspective. An agent with a harness that can run programs can solve a lot more than one without the harness. The AI system includes the harness, and it's not clear to me that AGI requires more than LLMs + code generation & execution are capable of.
reply
So AI is AGI in fields where code can't solve anything?

Is code omnipotent, I have been in software all my life and I would hard agree here.

Sure stuff LLMs can do with being good at parts of code reproduction is incredible. And honestly it's the new way to do a lot of things but I have not see an iota of proof that it can scale across the board.

For instance Maths is just code with different symbols and slightly less universally legible concepts.

AI is the best invention at figuring out or walking the search space and directionally doing logically computation over general software adjacent stuff.

But that's it, I am certain a bunch of companies will make a lot of money despite no AGI.

I think people either don't understand AGI or don't understand how real world works.

Until an LLM can bow it's head take responsibility for mistakes made and ensure they aren't repeated again with 100% confidence to the leadership it's inarguably a tool a rather questionable one at that.

reply
First, you're moving the goalposts. Second, it's not actually true that any existing frontier AI can write a chess bot program that can beat a 1600 player ... not unless the program is derived from Stockfish or some other leading engine that has been in development for decades.

> The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.

These comments indicate a complete failure to understand the technology.

I won't respond again.

reply
> I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access.

I don't believe this.

You refer to "subagents", so this is not just an LLM but an LLM with some kind of agentic harness. Any reasonable harness and prompt, given internet access and appropriately prompted to succeed on this task, is more than capable of firing up Lichess or chess.com and relaying moves back to you. The free levels will be enough to beat you.

A frontier model can also likely one shot a chess engine that plays at your level, again if given an environment in which it can do that.

I completely believe the LLM on its own can't play a full game of chess at your level. Though I'd bet that with enough reinforcement learning it is possible to train a pure transformer architecture to do that. We just don't do it because there are other approaches that play chess much better.

reply
Even if you take that website at face value, the ELO scores shown are relative to the other AI models tested, and not comparable to the ELO scores of humans who play against other humans.
reply
I wonder why they didn’t throw a real chess engine in there for a baseline. There are engines where you can set the elo in the settings, so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other.
reply
> so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other

As a 1500 elo human I can tell you that a 1500 elo chess engine doesn't play like anything like a 1500 elo human.

reply
This is true, but I'm not sure it matters? I was poking around at the lichess database recently and those elo calibrated bots are remarkably well calibrated, their rating variance sticks out like a sore thumb compared to human players even at similar game volumes. So it should still be a decent predictor of how good a human at that level is, even if the playstyle seems alien.
reply
I feel like every position is in the database so you could just lookup the most popular move for an arbitrary elo and that's the bot.
reply
These ratings seems very wrong, i have beaten GPT Astra max thinking in chess and my rating is close to 1500. The ratings here seem more accurate: https://chessbenchllm.onrender.com/

GPT-6 almost never suggests an illegal move anymore while even Sol still did so time to time

reply
"Elo is relative to the ChessBench field."

They are of course "wrong" if you don't read the faint fine print and sensibly interpret them as FIDE or similar ratings.

reply
Probably tells us that without labs explicitly training/tuning the models or designing the harness (with fast oracle) the LLMs aren't going to get good at those areas.
reply
If it's a GM then I'm Magnus Carlsen, https://lichess.org/study/27lCQqDa.
reply
deleted
reply
more like you lose intelligence in chess by maxing for coding... hence knocking back the claims of emergent intelligence
reply
Dumbest thing I’ve seen today
reply
Please do not post misinformation. They are not playing anywhere near GM level.

"Elo is relative to the ChessBench field."

reply
If you give the same task to an exceptionally intelligent human, who does not play chess and has only heard about it in passing, then they would be beaten by every child who has looked at the rules for more than 10 minutes.

What kind of intelligence is "playing <____> but we don't tell you the rules" supposed to test?

reply
I suspect (in a probably ignorant fashion) that this is because learning process has been reading a lot of algebraic chess notation (such as "1. e4 e5 2. Nf3 f6 3. Nxf6 gxf6 4. Qh5! +-") then, to play, generating more of it without considering the rules of the game. This is exactly how it's always felt to me when playing chess against LLMs. Sure, "1. e4 e5 2. Nf3 Nc3" looks innocent to somebody simply learning the syntax of algebraic notation, but that Nc3 by black is an illegal move.

An LLM is the wrong approach for playing chess.

reply
1. It’s hard to trust a 2026 paper that’s showing results for such old models.

2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.

3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

reply
Good science, properly digested and presented takes time.

The idea that anything other than a breathless blog post about the latest model snapshot is useless is really poisonous to proper debate on AI issues

reply
so prove it! get a public repo out there, have it play against some open source engines

also I think the operative letter in AGI is the G - and if the G is short for 'variably competent savant-like hyperfocus on certain kinds of software coding and not any other general skill' then its not really G at all, is it?

reply
I suck at chess. Are you saying I can't be intelligent?
reply
If you read all chess tutorials, strategy documentation and game archives on the internet and then would still suck at chess: yes.
reply
Declarative knowledge is not the same as procedural knowledge. You can read as many chess tutorials, strategy documentation and game archives as you like, they won't make you good at chess until you actually start practicing chess.
reply
is that what I'm saying? or am I talking about AGI? perhaps there's some irony here to be explored when it comes to basic reading comprehension gaps
reply
That's a polite way to put it. :-)
reply
> Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

A bash script can clone and build stockfish, feed in human moves, and reply. By your standard, this bash script would "destroy any human at chess."

Are you interested in assessing the intelligence of the model, or the intelligence of the tools the model can use?

reply
Maybe practically it doesn’t matter? Perhaps AGI is not the model but the model plus everything it’s got access to. If we’re modelling intelligence in the way we seem to have to to have any coherent definition of AGI, it seems to me <model + everything it can access> is always going to be more “intelligent” than <model> alone.
reply
That would mean we should consider any human with coding knowledge a chess grandmaster, which is obviously not the case.
reply
> People who are good at it rely more on experience and deep domain expertise

People are good are 1900 or 2100 above and the top ones who spend decades in the field i.e. deep expertise are well in the 2200-2700 range.

A 1100 player is none of these things, they are purely relying on strategic reasoning there is a good chance they cannot name a single opening or articulate clearly why a move was appropriate. 1100 is quite low bar.

reply
1100 at online speed chess or something, could be. I'm not that deep in the chess world but everyone I know that can make 1100 in official rating can name a dozen openings and most of the known tactics, and is pretty good at applying at least one opening.
reply
1100 lichess/chess.com does not represent real elo. I'm around 1400 online, I would still be unranked in the real world. The fact that I easily beat any model publicly available is not a great look for AGI.
reply
1100 is literally below the ELO you get by default as a beginner.
reply
> Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

Delusional, but then Claude fable also isn’t beating any human at chess, the engine is.

reply
Llm systems are not really build for adhering to a grammar (other than "a string og tokens").

It is also not clear whether the llm adhering to a grammar is necessary for intelligent agents.

Certainly,a harness can easily correct for it.

reply
It seems absolutely crazy to me to expect an LLM to code a solution to a problem while also not expecting it to be able to adhere to a grammar.
reply
How much support do we as humans need to get rules right?

I'm an expert in my field, read my comments, my gramma is shit.

reply
Why?

You might never have tried to program before, so I don't blame it on you.

But most programmers, even experienced ones, see grammar and type errors regularly.

reply
By the promise of it, llms should be able to both adhere to grammars, or go free form where necessary. I mean, doing math is supposed to be strict but in practice it's a somewhat educated random walk in the space of correct lean theorems.

Harnesses do correct things, sure.

reply
You are right. I am imprecise.

Languages allow a certain flexibility in their grammars - you can read a sentence without that adhering it exactly to the grammar.

Games and programming languages (including lean) does not allow this flexibility.

A very intelligent person would likely also reason in terms of probably outcomes before correcting a statement to adhering entirely to the grammar.

Certainly it must be like that, otherwise reviews in math was rendered moot.

Do we blame research mathematicians for not adhering to the grammar?

reply
Well, yes, PGN files have structure... But still, playing Chess with an LLM is so weird that I impulsively question the sanity of people attempting to do so. Do some people really believe training on TWIC PGNs would make an LLM a good chess player?
reply
> Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

It can be very interesting and even entertaining to know where models don't do well. I don't find something like chess to be very instructive about anything though, nobody is going to pay for AI to play chess at any significant scale even if it could do it perfectly.

> The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.

I have found that not to be the case, and I am not an AI power user or cheerleader by any means. There's probably a bunch of even "simplest" tasks where AI doesn't do well and might never. That doesn't take away from the cases where it works well and is a productivity booster. It doesn't even have to be solving millennium prizes or any other breakthrough creativity or reasearch, it's still very useful in places.

reply
I wonder how current models would fare. The ones they tested are fairly old now.
reply
This isn't how intelligence works. The LLM may not be able to play chess directly through inference, but it can write a program to do it and execute that program. Same as how human intelligence works. We can't fly, but we can build planes.
reply
Human beings can play chess directly without coding up a tool.
reply
Asking an LLM to play chess by writing algebraic notation is like asking a human to play chess blindfolded.

Yes some people can do it but most people can't even if they're unusually intelligent.

You really need to be giving the LLM a board representation.

EDIT: I see that they actually were giving the LLMs a board representation and they still played badly. Fair enough then.

reply
If they wanted to train an LLM to play chess they could easily do so.

But nobody wants that.

reply
Very poorly compared to the tools we have built. Similar to the LLM.
reply
Comparing to raw LLMs? Much much better.
reply
Poorly in what sense? I think human chess leagues are way more popular and fun than just playing a computer by yourself. Human oriented communities are always a vastly better experience than their digital counterparts.

There's more to games than simply winning you know.

reply
Thanks for saying this, feels like everyone has gone insane over this stuff.
reply
Humans don’t code a $game engine to play $game, they can just play it. It seems like you are the one that has gone insane.
reply
And how many years of direct play and study does it take for a human to get good at chess or any other game? Absolutely no human ever could be good at chess just by reading a few books, or even every book on chess. That's just not how the brain works. If LLMs could do that they would truly be superintelligence.
reply
> Absolutely no human ever could be good at chess just by reading a few books, or even every book on chess.

Maybe not, but you'd be surprised how little it takes.

A six year old child can learn the rules of chess well enough to be able to play legal moves only in a single day. And they can improve their game at a pace which is almost frightening to behold. I have taught children, and I've witnessed significant improvement materialise in a single game. LLMs have probably thousands of chess books, games, videos, etc in their training data, yet they are unable to even follow the rules.

This is, at the very least, interesting. It illustrates many of the things brains can do, which current ML systems in general, and LLMs in particular, can't.

reply
> And how many years of direct play and study does it take for a human to get good at chess or any other game?

Time is irrelevant to training; the more relevant comparison is "how many games does a human need to play to get diminishing returns".

reply
A week. My brother learned and was above 1100 online within 12 hours, after a few hundred games.
reply
No, learning is definitely not a sign of super intelligence. I know words don’t mean anything anymore, but that is simply general intelligence, despite the claims we have reached this milestone.
reply
No, but superhuman capabilities derived from ordinary learning is, which is what the parent comment described. Why is that not obvious?
reply
The story isn't so clear cut.

The caveat is: It depends on the task.

Are there reams of chess moves that the model can train off of? No.

Are there reams of math papers the model can train off of? Yes.

reply
> The caveat is: It depends on the task.

I think the line of criticism around LLMs sucking at chess makes more sense when you understand what the AI companies are saying about the future trajectory of these models.

The entire recursive self improvement story falls apart once you point out that there is not much "cross domain transfer learning". Meaning that training an LLM to become good at coding, math, etc, will eventually transfer into them being good at other skills that were not explicitly trained for.

Using games like chess which have little economic value is actually a good test for this. What's even more surprising about them sucking at chess is how much information about chess strategy exists in the training data.

reply
There’s multiple databases of games in algebraic notation. You can also, very easily rl train on pitting models against one another, even without mcts.
reply
> Are there reams of chess moves that the model can train off of? No.

This is as false as something can possibly be. There are open databases of millions of chess games spanning hundreds of years.

reply
It is even worse.. This is a classical reinforcement problem where data generation is easy because the rule set is pre-defined. So you really don't even need any data to start with (but would help).
reply
There are more possible game combinations than atoms in the universe, even those generation of valid game states are as you say pre-defined. that is why models cannot go this route and therefore are poor at chess
reply
Isn’t this exactly how AlphaZero was trained? The rules are known and well defined so the training process can generate games without any outside data.

The only reason LLMs are this bad at chess is because the labs don’t care about chess performance so they’re not going out of their way to train the models for it. The ability they do have is from what chess information happens to be in the training data, plus whatever general reasoning abilities they may be able to apply.

reply
>Are there reams of chess moves that the model can train off of? No.

For real??

reply
deleted
reply
[flagged]
reply
deleted
reply
Frontier labs don't care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. In fact there's a google paper on grandmaster level chess without search with a 270M transformer. Outside that, there was gpt-3.5-turbo instruct which was incidentally a 1800 lichess elo player that didn't make any illegal moves even after a few thousand moves. Frontier labs care deeply about automating knowledge work and computer use. They are working hard on getting models better and better, and they are succeeding. Astra is a step change on that front. So good luck i guess, if chess performance is your barometer.
reply
> Frontier labs don't care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player.

If the models were actually intelligent, the way that the boosters claim, they wouldn't need to be tuned to play chess in order to be good at it. That's kind of the point of intelligence, that it is generically applicable to whichever task one wishes.

reply
Thats just absolutly not true.

A human being has general intelligence and needs A LOT of training and finetuning to become good in chess.

And there is a relevant and significant difference between the expectation of an AGI and an ASI system.

reply
Pretty much this. Feed it a book or two on chess, and you should have a decent (or good) player. That's the generic intelligence people have. The aims is not to be supremely talented at something, but being able to read a manual and figure how to use/play something. Mastery can be gained overtime.
reply
If you gave a human a book or two on chess they would not become a decent player (they would be closer to 500-600 than 1100 ELO) and they would only get better after playing hundreds or thousands of games (often making illegal moves and moves that violate the rules of chess as they learn).

Your assumptions/intuition about generic human intelligence feels quite incorrect, considering LLMs currently play better than a brand new human player would (presumably without any attempt to fine tune them specific on chess, such as playing thousands of games).

reply
That's quite untrue. I taught my (adult) brother the moves, the only illegal move he ever tried against me (over his 6 first games) was a castle with a rook that already moved twice. Within a few hundred games (less than 500 for sure, he played 3 minutes blitz but always took at least 10 minutes analyzing his games) he was rated 1100 on lichess (which is like 1050 on chess.com and unranked in the real world).
reply
So your brother tried to make illegal moves while learning the game and it took your brother hundreds of games to get to be a decent player? I don't see how this contradicts anything I said...
reply
The _only_ illegal move a human might make as a beginner is a failed en passant or a bad castle. And yes, a few hundred games is all it takes to be better than any publicly available LLM at the moment.
reply
the discussion isn’t really about whether language models can become strong chess players though, the point is they seem to struggle to consistently make valid moves. Most humans don’t need to read two books to pick that up, just a couple lines of basic instructions
reply
That has not been my experience with new players, they regularly make invalid or incorrect moves even after detailed instructions especially in novel situations.
reply
> considering LLMs currently play better than a brand new human player would

They’ve ingested all the literature on playing chess, a brand new human player has not.

reply
Yes, but my point is that humans can’t even do the thing that the above comments are claiming humans can do (read a book or two and be decent at chess), and then they complain that LLMs can’t do the same thing (that humans can’t do either).

We seem to be moving goalposts to the point that humans don’t even live up to the expectations of the AI critics. The only way you get better at chess is by playing a lot of games and learning from mistakes, that goes for humans or AI agents, not simply by reading about chess.

reply
> The only way you get better at chess is by playing a lot of games and learning from mistakes

How can you play without being aware of the rules and how can you learn from your mistakes without knowing they are mistakes? That’s what I said about reading a book of two. It is to kickstart the process. Then mastery is gained over time through practice.

This kickstarting then gradual refinement is how most people learn. And the foundational knowledge stays. Even a basic player knows to not do illegal moves.

reply
Reading can kickstart the process, but you can also make random moves guided by some sort of system (such as a computer GUI) or learn by watching other players play. The overall point is that you learn through observation and lots of trial and error (whether you are a human or a computer). And beginners in chess often make illegal moves even after learning the rules, it's fairly common.

It feels like you're trying to say that humans never make illegal moves while learning chess, which doesn't match with my experience. I'm trying to understand your overall point.

reply
If humans were actually intelligent, they wouldn't need to train and practice to play good chess. I mean, what level do you think people without any practice or training are ?
reply
Except all these LLMs were already trained with hundreds of chess book and game databases and they still suck
reply
If all you do is read chess books, you'll be a shit player. Training and practice is what it takes to be great.
reply
Oh right. But if all you do is reading programming books you are an amazing programmer? Where is all the training and practice LLMs did to become so good at coding?
reply
It's called post-training, typically through some form of reinforcement learning, and is a significant part of modern LLM development.

You have the first stage, pre-training, which is learning from next token prediction. That's where the model memorises a lot of facts about things and generally gets good at forms of writing. It's like reading a lot of books on programming and reading through a lot of source code. It's learning how to autocomplete code, essentially. Doing that requires a developing a reasonable understanding of code, but it's also learning how to autocomplete bad code as well as good, and won't make it a "good" programmer.

Pre-training uses a method called Cross-Entropy Loss to update the weights of the network.

Then comes post-training. This is where the model is trained against huge sets of example problems, like fixing a bug, adding a new feature based on a spec, etc. They are set the task and try to complete it inside a training environment. Once they're done, their complete solution is evaluated (either by humans, or by some separate evaluation model that was developed based on human feedback) and they are updated based on whether the solution was good or not.

Post-training uses a different method called Proximal policy optimization to update the weights of the network.

So these really are very different forms of learning, and mainstream LLMs are not post-trained to be good at chess. They could be. You could easily create a reinforcement learning environment that evaluated and improved their ability to play and win at chess. The result would be a very strong chess playing AI, something we know is possible because the strongest chess playing programs we have are neural network based, but it is not a priority for AI companies.

reply
>Where is all the training and practice LLMs did to become so good at coding?

Coding is a matter of translating the natural language description of a problem to the code specification while keeping the semantics fixed (and imputing the unspecified semantics as necessary). It is not considerably more difficult than translating between two dissimilar natural languages. Chess isn't a matter of language translation, but a compute heavy game of finding the best move out of many possibilities with wide variation in the quality of each move. Chess takes directed practice and reinforcement whereas language translation does not.

reply
LLMs (and Humans) don't get really good from programming books lol. The training and practice is the actual code they predict and learn from in the process of predicting.
reply
Oh I see. So if someone just reads books AND actual code then they can become experts, got it. And by the way LLMs are also trained with probably hundreds of thousands of actual games not just books
reply
WTF even is this post?
reply
Contrary to popular belief, you need a lot of training on something for an LLM to be good and consistent with it.

People think that if one mention exists in the training set, then the LLM is perfect at it.

reply
Not one mention. Hundreds of books, articles and databases of games.
reply
OpenAI making the next model good at chess is not analogous to a human training to get good at chess. It is analogous to God creating Human 2.0 which now has increased chess playing ability. If LLMs were intelligent the way humans are, then the models that exist right now would be able to spend time improving themselves at chess and become good at it. They can't do this because they are not, in fact, intelligent.
reply
why cant models make a tool call to stockfish? its like saying model can't execute python for complex math calculations
reply
Because then it’s not playing chess, stockfish is?
reply
Exactly. All these nerds saying cars make bad submarines. Well duh.
reply
The last post on HN I read was about someone using LLMs to reverse engineer an Apple GPU driver for linux in a month. The top comment points out how the poster must have had specialist internal domain specific contact with Apple. But then the thread concludes that wasn't the case and that this would take domain experts years to do.

> "current frontier models need laborious oversight and guardrails on even the simplest tasks"

I feel this statement is extreme. I can't personally reconcile it with any of the projects we're regularly seeing get delivered largely by LLMs now.

What are you thoughts? Like, what's your position here? Even if you sincerely believe frontier models need laborious oversight on even the simplest of tasks, do you think that accurately captures and reflects the current state and progress of frontier LLMs?

Don't get me wrong, there's lots of things LLMs can't do well, but the idea that they're basically not helpful for even the simplest of tasks seems... disingenuous?

reply
I don't see why this is such a big deal. Nobody's using LLMs for chess, but even if they are, just give them Stockfish as part of their harness. They don't need to do everything themselves as long as they're intelligent enough to use tools.
reply
The fact that they can play chess at all despite having no specific training for it blows my mind, and the fact it doesn’t do the same for many others shows just how far they’ve come and how fast.
reply