upvote
HN is no different than Reddit, or any social media for that matter, in that commenters pretend to read articles.
reply
Back in 2001, our social medium was Slashdot and no one ever pretended to read the article. No one read the article either. It was slashdotted most of the time anyways.
reply
that is if it even a human commenter at all
reply
State-sponsored psyop meta comments aside, the models obviously continue to get better, but there is still a lot of 'guard railing' required to keep even the latest models completely on-task. The chess example is interesting because it's clearly a well-studied and established domain so the rules, strategies, and whatever else is in the training data should make yield excellent results; but clearly there is some behavior in these systems that's difficult to engineer out.
reply
Exactly, especially when you touch “forbidden things”, like questioning why rust IS not the best system programming language, you will be punished so hard by “expert”s.
reply
HackerNews is Gell-Mann amnesia that refreshes on every comment on every thread.
reply
What levels are they actually at in your experience?
reply
Sub 1300 that's my rating in the singular official tournament I participated at.

But given how easily I can crush them and how often they want to make illegal moves (btw above bench seems to use a harness that pokea the model until it gives valid moves).

I would rate them around 500-800 big range but at that level it's all about if the model can recall an opening or not. If it plays good first 4-8 moves the person on the end will fumble for certain and they win.

I can play good/best moves till 14-15 moves if I remember the lines and find someone who falls for it.

If you could give them the lines as prompts like the best 20-30 openings then they will be around 700-800.

700 is around the rating for a human who doesn't know the tricks but can do bare minimum calculations and understands the rules thoroughly.

reply
As someone who used to compete for years and plays currently as a hobbyist, you’re absolutely correct. LLM’s are terrible at chess and if anyone wants to sober up their view on AI, try it yourself.

Anyone who casually plays on a regular basis can beat them more often than they lose. As you said if you just know the core openings (and end games, both of which you can get a handle on with modest effort) you will generally win.

Edit: reminder we had computers beating the best players in the world literally decades ago. LLM’s are remarkable tools but the current promises and expectations are ridiculous

reply
So you can see an actual game on that website, and the play seems pretty decent to me for a while (~1700 lichess = 1300 elo) until move 28 when black throws away their queen for absolutely no reason in an incomprehensible blunder.

In some ways this is reflective of the AI experience at large, sometimes shockingly competent but then also sometimes ludicrously incompetent.

reply
> even if I give them literal infinite time and all the subagents and internet access..

Don't use the word infinite in any CS claims. They can recreate or approximate monte Carlo tree search and it technically is still a correct solution in your framing of the problem so long they defeat you.

reply
The AI can write a chess bot program that will beat you.

You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.

We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing's flight guidance system to do so.

reply
I can write a chess bot program that will beat you. Does that mean I’m good at chess?

>If they cared to have it perform well in chess games, you'd see a different shape and behavior.

So the things they claim are on the verge of AGI actually aren’t? They need to be trained for specific tasks?

reply
They’ll never be AGI simply because the definition will be constantly updated to be some steps ahead of them.
reply
I'm pretty sure "competent at chess without external aids" has been on the standard AGI checklist since before personal computers were a thing. How can you claim an intelligence is general if it can't make sense of such a highly constrained board game? This is solidly table stakes.
reply
It's not quite the same, but the in-flight chess game provided by Delta was known to be absurdly hard: https://news.ycombinator.com/item?id=46593395
reply
I believe I remember reading it was based on Glaurung's code (which eventually evolved into what we now know as the juggernaut Stockfish).
reply
So AGI needs to be trained on something to work well on it. Lovely reasoning we have right here.

Delusion runs deep in HN circles.

I say that as someone heavily invested in AI startups and projects and as someone working in the field.

I think most people on HN should touch grass and find real human contact. Lmao

Incredible reasoning all around here.

reply
An AGI doesn't stand for 'perfect intelligence' it stands for artificial general intelligence.

And no an AGI system doesn't need to play chess on a certain level to be disruptive to you and me and whole industries. It only needs to be as good as a person and cheaper.

Just because you define AGI as something it doesn't has to be,doesn't mean i need to touch grass.

This chess comparision is one of the most ignorant and stupid arguments i have heard after the parrot thing

reply
AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI

Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count

reply
>LLM can’t beat an avg chess player.

Why should that matter?

reply
If something has general intelligence it should be able to read the rules of a game and follow them. Therefore an artificial general intelligence (AGI) should be able to do this.

So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.

reply
So we humans are not a general intelligence then?

And the stuff i'm using LLMs daily is just fake?

I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.

reply
> And the stuff i'm using LLMs daily is just fake?

It simply means that LLMs are smarter than you, but not smarter than the average person

reply
I would bet a lot of money that Astra can follow the rules of chess (perhaps if repeated within the context window). Also, this is a different argument than what I responded to.
reply
I can write you a benchmark to prove it even with a heavy handed system prompt Astra will make an illegal move during the course of the games first few moves are generally ok since it's just throwing out learned moves.
reply
I'd genuinely like to see the results of that.
reply
I would definitely take you up on that.
reply
> Why should that matter?

Because we want to use this as a replacement for humans, and the average human can learn the rules of chess without needing to see the rules explained hundreds of thousands of times in millions of games.

So, yeah, it matters if a model has millions of examples of something in its training set and still cannot follow the rules.

reply
We're not talking about learning the rules of chess here, but playing a competent game from just being shown the rules. Why is it so hard for people to keep track of the thread of discussion?
reply
But we are. The models can't even follow the rules: they try illegal moves all the time.
reply
The fact that LLMs can play chess at any level is a strong indication we are in AGI.
reply
Can they if they frequently make illegal moves?
reply
No it isn't. Computers could play chess long before LLMs, better than LLMs can in fact. That didn't make them AGI.
reply
I'm stating that certain folks are trying to use the software-generating product as an AGI/ASI and then complaining when it doesn't play chess very well.

People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.

reply
It's not even a "software-generating product". It's only half of it. Most of the heavy lifting is done by absolutely not-AI compilers, analyzers and the like. If not for these programs, written well before AI boom, them LLMs would be no better at programming than they are are at pure LLM based calculations or writing.
reply
Then why respond at all for the sake of responding?

We all know AI can code, but the question it all stemmed from what if it's AGI or GM level in chess on it's own.

You can't just back pedal from the statement that apparently being able to code a chess engine is the same as being good at chess.

I can write a chess engine that beats Magnus Carlson without AI that alone neither makes me GM level or AGI or any of the other claims the above comments seem to be making?

reply
He keeps posting with a particular type of tone.

He definitely needs to touch grass.

reply
Try to embrace hacker ethos and stop hating.

Y'all seem to miss the point of this forum. Building and hacking and science and engineering.

I swear there's a whole lot of you who just like to look down instead of up. There's a whole universe up there.

reply
I agree w/ this perspective. An agent with a harness that can run programs can solve a lot more than one without the harness. The AI system includes the harness, and it's not clear to me that AGI requires more than LLMs + code generation & execution are capable of.
reply
So AI is AGI in fields where code can't solve anything?

Is code omnipotent, I have been in software all my life and I would hard agree here.

Sure stuff LLMs can do with being good at parts of code reproduction is incredible. And honestly it's the new way to do a lot of things but I have not see an iota of proof that it can scale across the board.

For instance Maths is just code with different symbols and slightly less universally legible concepts.

AI is the best invention at figuring out or walking the search space and directionally doing logically computation over general software adjacent stuff.

But that's it, I am certain a bunch of companies will make a lot of money despite no AGI.

I think people either don't understand AGI or don't understand how real world works.

Until an LLM can bow it's head take responsibility for mistakes made and ensure they aren't repeated again with 100% confidence to the leadership it's inarguably a tool a rather questionable one at that.

reply
First, you're moving the goalposts. Second, it's not actually true that any existing frontier AI can write a chess bot program that can beat a 1600 player ... not unless the program is derived from Stockfish or some other leading engine that has been in development for decades.

> The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.

These comments indicate a complete failure to understand the technology.

I won't respond again.

reply
> I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access.

I don't believe this.

You refer to "subagents", so this is not just an LLM but an LLM with some kind of agentic harness. Any reasonable harness and prompt, given internet access and appropriately prompted to succeed on this task, is more than capable of firing up Lichess or chess.com and relaying moves back to you. The free levels will be enough to beat you.

A frontier model can also likely one shot a chess engine that plays at your level, again if given an environment in which it can do that.

I completely believe the LLM on its own can't play a full game of chess at your level. Though I'd bet that with enough reinforcement learning it is possible to train a pure transformer architecture to do that. We just don't do it because there are other approaches that play chess much better.

reply