upvote
Unless you think that AI will quickly hit a wall (which seems odd considering that only a few years ago the best models had trouble doing basic math or counting the number of Rs in "strawberry"), I don't see how that's reassuring. The models will only get cheaper and more capable over time. It seems quite likely that at some point (probably before I hit retirement age) they'll be able to fully replace me at my job.

Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.

reply
Don't frontier models still have trouble counting letters? Or am I out of date? Either way, it doesn't seem to be the same amazing rate of progress we're seeing in other areas.

It's easy to look at a fire burning through a forest and extrapolate that rate of progress across the whole world. But fire doesn't burn everything equally fast.

What other cognitive tasks will be a struggle to make progress on? I suspect there will be some, though which ones they are is anyone's guess.

reply
Your information is out of date by several years. The letter counting issue was due to how LLMs split text input into tokens (usually using BPE). Since 2024, frontier LLMs have used chain-of-thought reasoning to spell out the letters and count them.
reply
> The letter counting issue was due to how LLMs split text input into tokens

No - this is provably not the issue.

Take any model that fails to correctly count the letters in a word, and ask it instead to spell the word (even a made up word), and it will be successful - they have no problem predicting the letter sequence from the token sequence (and would be shocking if they did - this is what they are built for: seq -> seq prediction).

The reason LLMs can fail at the letter counting task (depending on model training, prompting) is because of the counting part, not because of any difficulty correctly mapping the input token sequence to the letter sequence.

reply
The letter counting issue is due to tokenization. And most models still get this wrong often enough, even with reasoning. Probably less so on strawberry given how prevalent it is, and less so than without reasoning, but this not a historical issue. It’s becoming less of one though.
reply
My apologies, I got my info from an LLM. I guess they still have a ways to go in understanding current events.
reply
I just asked Opus 5.5 if any AI driven advances in mathematics have been announced in the last day or so and it gave me a summary of this OpenAI announcement. https://claude.ai/share/319b437a-c1f2-4119-8dc5-45d36545fed9
reply
Which LLM specifically? If it’s cloud based, you should be able to share the chat, right?

But seriously, I am still waiting for someone to wager that AI won’t be able to do a specific cognitive task in the next 5 years. This fact should be evidence enough that we have no idea how far AI capabilities will continue to advance.

reply
I asked both gemini and chatgpt "do frontier models still have trouble counting letters?" and the first word of both responses was yes.

The fact that no one is taking you up on that bet I don't find to be particularly persuasive. I suspect there will be plenty of cognitive tasks LLMs struggle with in 5 years, maybe even 20. But I wouldn't hazard to guess which, I don't think anyone is capable of that level of foresight.

reply
This is so strange I tried it on Gemini Flash: "Yes, but significantly less than before." When you read beyond the first word it explains where LLMs might fail and why.
reply
Ok, share links to the conversations with both models. I asked ChatGPT Astra 6 medium effort and it said, "Much less than they used to." and provided stats showing how accurate they are.[1]

1. https://chatgpt.com/share/6ac5e4cc-02f0-83e8-8f05-99a7ea2bf9...

reply
https://share.google/aimode/iKkrZtVYo4DSielWs

https://chatgpt.com/share/6ac63b6a-481c-83e9-a8fa-a13ce7402d...

I used whatever the default free model and thinking time was. If progress was really as fast and continually cheaper as some worry it is, wouldn't we expect free models by now to know (or even perform) what frontier models were capable of as much as 2 year ago?

This deep in the "comparing logs" tangent we risk missing the point. It's not what exactly frontier models are capable of at this particular point in time. But that there's entire categories of problems that seem easy to us which LLMs really struggle with. We've stumbled on several just a few replies into casual conversation. (Can they count? Can they know if they can count? Can they reproduce results? How quickly do new capabilities filter into free models? And that's just what's come up naturally, if we wanted to pick adversarial examples there's more to choose from.)

So while there's a number of difficult problems that are easy for LLMs (like bulk generating lean proofs), there are plenty of things where progress is not so impressive.

If LLMs can struggle so much with such easy problems, what hard problems have we yet to discover that they'll struggle with? The fact that no one knows, 5 years in advance, what those problems will be does not mean the chance of them is zero.

So far progress on the things LLMs are good at is fast and easy. It's like fire in a room full of oxygen. But once the low hanging fruit is gone, and the oxygen is out of the room. How fast will the fire burn through steel walls?

In my opinion it's a mistake to look at only rate of progress on one type of problem (whether it be what LLMs are good at OR what they're bad at) and assume progress on all tasks will progress at that rate indefinitely. Isn't there a saying about exponential curves, in nature, all being sigmoids eventually?

I guess we'll just have to see. I wish you good luck with your wagers.

reply
You are extrapolating from the mistakes made by free versions of smaller models to claim that frontier models struggle with easy problems. This is an obvious mistake in reasoning because as you can see from my shared Astra conversation, frontier models don't have the same limitation. (They can count letters and they know they can count letters.)

Many people in this thread have made claims about limitations of frontier models, but I'm the only one who has shared a conversation with one. Everyone else is either sharing conversations of smaller models making mistakes, or they're making claims about frontier models but not linking to examples of them falling over. If frontier models were so easily fooled, you'd think someone would link to a conversation showing that.

Why look at the rate of improvement of free models when you can look at token pricing? Back in 2023, GPT-3.5 cost around $20 per million tokens. Astra costs half that.

The worry is not that smaller free models will replace people's jobs. The worry is that future models will. We are talking about the capabilities of frontier models because those put a lower bound on the capabilities of future models. Extrapolating from smaller models is a waste of time, as you can interact with the frontier model to figure out its capabilities and limitations.

Also the timestamps on the shared conversations show that you asked Gemini 10 hours after ChatGPT, which means you asked it after your comment claiming you asked both models.

reply
I didn't save the original query so I asked again this morning and ended up getting the same response -- points for consistency, though it might have been more reassuring with the correct answer.

I don't think my point is really landing so I'll try once more and then give up.

Let's say frontier models today have no problem counting letters, I never really disputed that but only asked about it. It seems based on the other replies in this thread, it's a bit of a "who you ask" kind of thing, but let's grant that they have no issues with it now.

The first version of chatgpt was released 4 years ago next month. Which is not quite 5 years but close. In that time we've just barely managed to get spelling down. If we extrapolate that rate of progress forward 5 more years, are you still afraid for your job?

I think we're all more likely to lose our jobs from a downturn in the economy caused by the capex/debt bubble bursting than being made redundant by AI. (And the continual pricing reductions only seem to make this result more likely.) Hopefully neither happens and in 5 years we'll all still be gainfully employed.

reply
Solve the Collatz conjecture in the next five years? If humans publish significant advances during that time, and A.I. copies it, then yes. Otherwise, I'd definitely bet money it won't happen. I'll give you 10,000 brownie points if I'm wrong.
reply
They still have issues with problems like this actually, and I use all the frontier models from all the major labs, so it's not solved.
reply
I'd love to see some examples of frontier models getting letter counting wrong. Can you share some?
reply
> out of date by several years

This is delirious exaggeration. The problem has not even been widely recognized for several years. Fable reported "two rs in raspberry" to me as recently as August. There is some randomness, it's hard to predict which words will trip up the machine, and I haven't been able to do it at all since August. But it was absolutely happening until very recently, and probably still is.

reply
But doesn't that just amount to labs intervening to teach the models to use a particular strategy to mask this one very obvious marker of the difference between their intelligence and biological intelligence? (And similar surface issues like using tool calls / reasoning for arithmetic, even though humans writing on the internet don't typically break show their work for multiplying two numbers)

The deeper architectural difference is still there, which manifests whenever you try to get the models to apply known techniques to modalities and problems outside their training data.

reply
Chain-of-thought reasoning was added for general purposes, not to fix letter counting specifically. It just happens to solve that problem in addition to many others.

You're in the discussion section of a post about OpenAI releasing hundreds of novel mathematical proofs, and you're claiming that AIs can't apply known techniques to modalities & problems outside their training data? I'm not sure what else would convince you.

reply
LLMs are very useful, I use them every day as a software engineer to solve problems and search for information represented within the data available to them. But they are a specific type of intelligence, with many advantages and disadvantages vs human intelligence and it's not clear that just scaling or tweaking them without a theoretical, architectural change will make them more generally intelligent than humans (despite all US AI companies promising exactly that).

They are fundamentally based in language, and achieving deeper models of the world through language alone is deeply inefficient compared to the way humans model the world for years without any language at all. They do not learn at inference time. They don't have semantic understanding of the difference between their own output and other sources. etc etc.

That depth is the key for me. Of course they are capable of producing novel sentences that aren't in their training data, but the depth of that novelty is basically within the bounds of language itself. They are capable of more serious depth and more abstract reasoning than that, but I have experienced limits, which it then tries to surpass with tools to convert things it can't understand back into language (unit tests, LEAN) upon which it is trained.

Because I'm not an AI booster, my account is limited to 5 comments a day. So this is the last reply I'll be able to make today, if you want to continue the conversation we'll have to wait for tomorrow.

reply
Also think it depends on language, literally asked 2min ago from chatGPT (no login so maybe it's a shittier model?)

> Hur många 'r' I abborre, använd inte web search? Det finns 3 r i abborre.

And I explicitly had to say not to search the web, because that's what it did by default, to count letters in a word...

reply
The free models for ChatGPT, especially without login, do very little reasoning. You should at least log in to set any level of reasoning above Instant, which uses virtually none.
reply
There is no such thing as reasoning in models. Any "reasoning" is invented afterwards.
reply
I was trialing MiMo-V2.6-Pro recently due to its high benchmark scores, and it argued that substring matching the names of audio codecs in a search field was a mistake because "a user searching for 'aac' would get unwanted results for 'alac'." Which isn't exactly counting letters per se, but there are still weird issues with understanding words as strings rather than as tokens.
reply
LLMs already hit a wall. Now it 80% of marketing hype and 20% of retooling and benchmaxing.
reply
It has limitations for sure, I just don't expect those limitations to last. What probability would you put on the limitations being overcome in the next 5 years?
reply