1. https://chatgpt.com/share/6ac5e4cc-02f0-83e8-8f05-99a7ea2bf9...
https://chatgpt.com/share/6ac63b6a-481c-83e9-a8fa-a13ce7402d...
I used whatever the default free model and thinking time was. If progress was really as fast and continually cheaper as some worry it is, wouldn't we expect free models by now to know (or even perform) what frontier models were capable of as much as 2 year ago?
This deep in the "comparing logs" tangent we risk missing the point. It's not what exactly frontier models are capable of at this particular point in time. But that there's entire categories of problems that seem easy to us which LLMs really struggle with. We've stumbled on several just a few replies into casual conversation. (Can they count? Can they know if they can count? Can they reproduce results? How quickly do new capabilities filter into free models? And that's just what's come up naturally, if we wanted to pick adversarial examples there's more to choose from.)
So while there's a number of difficult problems that are easy for LLMs (like bulk generating lean proofs), there are plenty of things where progress is not so impressive.
If LLMs can struggle so much with such easy problems, what hard problems have we yet to discover that they'll struggle with? The fact that no one knows, 5 years in advance, what those problems will be does not mean the chance of them is zero.
So far progress on the things LLMs are good at is fast and easy. It's like fire in a room full of oxygen. But once the low hanging fruit is gone, and the oxygen is out of the room. How fast will the fire burn through steel walls?
In my opinion it's a mistake to look at only rate of progress on one type of problem (whether it be what LLMs are good at OR what they're bad at) and assume progress on all tasks will progress at that rate indefinitely. Isn't there a saying about exponential curves, in nature, all being sigmoids eventually?
I guess we'll just have to see. I wish you good luck with your wagers.
Many people in this thread have made claims about limitations of frontier models, but I'm the only one who has shared a conversation with one. Everyone else is either sharing conversations of smaller models making mistakes, or they're making claims about frontier models but not linking to examples of them falling over. If frontier models were so easily fooled, you'd think someone would link to a conversation showing that.
Why look at the rate of improvement of free models when you can look at token pricing? Back in 2023, GPT-3.5 cost around $20 per million tokens. Astra costs half that.
The worry is not that smaller free models will replace people's jobs. The worry is that future models will. We are talking about the capabilities of frontier models because those put a lower bound on the capabilities of future models. Extrapolating from smaller models is a waste of time, as you can interact with the frontier model to figure out its capabilities and limitations.
Also the timestamps on the shared conversations show that you asked Gemini 10 hours after ChatGPT, which means you asked it after your comment claiming you asked both models.
I don't think my point is really landing so I'll try once more and then give up.
Let's say frontier models today have no problem counting letters, I never really disputed that but only asked about it. It seems based on the other replies in this thread, it's a bit of a "who you ask" kind of thing, but let's grant that they have no issues with it now.
The first version of chatgpt was released 4 years ago next month. Which is not quite 5 years but close. In that time we've just barely managed to get spelling down. If we extrapolate that rate of progress forward 5 more years, are you still afraid for your job?
I think we're all more likely to lose our jobs from a downturn in the economy caused by the capex/debt bubble bursting than being made redundant by AI. (And the continual pricing reductions only seem to make this result more likely.) Hopefully neither happens and in 5 years we'll all still be gainfully employed.