upvote
Agree but it's helpful to remember how we were personally benchmarking. I remember people saying stuff like "haha I asked gpt4 for xyz function and the typescript didn't even compile". We're so far beyond that now, we just adapt quickly.
reply
Oops, I might have been misremembering then. Maybe I meant 4 to 5
reply
No, no, I also remember 3.5 -> 4 and the general sentiment was that it was underwhelming. I guess we all expected absolute miracles from the models. I think our expectations sobered up a little since then.
reply
4.1 was the first decent 4-series model. It was significantly better than previous generations at tool calling if I'm recalling correctly.
reply
Yeah 5 was very underwhelming.
reply
The couldn't even get the bar chart right, iirc. [0]

[0] https://www.reddit.com/r/singularity/comments/1mk8tm8/gpt5_c...

reply
Yeah, GPT4 was one-shotting utilities that GPT3 Davinci couldn't. So, I'd have my limited tokens on GPT4 crank out the initial program before iterating with my abundant, GPT3 tokens.
reply
GPT 4 to 5.5 felt about the same as 3.5 to 4 to me.
reply