upvote
> a tiny group is watching the curve go vertical

Caveat: that same tiny group is employed by the AI vendors, meaning it's in their financial best interest to make it sound like the curve is going vertical.

reply
Also, part of "seeing the curve go vertical" is that the projects that show those results are rather expensive. If you work at one of the AI companies or certain well funded customer companies then you can pay for the tokens. But, like it would make no sense for me personally to spin up a bunch of agents to try to write a browser or solve math problems.
reply
Down is a perfectly valid direction for a vertical line when not given a ± in the vector.

Considering the investment in AI, the lack of moat, and the increased inability of any of the big players to come even close to profitability (with OpenAI already breaking the "ads" emergency glass option)... perhaps he meant a tiny group is already seeing the line crater

reply
It can be both and it is.
reply
Another caveat: many of that same group seem to have a shared delusion that they’re birthing a super intelligence, and those are the same ones claiming the vertical curve.
reply
> and those are the same ones claiming the vertical curve.

Do you still write code by hand?

reply
I work for a typical software product company, and along with most of my colleagues, I clearly see that the curve is so steep that our software development process has already changed tremendously and will be different next year, and probably completely different in the following years. The revolution is real, undeniable, and the old days are gone. Some companies adapt to changes more slowly, some faster. LLMs, agents and harnesses are just a part of the bigger picture.
reply
I hate to say it but this is cope. I believed this in the past but it's over.

AI models can dismantle billion dollar industries. They can reverse engineer Adobe and Microsoft products that once were their moats and titans of their industry.

Why do you think they'd need to lie?

reply
Because if they can, why haven't they?

LLMs are great and will get better but we have not seen an open source office suite built by LLMs come online yet, and even if there was one I'm pretty sure few would adopt it, cause that's not the only moat, LibreOffice has been around for decades.

reply
Have you not seen WordCraft[0] or EffectCraft[1]?

Sure they didn't build it from scratch but what I'm saying is that there's no moats anymore. If you build something, they will tear it down to its composite pieces and assimilate it.

[0] https://getartcraft.com/apps/wordcraft

[1] https://getartcraft.com/apps/effectcraft

reply
Well, in the case of Microsoft software I'd assume it's because Microsoft owns significant stakes in both OpenAI and Anthropic.
reply
The curve go vertical for what, precisely? For all the chest-puffing and ominous and cryptic comments from 'insiders' about these incredible capabilities every time they put something public it turns out to be a pale shadow of what was trumpeted. These are powerful useful tools but the quasi-cultish behaviour around them is getting old.
reply
Congratulations, you've parroted something said by someone else.
reply
I use models all day everyday, have unlimited access to all models. The curve is not going "verticle". I have all the workflows and meta agentic tooling, im not holding it wrong. Its bad, not everything is a 20th percentile problem.

There is in fact no indication of this, not evem the precious benchmaxxed benchmarks ya'll love to reference.

There is however a exponential curve of slop, and an ever increasing number of people who's minds are completely captured by these things.

reply
As someone who has done software verification professionally for many years the last 6 months or so have looked extremely vertical. The robots are better proof authors than I probably ever could be even if I dedicated the rest of my days to the practice, and projects that once would have taken months now take a day or two.
reply
Then you'd know well that 2hr of LLM code generation can easily be about 4-8hrs of review, and that review can be brutal.

I'm not arguing that they cant write code, or write a proof. Its just not written or designed well and is absolutely brutal and soul crushing to work with. Look at these proofs they're producing also, they're millions of lines of Lean that are impossible to reason about.

The way we're using the term 'verticle' to describe a curve means we're not being honest about this. This curve can actually be plotted, you can go look at the curve. It is not in fact 'verticle'. Each model release is climbing single digits on benchmarks it was overfit for.

reply
You don't need to review proof code.

In the last 9 months or so llms have gone from just another useful proof tactic (like grind or sledgehammer or sat solvers) to being so good at writing proofs that I don't even bother to try myself anymore.

reply
I believe this, but it is also a unique case where the pitfalls of LLMs (producing weird errors that a human wouldn't) are zeroed out. Since you have a proof checker that tells you if the LLM did it right.
reply
I'm pretty convinced most serious software will have some kind of proof system inside within the next few years.
reply