upvote
As someone who has done software verification professionally for many years the last 6 months or so have looked extremely vertical. The robots are better proof authors than I probably ever could be even if I dedicated the rest of my days to the practice, and projects that once would have taken months now take a day or two.
reply
Then you'd know well that 2hr of LLM code generation can easily be about 4-8hrs of review, and that review can be brutal.

I'm not arguing that they cant write code, or write a proof. Its just not written or designed well and is absolutely brutal and soul crushing to work with. Look at these proofs they're producing also, they're millions of lines of Lean that are impossible to reason about.

The way we're using the term 'verticle' to describe a curve means we're not being honest about this. This curve can actually be plotted, you can go look at the curve. It is not in fact 'verticle'. Each model release is climbing single digits on benchmarks it was overfit for.

reply
You don't need to review proof code.

In the last 9 months or so llms have gone from just another useful proof tactic (like grind or sledgehammer or sat solvers) to being so good at writing proofs that I don't even bother to try myself anymore.

reply
I believe this, but it is also a unique case where the pitfalls of LLMs (producing weird errors that a human wouldn't) are zeroed out. Since you have a proof checker that tells you if the LLM did it right.
reply
I'm pretty convinced most serious software will have some kind of proof system inside within the next few years.
reply