I had similar thoughts recently, that now that machines are good at writing proofs, this could help with their reliability in software development.
Then I had a funny incident where an LLM implemented a feature completely backwards. Plenty of tests were supplied which demonstrated that the completely broken feature was correctly implemented.
I realized that formal verification would not have helped here, if I had left the task to the machine. It would simply have written a mathematical proof of the correctness of the incorrect feature!
Apparently this is an issue for humans as well, called the "spec gap" or something like that.