upvote
I recently had this discussion with an energetic junior coworker who just learned about TLA+ and thought it would solve all the problems. I think it probably did give the LLM that he was using to do the actual implementation a better starting off point, but the obvious gap remains, which is whether or not the spec (TLA) matches the implementation (elixir) which there's just not a good answer for. I encouraged him to put his TLA code in our docs folder, because as far as I'm concerned, it's just a suggestion, and then write (and understand...) some property tests about the feature.

I think there probably is some value in vibecoding TLA specs and not actually understanding the invariants yourself, but it's way oversold by the talking heads of the tech world, and the gaps need to be filled in some other way if you refuse to write your own code.

reply
These probabilistic guessing machines are pretty great for creating these formal models, e.g. TLA+, and then guessing if the implementation aligns with the spec.. In fact, it's my go-to tool for constructing soft guardrails for the model, so the design it's going to implement is logically sound. Same as for people: it's easier to make something working when you have a spec that is working.

Of course, it still allows the risk that you don't actually get to understand it.

reply
> People can't escape the need to actually understand the things they are building.

While on the one hand, you do need some kind of grounding in human specification for what to build and what good looks like, any particular defect humans can find should be findable via software.

reply
I wonder if asking an LLM to model their implementation in TLA+ first would improve their implementations.
reply
It does somewhat, or at least used to for a few months. Nowadays it kinda seems that the models have internalized something like TLA and are thinking in it in parallel to thinking in the language they’re writing, so it doesn’t help as much. This is all educated guesses from me, I’ve been telling models to do TLA back in the stone age around February and stopped seeing improvements when telling them to start with specs.

I’m however pretty sure that if you push a good model hard enough on a code base complex enough it’ll find stuff it wouldn’t have otherwise, the Specula folks have some experience with this.

https://github.com/specula-org/Specula

reply
A TLA+ spec defines both a model and properties (global invariants). How do you know that the properties the LLM specifies are the ones you care about?
reply
How do managers build software?

The fact is when LLMs get good enough you WILL be able to build software without reading/understanding the code.

Whether or not you think we are already at that point is kind of an unimportant detail.

I would say we are quite close, depending on the type of software you are building.

reply