1. User request understanding: converting natural language to a more rigorous form, in my case Datalog.
2. Query result interpretation: converting facts, derived facts back into natural language.
Between those terminals should be pure mechanical reasoning over an ontology of some sort.
That connects to another principle I've been thinking about, which I call weathering: useful reasoning should leave durable residue in the system. If an LLM had to infer a relation, mapping, rule, or abstraction once, the next similar request shouldn't require it to rediscover the same thing from scratch.
In fact with continuous use, the system should require less and less probabilistic intelligence over time.
I ran into the very same problem of the LLM forgetting that we ruled out a conclusion that was verified not to be the cause as it came up further in the conversation history while I was exploring possibilities.
I had to keep reminding we ruled out that conclusion prior.. I just carried on with having the LLM capture some of the supporting sources of other people experiencing the same problem and kept having to refine those sources because it was focused only on summaries, but eventually i got the sources to a point where they were good enough hypothesis that we could formulate a better conclusion on what the potential cause was.
(See also Cyc: https://en.wikipedia.org/wiki/Cyc)
I think approaches like this are going to be (or maybe already are?) the basis of effective grounding of LLM responses in authoritative data sources. It should be possible to pinpoint any error to an incorrect traversal or an incorrect "fact." This would work best for concrete, unambiguous facts, however; fuzzy, ambiguous or opinion-based information will probably remain the purview of LLMs.
But IRL it's too vague. The exploit hunting is a better use case.