RAG was supposed to be the way out on that, and ended up being mostly abandoned.
However, this is not something that is inherently part of models or inference engine, but part of the harness.
Harnesses are very hit and miss, and are not integrated into the stack, and I think that will have to happen eventually. Like, conceptually similar to an LLM performing a tool call that just calls itself recursively, I think this would go a long way to making LLMs more viable for being an actual product people could conceivably want.