This was definitely truer with older models but isn't necessarily the case now.
They frequently do other things apart from navel-gazing that take a lot of time but get good results, like spinning up subagents to solve some hairy task in a loop.
Stopping/distrusting long-running AIs is a habit I've had to unlearn myself.
I frequently get good results from a 30+ minute Fable session, when I've asked it to do something complex (e.g. run the QA tester in a loop and eliminate all crashes, one commit per crash fixed)
The main failure point for me with long running Fable sessions now is just that it might hit a safety guardrail and downgrade to Opus midway.
I've found 6 minutes or so the sweet spot for upper bound with 5.6 Sol.
And it sounds like the OPs query above requires scanning throughout a large portion of the codebase, which will inherently consume a large number of input tokens. No locality to it.