when you want a machine to reason about the prompt and generate a structured output not using an actual LLM makes no sense. I have been doing it since the first chain of thought open models became available.
perhaps there may be a way to get a Jev-type model to think for a very specific number of steps to gain control over its latency, if so that would be the next step. truncating LLM thinking like this does not work well, and its thinking isn't efficient anyway.