If you had some use case with very small output sequences they could be interesting to try. I think dropping down to a 9B-class model would produce better results for most cases.
Show HN: Forge – Guardrails take an 8B model from 53% to 99% on agentic tasks