There are no configurations even close to running something comparable to frontier model variants, they're simply far too large, but something like full precision Qwen 35b or DeepSeek 70b at 50+ t/s is well within available configuration, and potential for plenty of room for large context sizes.
I'm using Flash heavily, and I would describe it as nearly as intelligent as Sonnet-class in agentic coding, but more usable. Less world knowledge of course, and definitely a bit less intelligent; but not _that_ much.
On usability: Takes less handholding, less likely to make unsolicited refactors or whatever, and the writing style is readable.
It's not great at super-long-horizon goals as the Claude 5 models are; but if you have a good harness, you can get around that.