The current benchmark suites that frontier AI labs use are probably a good fit, e.g.
https://z.ai/blog/glm-5.3#:~:text=Performance%20across%20com...
https://www.kimi.ai/ai-models/kimi-k3#:~:text=Performance%20...
https://www.anthropic.com/news/claude-opus-5
https://openai.com/index/gpt-5-6/
But guessing from your current benchmarks, I assume that you are severely compute-constrained. What is your time budget?