FrontierCode is prob the closest. [1] It's closed source (so no direct benchmaxxing), and it was calibrated by 20+ open source maintainers. It shows Opus 5 (medium), beating out the other reasoning levels by a large margin. i.e. Opus 5 w/ higher reasoning levels actually reduces performance. [2]
However, you'll have to gauge for yourself how closely their tasks resemble your tasks.
1/ https://cognition.com/blog/frontier-code
2/ https://cognition.com/frontiercode