You can see this in the Artificial Analysis benchmarks - GPT 6.1 Sol on medium thinking scores 10 points higher than DeepSeek 4.1, but is actually cheaper per task, as it only outputs 15 million tokens instead of 250 million.
Though for longer sessions I think DS4.1 would still come out cheaper... it's hard to beat that 98% cache discount