That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user.
Regardless what you call it, it's a divergence between what the power user wants and what the model developers want, I think.