Caching is the simplest one to understand, cloud providers often reach a 90% cache hit rate, so hosting the same request locally on the exact same model on the same hardware is often way less efficient than on the cloud where a group of users generates a healthy cache.
The benefits of scale are on the token generation side, you can batch rounds and generate tokens for multiple conversations per pass instead of just one token per pass.
It seems damaging since most folks (who lack insider knowledge) will naturally wonder if it’s due to plateauing performance per $ or some other non “alignment” reason.