That's not an emerging practice, it's a tested strategy that is these days only used as a last resort by those desperate to fit a model in memory. Some models do better than others, but generally the model quality suffers greatly under those conditions.
But it's a moot point, because for local inference on consumer hardware, the MoE is so much faster.