Longer context also slows token prediction proportional to the context size. If it wasn't regularly referenced then there would be no need to keep it in RAM.
Usually the pitch for more memory is "I can run a massive model/context and get my answer in a while instead of next weekend from disk".
Not necessarily for MoE
I think most people are getting 512 for running Chrome with a bunch of tabs open. /s