upvote
I guess it's more performant to stuff in a bigger system prompt now that models can support larger input sizes
reply
I would expect this only to be true for linear architectures like Mamba or Gated DeltaNet. Transformers and hybrid architectures do not have constant compute cost per token.
reply
Performant could certainly mean “higher performing” and not “quicker”.
reply
It reminds me a bit of building codes and boilerplate contracts: they start out small and simple, then accrete over time in response to mishaps and exploitation of loopholes. They say the building and electrical code was written in blood.
reply