I would expect this only to be true for linear architectures like Mamba or Gated DeltaNet. Transformers and hybrid architectures do not have constant compute cost per token.
replyPerformant could certainly mean “higher performing” and not “quicker”.
reply