upvote
I'm thinking super long context length or something to that effect.

I can imagine schemes for instance where context is compressed into chunks and then chunks that are ranked highly relevant for the token are decompressed. Which would sort of be between a long context and a memory retrieval scheme...

reply
That's how I interpreted it, but now I'm wondering if they mean "this model gives up far less often"..
reply