upvote
Gemini uses MoE and context caching, which is a similar approach.

You are not really accessing the biggest frontier model every time, and you're not really doing an end-to-end LLM request on each prompt.

I would go so far to say frontier models have peaked and improvements from here come from clever (or very elaborate) harnessing. "LLLMHs" - Large Large Language Model Harnessing !

reply
This is an odd comment: the project is right there for you to use, so just try it and see if it holds up to the claims? Then you can comment about the fact that it either doesn't hold up, with numbers to back that up, or on how awesome it is because it works =)
reply