How is this idea generally working out in comparison with Jev? I'm curious, from what I read so far it seems like Jev is still beating this kind of thing.
But it's curious because it's not entirely clear why, from an architecture point of view for all we know that's exactly what they're doing. So it must come down to the quality of those logits, ie., model size and training details.
It seems to me that what most of these single-token-prediction projects are missing is that Jev seems to be claiming they predict well-calibrated probabilities. This is an incredibly valuable thing that LLMs simply can't deliver unless they are trained specially for it.
Jev clearly has _some_ secret sauce compared to doing the dumbest thing that could possibly work with Qwen. It's not clear how durable that advantage is against OpenAI wiring up Luna-5.6 and doing a minimum amount of tweaking, but I presume we'll know in a week or two.
https://docs.typesafe.ai/model-jaggedness/jev-1.13#generatio...
I think the real difference Jev makes is the fast parallel decode, it just seems rather bizzare how that works.
Slack is glorified IRC yet they're worth billions.
Dropbox can be trivially implemented via rsync yet they're worth billions.
God I’ old
Similar to how we upgraded computers for decades and the software bloated to fill the specs
I quite enjoyed handwriting my article to be honest.