upvote
Yes, it's easy. The thing people are missing (especially those believing AI is a "dead end" and "not transformative") is that the field has been advancing so fast in the past few years, that there's lots of such unexplored avenues, unpicked low-hanging fruits, that everyone just raced past. We've barely begun exploring the capabilities ML brought us - patterns, applications, and architectures.

Now that we're hitting against the hardware supply limits of global economy, I expect more people to go back and revisit the things left along the way in the mad rush to "just throw more compute at it / make a bigger model" - and thus many more cases like Jev to show up in the next few years.

reply
The basics are pretty simple. And depending on what your specific need is, the model can be really really basic, fast and super effective (ie. run on a mobile device and process thousands of requests in <100ms)

I've been playing with this for the last year or so. Started with a personal email classifier, also did benchmarks with some public datasets, then created a couple classifiers that could play Doom, and now I've been trying out some other experiments, like a request proxy/router to automatically choose a classifier and fallback to LLM to handle unseen requests

Jev did a great job at creating hype, but also at shaping the concept and space of "decision engine" or "decision model". People were already doing this with LLMs, which is very inefficient for most tasks like that, and the Jev guys figured there was a market there. It seems like they were right, and now there's a rush to flood the space, taking advantage of the hype window

reply
There is a bunch of stuff to tease apart.

In general, training a general purpose classifier is something lots of people have worked on for a long time. Large Transformer models themselves are typically "generalists" already, so structured generation and constrained decoding have given you the ability to use an LLM as a general classifier for years. It's an incredibly common pattern for working with LLM judges or any sort of branched decision making workflow.

A lot of people who are a bit less familiar with the field saw the hype around Jev and presumed that the reason it was so exciting was that it was a fundamentally new interface for working with an LLM. And that additional excitement drove even more attention to Jev. But fundamentally, TypeSafe's announcement was that they found a particular architecture/training paradigm that resulted in a model for this particular interface that had incredible accuracy, very low latency, and for which they could offer inference at a super low cost.

I've not kept up with the flood of Jev clones that have been released, but I think this is just typical for any new component in deep learning that gets popular. There are an absurd number of open source autoregressive LLMs and fine tunes you can use. The thing that makes one more popular than the other is typically the general performance of the individual model.

But training a model for this purpose, or emulating the procedures described in Jev's papers, isn't something that would be beyond the capabilities of any lab. It's not an entirely alien architecture or approach.

The bigger question for TypeSafe as a company would be if other teams are producing Jev-like models that win on performance or cost. Like I said, I haven't followed the reports super closely, so no idea if that's the case or not.

reply
> The bigger question for TypeSafe as a company would be if other teams are producing Jev-like models that win on performance or cost. Like I said, I haven't followed the reports super closely, so no idea if that's the case or not.

If CF's benchmark is representative and sufficient, Clef outperforms Jev!

Models by themselves don't guarantee market capture. Rather, its how they integrate. I think a lot of folks are burned by the closed nature of many models.

reply
Live coding Jev from Scratch | Understanding Qwen architecture

https://www.youtube.com/watch?v=AzxoU7kxjig

reply
Yes it's easy for an established shop, all they need to do is to tweak the post-training workflow. "Decision model" is the same kind of marketing as "LRM" attempted by OpenAI when RL CoT was new (to hyped up crowd). It's still fundamentally a classifier used for "decision making", games and RP were using generalist models and constrained outputs to do what the DOOM demo does for years.
reply
The interesting part is also the easy part. The model and architecture are not hard for an experienced machine learning engineer to build.

The hard part is the data and evaluation. Sure, it’s not that hard to build a fast model with good predictive power. But fast at doing what? You probably don’t care about classifying whether a hotdog is a sandwich (which is the Jev demo).

reply
You can make a basic one in minutes based on existing open-source models.

Latency won't be that good, but could still work similarly. Simply force the structured output of a LLM to the given schema.

Probably also easy to train because we can use stronget LLMs to generate input/output data, or even synthetic data is easy to generate.

It's not really a new technology, it's more like a new use-case.

reply
What even are these new "decision models?" Take an existing LLM, feed it a prompt, force it to pick a choice; decode is 1 token (or rather, the whole logit set for only that last token; token implies selecting one logit) so you made a choice. That's it?
reply
Yes, although you probably want to calibrate your model if you want the probabilities to actually be meaningful.
reply
Yes but optimized specifically for the purpose. Using that for "decision making" is also not a new use case, but turned out to be new to many people. Which is great, I hope they make something cool with it!
reply
Yes it's very easy if you have fairly basic ML knowledge.
reply