There's the potential for an inverse LLM play. In contrast with LLMs, all that seems to matter is the model, and the products the labs build around the models are all really samey and boring; the same left panel list of agents, main view agent conversation, right hand extra context, and we're now in the era of everyone creating the same cutesey furry friend on top of all this tech.
This is a very important insight. And it applies to LLMs as well. Very few people were impressed with the capabilities of GPT 3, it was mostly a techie novelty
But then when they added chat on top of gpt 3.5, all of a sudden it was a huge hit. Sure there were improvements in the model from 3 to 3.5, but the biggest impact was from the chat experience
Conversely, when they created Eliza, a basic chatbot more than 50 years ago, people even got addicted to it, despite it’s ai model being something super rudimentary and basic compared to what we have now. The model capabilities didn’t matter as much as the experience the chat created
Why?
But even if others surpass them and make better solutions, the fact that nobody was able to see it before typesafe is a testament of what they might be able to come up with next.
I sound like a fanboy but I swear I 'm unaffiliated with typesafe. I was building my own version of this way before they announced JEV (mine was ALE and it was mentioned here on HN for a bit), in use for VR gaming (so one can give commands to NPCs with voice and supports multiple commands in sequence in a single pass), but I missed the "killer usecase" of being a new primitive, like everyone else.
TLDR: There's a lot of value in thinking ahead and seeing the future. The clones are nice and exciting but they give me "I could have built this first, yes but you didn't" vibes. I hope they manage to keep it up and push the space forward again
Counterpoint: there are a lot of one-hit wonders, and they vastly outnumber the idea-factory people. This is not to minimize those people, a single idea can be very successful (see Zuckerberg), but it doesn't mean your subsequent ideas will also be great (see Zuckerberg)
Here if anyone is interested: https://pantel.is/projects/ai-gaming-companion/
Yes, this generates probabilities over a set of given options instead of the whole token vocabulary.
I don't see what's so exciting about it, compared to regular LLMs which support structured output options in the API.
I keep evaluating Jev for my product because of the hype, but the reality is that for my tasks, Luna is more accurate, only a little more expensive, and the latency doesn’t matter. I’d rather spend the extra $50 / month or whatever than have to shoehorn in another API and provider, and also lose the ability to change reasoning level and get reasoning summaries for eval purposes.
This all sounds intelligent and likely, and yet we can come up with countless counter examples where the first mover is not the big winner, and nobody cares about who did it first. There is typically way more value in nailing the execution of a big idea someone else came up with, rather than "seeing the future".
Precisely this. Should be top comment.
Also, the model moat is understated as training data for these purposes also accrues to the winner, which due to the first mover advantage as well as the distribution advantage you speak of, is typesafe. In contrast to relatively open coding data. Openai anthropic also have that, but like you say its a different business.
If the word “calibration” in probability doesn’t mean anything to you, the difference isn’t very apparent, but that doesn’t mean it doesn’t exist.
https://seldon-ai.com/blog/fronter-llms-are-semantic-interpr...
I give it to them for creating the hype (good marketing), and for making a useful classifier. Not sure what they would scale rapidly though.
Jev has been around for a couple weeks. Cost and performance matter more than anything. Staying with an existing provider (the # of people choosing Jev without already having a frontier API key is probably zero?) is way easier than this.
What the hell are we even talking about.
Actually they stolen the idea from a paper.
That’s a different set of properties in terms of size, cost, and latency. Jev (apparently, not like I’ve seen its insides) is extremely cheap, extremely fast, doesn’t cost any output tokens as it speaks the output natively, can’t get the output wrong because it speaks the format natively, and (presumably based on the docs), the context is separate from the question, meaning it should be immune (or at least highly resistant) to prompt injection attacks.
It’s not just about the accuracy of the result, it’s a collection of all the properties that make Jev interesting.
Jev took years to develop, I strongly doubt that a copycat that was put together within days after Jev’s release will be able to match it on a sun of its properties. Even if fine tuned LLMs can outperform it on raw accuracy.
Also Jevs purpose isnt to become its own thing. It will get aquired in 18 months by one of Andressen Horowitz's incestuous circle of companies and everyone will make money, and the person who buys it wont necessarily care if Jev itself makes them a ton of money. They're just passing chips around the table.
"SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization"
Jev is a general-purpose thing. That is a specific-purpose thing. General-purpose thing is not the same as specific-purpose thing. What makes people think these are the same thing? I don't get it.
Copypasta:
• A reinforcement learning architecture specifically designed for sales conversation analysis and conversion prediction
• A synthetic data generation pipeline leveraging GPT-4O to create diverse and realistic sales conversations
• Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific features
• A meta-learning approach enabling the system to express confidence in its predictions based on conversation similarity to training data
• Integration mechanisms providing real-time guidance within existing sales platforms
• Extensive comparative evaluation demonstrating significant performance improvements over LLM-based approaches
Jev is basically the embeddings side of an LLM. Yes, it's a good idea, but the moat is non-existent.
Literally misses the point of Jev, which you don't need to fine-tune to get accuracy nor - and no other model has this - some sort of out of sample calibration
But why? If the simplest way to achieve Jev's capabilities (accuracy, cost, latency) is by fine tuning a small model, what makes you think that this isn't exactly what Typesafe did?
And even if they did something different - what makes you think it was a good idea in the first place, given how easy their results were replicated without any "secret sauce"?
Plenty has been said about this claim. If you're still falling for this, I feel sorry for you.
If you remove wheels from your car, your car will get a "no speeding ticket" property, and yet there's nothing exciting about it.
> You can’t get that with a fine tuned LLM
Of course you can. All these claims are nothing but marketing.
The only difference I am aware of is that probabilities are better calibrated with these decision models compared to regular LLMs which can output hallucinated numbers where your schema allows a number.
Fortunately Jev is cheap, so I dont think it matters too much, but I think its robbing people of the opportunity to learn and implement this themselves.
Also, I dont really want 3 companies responsible for censorship/classification.
Have you heard that from a different source than OpenAI? From what I'd heard other models haven't gotten close, and the open source ones are like running gemma4 E2B against Opus 5.5- sure, the API calls go in and are returned the same but the quality isn't close.
no, you can't, and it's unclear why you would think this.
*if you have a sufficiently sized and quality dataset for the specific classifications you're targeting
So they clearly have a product, a strategy around it and perhaps the compliance scaffolding (SOC2 Type II etc) that may be needed before actually being able to charge money for it. They also have the right brand names associated with the founding team. Execution, so far, seems good enough to create a splash, at least.
As an investor, the question(s) to ask is (in my view): "How do they make money? Will that way to make money survive?". The answer to the first: selling input tokens and perhaps subscriptions/credits eventually.The answer to the second: "Yes, but with the risk of unit revenues declining faster than their unit costs". How can they mitigate this problem: by being big (scale / mindshare etc) so that their unit costs (including for customer acq) fall faster than their unit revenues will - I believe that is the question most AI companies are trying to tackle these days. Any new competitor will have to tackle basic fixed costs (of getting started) first before even getting to the stage of having the luxury of worrying about unit-economics.
So yes, they might eventually be competed away but whoever is in their team is trying hard to make a useful product/ecosystem and that should be applauded, not ridiculed with "it's all marketing". This is way more than a simple github/huggingface-based open-source replica solution can hope to achieve without institutional backing (either big-tech or system-integrators).
What should rightly be questioned, of course, are the valuations the VCs are providing to them in hopes of passing this hot potato to a willing buyer (say a hardware maker like NVidia) - the incentives there are very well defined and depend very much on perceived TAM (which lately is on very shaky ground given how far token pricing has fallen causing, among other things, OpenAI to "miss" on the market's expectations for annualized revenues, even before they're listed!) [1]
[1]: https://www.ft.com/content/b66a9858-f8fb-46cb-b506-44bfe26fc...
Here's a hint: confidence is not generated by a model.
But confidence value is just a function applied to probabilities. It is not coming from the model, and it carries no additional information.
It is documented btw, and yet you will see plenty of claims that Jev is better than LLM because it returns both.
Do you think though that Luna being a model post-trained for chat produces over-confidence in logprobs?
Typesafe claims that Jev is calibrated, but there are plenty of examples where it completely fails (predicting die roll being the most obvious one).
Unfortunately calibration is hard to benchmark.
If you instead give it a list of probability for each number and ask it whats the probability of each number, the result will be accurate.
Did i hear that correctly? In order for Jev to be accurate you have to give it the answer before asking for the answer?
(btw this is exactly how Jev is playing games).
But, for OpenAI this is not a primary business, for open source models as well, so they will not be chasing the market and customers to buy their product and promise them to maintain it.
TypeSafe will do all this, they will try to understand your use cases and then solve your pain point, while others are providing raw material.
Guess the (investment) market has spoken.
The Microsoft article GP linked even shows the MS model having 95ms latency.
Even on price Jev being matched (the same MS model is "Input tokens cost $0.042 USD per million tokens. Output tokens are free.", same as Jev).
Cloudflare's Clef-flash model is actually slightly cheaper: "$0.038 in / $0 out per 1M" too.
Jev is being matched or exceeded in performance and price within a month of them going public.
such an arrangement can end up beneficial to the VC firm
Jev is 42$/B but OpenAI is 100$/B token.
By all means, become an A16Z LP.
To run the SDK examples below, use these OpenAI SDK versions or later: Python 3.26.0,
I thought Pythin 3.15.0 just came out, 3.26.0 must be really far off?