For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc.
These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k.
In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.
Open weights models giving a distant salute from afar
If you want to make it about 99% of real world companies, they are all on Gemini or Copilot anyway, nobody is going through legal and procurement to get models from dubious silicon valley startups when you have relations with Microsoft or Google or Amazon from ages because some benchmark is showing some minor digit benefit when vibe coding GTA 6.
You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc
So, it is might be even worse.
I use Luna for this day in and out and its excellent - if Haiku is that much better I will be changing things up.
"less context is better and if you can't get stuff done with less yur bad" is the worst argument ever.
it might be pure luxury to your eyes, but it's great to not require the use of a special custom harness that transcribes everything into emoji and compresses everything into barcode images.
it's great to have a million token context to throw a large project into. If I need 100k just about any current gen consumer GPU in the world has very good models that I can self host for 100k context, limiting myself to 100k on someone elses machine seems to be missing a lot of the point unless the model itself is extraordinary.
Neither encode nor decode are linear in compute, so providers need to price for average expected length.
This is just getting closer to the true cost of generating tokens.
For this application 100K token input is plenty.
Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.
For other tasks like summaries (another suggested usage) it's good to see Haiku and Luna now competing against each other on cost.
I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases.
From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.
Fixed.
The benchmark in the article showed it as lower per completed task than luna, but I guess we'll find out how representative that is. Anthropic has generally been fairly honest in their benchmarking though.
With that said, the real reason to use Haiku is that it's faster than all of these models. OpenRouter is showing an average so far of 93 tokens/sec, and AA got at least 137 in each of their benchmarks. So it might be valuable for speed at lower thinking levels. (At higher thinking levels, it's likely going to take longer to produce results than Sol on low/medium.)
https://artificialanalysis.ai/models/releases/comparisons/cl...
I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.
Your vibes don't appear to be supported by facts. From the announcement:
>> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.
I'm asking to learn for a similar project, not to discount anything you're saying.
There are plenty of workflows like translations where you'd easily be under the cap.
AAI Index // Input // Output
Haiku 5.5: 43 // $0.10 // $0.50
Mimo 2.6 Pro: 46 // $0.43 // $0.87
Mimo 2.6 Flash: 38 // $0.10 // $0.28
Seems competitive to me? Plus then I don't have to manage multiple providers
And Opus 5.5 is really good.