The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years.
Cerebras will win in terms of approach.
It's 1998: hey, I can drastically speed up your web service, let's etch it right to silicon.
Qwwen3.5 122b was released 6 months ago and is still best in class overall 100-140 B param model.
....
:T
Considering the 8B model uses 53 billion transistors, that's 6.625 transistors per parameter.
Assuming they can get it down to 3 (somehow), that's still 300 transistors, or 5.565 RX 9070s.
https://www.techpowerup.com/gpu-specs/radeon-rx-9070.c4250
You're looking at
1) waiting for another 3-5 generations of transistor improvements before it can fit into a single conventional chip, or
2) another generation before getting a monster of a chip (1000+ mm^2), and prices for flawless etching scale quadraticly (likely $1000+ for manufacturing costs alone).
Could happen, but it's a long shot for a market that could be satiated by specialized accelerators.
Slightly disagree. It really depends on the price-point at which they can do that etching. ~1k usd / ~30B model in a hdd-sized case that fits on your desk? I'd buy one right now, even knowing that I'm "stuck" with whatever model of the day is.
> From the moment a previously unseen model is received, it can be realized in hardware in only two months ( https://taalas.com/the-path-to-ubiquitous-ai/ )
I distinctly remember 32-bit/33 MHz PCI accelerator cards for SSL being a real thing (for use on OpenBSD or FreeBSD), in an era when something like a single core 700 MHz Pentium 3 1U system was a relatively powerful individual bare metal httpd box.
http://www.aster.si/partnerji/compaq/atalla/axl200.html
The CPU load of doing a lot of SSL purely in software was a problem in terms of scaling things up, so this was one attempt at a (very short lived) solution. Note that this predated TLS1.0.
Yes, please!