upvote
> so 17 million 8080 equivalents

At, ballpark, 1000 times the clock frequency. It’s a pity we do not how to connect such a large number of tiny cores in a way so that it can perform meaningful work. One hurdle is that, even ignoring the insane amount of connections needed, it would not be possible to connect each of those cores to each other one because their address spaces are so tiny.

reply
Cerebras built the wafer scale engine so I'm not sure what you're pitying.
reply
> It’s a pity we do not how to connect such a large number of tiny cores in a way so that it can perform meaningful work

Every so often someone reinvents the transputer, and it turns out not to be quite as good as mainstream multi-core CPUs, but the wheel must turn.

reply
yeah, i have the notion that a grid would be good enough
reply
The ARM 1 only had 25,000 or so.
reply
wonders how ARM did that - the MIPS R2000 came in at ~115k and the 80286 ~134k

anyway - say you want 50% of the transistors for on chip RAM these days, then thats

100 billion / 50 thousand = 2 million ARM 1s

clocking at say 3GHz

6e15 MIPS = 6 peta MIPS

so if you could run code on it, you would get a 10,000x speed up ;-)

reply
GREAT classic question (with a well-defined answer)

In general, the reason why modern CPUs are so complex is because the gap in performance between CPU and memory has grown massively over time.

In the old days, something like a 6502 was running nearly synchronously with RAM.

That gap has grown massively over time; a modern CPU is orders of magnitude faster than RAM. So they have to jump through a lot of loops to avoid simply idling 99.99% of the time while they wait for some new data or instructions from RAM. On-CPU cache memory is one answer. Branch prediction and speculative execution are others. As you may imagine, speculatively executing code based on branch prediction is very complex because you must roll back any side effects from that execution if your prediction turns out wrong.

Example:

   # assume `i` is a value stored in main memory
   if i == 42
      j += 1
      k -= 1 
      l = 666
   else
      z = 123
      q = 5879873

Waiting for `i` to arrive from main memory might take thousands of CPU cycles. So instead we will execute one, and possibly both of those branches. But we'll need to undo those side effects if turns out we executed something with an invalid prediction. It's complex, and messy, but still better than sitting around doing nothing for thousands of cycles.

That's why we can't just take an R2000 and scale it up to 3ghz. I mean, we could, but it wouldn't work very well unless we also had low-latency 3ghz main memory to pair with it.

reply
> wonders how ARM did that

Acorn co-founder Hermann Hauser famously joked that he gave the original ARM microchip design team two distinct advantages: no time and no resources. With no money for a large crew or complex hardware, the tiny core team had to keep the processor design exceptionally simple, which ultimately birthed the revolutionary RISC architecture.

It also had to run in simulation on a BBC micro… Probably with a second processor via the Tube interface, but still.

reply
ARM1 has no cache memory, no hardware multiply and division, no MMU and no cache.
reply
Yeah — even ARM2 has no divide, as far as I recall from my days with an A310.
reply
This is understandable, considering that the M5 is the Ultimate Computer, powered by Dr. Richard Daystrom's very own engrams
reply
But the M5 is clearly not safe to use in Enterprise¹ computing environments.

¹ USS Enterprise, that is.

reply
excellent point
reply
Yup and Moore's law about transistors doubling every 18 to 24 months did basically hold, which is even more impressive.

Foreseen it was.

reply
Nitpicking, but since the M5 includes cache and GPU a more realistic comparision would probably be to count all transistors in an entire 8-bit home computer system against a modern integrated GCPU (e.g. from googling around a C64 might have around 0.75 million transistors, most in the memory chips).

And in a way, those old 8-bit home computers also worked like a single integrated 'super-chip', because the whole system was driven by a single clock and entirely 'hard realtime', while modern computers are much more asynchronous (and I guess this asynchronous design is what enabled most performance improvements that go beyond pure transistor count scaling).

reply