RE: https://mastodon.world/@anttipeltola/117092344692403567
I really like the Pentium 4 as an analogy here.
The NetBurst architecture in the P4 was designed to be able to scale to 10 GHz. It initially shipped in the 1-2 GHz range. The performance was not that impressive compared to even the previous generation but, Intel told us, that didn’t matter: it would rapidly reach clock speeds that were completely impossible with the P3.
On the face of it, this wasn’t unreasonable. The Pentium Pro and Pentium 2 shared a fairly similar microarchitecture (P2 added MMX and a few other things). The Pentium Pro was introduced at 150 MHz and scaled up to 450 MHz after process shrinks.
When the Pentium 4 was in early design stages, it was completely reasonable to expect that 3-5x clock speed increases were possible.
And, technically, they were. The clock speed for a part is determined by how long it takes a signal to propagate across everything in a pipeline stage. Make the things smaller and the signal goes further.
That kept happening. The thing that stopped (completely by about 2007, but it wasn’t an abrupt thing) was Dennard Scaling. Beyond a certain point, making the transistors smaller meant you had a lot more leakage current and so the power went up. And power was already polynomial in terms of clock speed.
All this meant that, yes, in theory, you could run a Pentium 4 at 10 GHz, as long as you didn’t mind all the fire. Someone did eventually run one at this speed with liquid nitrogen cooling, but that wasn’t appropriate for a consumer device. The range topped out at 4 GHz, and those parts had infeasible high cooling and power requirements (and were outperformed by other designs at lower speeds).
The entire ‘AI’ bubble is predicated on the same assumptions of scaling of something foundational that the people building the end systems are not really paying attention to.
LLMs have become technology that has failed to deliver on its promises the most since Pentium 4.
Failure of Pentium 4 however didn't put a loaded gun pointing at the head of the working class as failed CPU microarchitecture is incapable of destroying the economy as it didn't have USD 2 trillion of CapEx depending on its success.
Same 10 % hallucination rate despite how much we throw data and compute on these transformer models. Holy fuck we're in trouble.