It schedules to these transparently for you, that's known as superscalar execution. To maximize occupation, out-of-order execution and simultaneous multithreading are used.
There were a few new instructions, too, mostly closing holes. You could left or right shift with a constant, while the 8086 had only 1 or the CX register. I think mul also gained a constant.
The fact that Intel released a CPU that could not be put in a PC probably indicates how low they estimated the survivability of the PC.
In the x86 microarchitectures superscalar came in Pentium and OoO got introduced in Pentium Pro.
(Superscalar is just having >1 pipelines, which at its introduction meant needing to manually schedule your code very carefully to take advantage of it absent the OoO execution. For example the frequently posted-about Doom optimizations and talk of the u and v pipes are about this. The scheduling didn't happen transparently in early superscalars, at best the cpu automatically stalled, and some archs (eg MIPS, i860, TI C3x) even visibly punted hardware detection of pipeline hazards and required the code to just not go there, see load delay slots and branch delay slots. )
Vs. Ken's Blog says "up to 100 times", and Wikipedia gives a lower estimate.
Theory: Your 100X experience compared Intel's "exact emulation" code (noted in the article) with native x87. That emulation would have to cover the myriad x87 oddities and corner cases which Ken describes. Vs. Ken's & Wikipedia's are comparing x87 to various "good enough" 8088 floating point libraries - so naturally much faster than Intel's exact code.
(And yes, speed might have been a low priority for the team writing Intel's emulator.)