Ken Shirriff's reverse-engineering of the Intel 8087 FPU reveals that FPTAN uses a deliberate hybrid algorithm: 16 bits computed via CORDIC, then a Padé approximant, the ratio of two polynomials, to reach the full 64-bit result. That switching point is the key insight. Because CORDIC leaves only a small residual error, the polynomial step is both fast and accurate, avoiding the lookup table bloat and iteration cost of running CORDIC all the way to 64 bits.
Shirriff's timing breakdown is concrete: for a representative input, FPTAN spends 33% of its cycles on CORDIC pseudo-division, 47% on CORDIC pseudo-multiplication, and only 15% on the polynomial approximation, with 5% overhead. The full annotated microcode listing is in the original article, which is where the real value is. Reading the comments Intel's engineers left in that microcode is a direct line into 1980 design tradeoffs.
Intel abandoned CORDIC entirely with the Pentium, a signal that the algorithm hits a hard scaling wall as precision requirements grow. The 8087 analysis shows exactly where that wall is and why the hybrid approach bought time. With SIMD long since displacing x87, this is architectural history, but the algorithmic reasoning is still current for anyone doing constrained embedded math.
[READ ORIGINAL →]