| | |

New processors roll out at ISSCC 2014

February 2014, At ISSCC in San Francisco, the traditional launch of new microprocessors was done in session 5. The session features new architectures and designs in the Power PC, ARM, Xeon and x86 solution space from IBM, Intel and AMD. Large data centers and cloud service providers are increasingly becoming the point of computation in the chain, as front end devices become thinner and thinner clients. This resulted in the processor session this year introducing new architectures that are providing increased performance and power efficiency. The largest devices have over 4B transistors and double digits on the core counts. Also increasing are cache size and thread count as well as the integration of voltage regulators, and self-monitoring of temperature & clocks. They introduced many adaptive techniques to meet power and performance design goals of these new devices.

The Power8 implementation of the PowerPC lead off the session with 3 papers. The first introduced the architecture and process of the new 12-Core Server-Class Processor
in 22nm SOI with 7.6Tb/s Off-Chip Bandwidth. The process features 15 layers of metal in the BEOL, 3Vts of thin-oxide devices and multiple thick-oxide I/O & analog devices in the 22nm SOI process. The design uses the thin and thick oxide transistors as well as optimized SRAM and eDRAM cells, in addition to 2 CAM and 31 multi-ported registers. The chip has 4.2B transistors in the 649mm2 die size. For the growing needs of Enterprise computing, the chip features an advanced memory interface with up to 8 high speed DDR channels with up to 9.6 Gb/s per lane, which supports 1.84Tbit/s inbound BW and 1.30Tbit/s outbound BW, this results in a total chip capacity up to 1TB.

Following the architecture and process paper, other IBM presenters talked about the distributed micro-regulators per-core for the Power8 microprocessor, and the wide frequency reange resonant clock that managed the 29 separate clock domains used in the Power 8 design. These designs were both architecture and process specific to provide the highest performance and not impact I/O bandwidth of the chip.

The next set of talks were on x86 / IA based microprocessors from Intel and AMD. AMD presented two papers on thier next generation x86-64 bit core, built using 28nm Bulk CMOS the is called Steamroller. Like the IBM design and presentation, the multiple papers covered the architecture, process and then a separate paper on the adaptive clocking and power saving aspects of the design. Unlike the IBM paper, rather than discuss the finished chip, they did an overview of the core. The process has 12 metal layers and 63 unique macrosand features 236M transistors per core.

Intel was showing thier new 15 core Xeon processor for the Enterprise marketplace, their graphics core with adaptive clocking and their Haswell dual and quad core processors all in a 22nm Trigate FinFET process. The process features 9 layers of metal and an in-process MIM capacitor. The 15 core Xeon processor has a unique architecture in that is it “choppable” down to a 10 core or 6 core implementation. This allows the design to be built as a 4.31B transistor version with 15 cores, 2.89 B transistors for 10 cores and 1.86B transistors for 6 cores. Like the IBM design, the circuit has optimized SRAMs and eDRAM integrated into the archtiecture. The design is available in a 20 layer organic substrate that is organized as a 6-8-6 layer stacking with an integrated heat spreader with a total of 2011 landing pads that are on a 1.016mm hexagonal pitch. The package also features the decoupling capacitors on the package directly opposite the circuits.

The Haswell processor is the next generation of the consumer cores, and was shown in the same process but at dramatically lower power platforms for mobile devices such as Ultrabooks. These designs range from fan-less computers through desktops.

Finally, on a step back process, Applied Micro was showing their implementation of a 64bit ARM v8 processor in 40nm Bulk CMOS technology. The first generation product has 4 processor modules that consist of 2 identical CPU cores with a shared L2 cache, and a 4-wide out-of-order superscaler micro-architecture. This supports integer, floating point and 128b SIMD engines. The design is compatible with some levels of support for hardware virtualization. Once again, the design has multiple fine-grained clocks and power management using DVFS.
 

Similar Posts