Processor Overview at ISSCC2012
February 20, 2012 – The theme of the microprocessor session at ISSCC was no longer raw speed and performance, rather it was data processing throughput per watt of power. Low power and reduced stages for data processing were the key.
The session opened with Intel presenting the 22nm introduction of their new ivybridge processor which featured 4 IA-32 CPU cores built with 3D FinFET devices, and added 2 full Direct X11 compatible GPU, security processing, and a PCIe controller in a single die. The die will directly support up to 3 displays through a DisplayPort interface. The key for the device is the amount of data that is processed through the internal bus and caches. The low power implementation has a much lower leakage due to the 3D transistors than the 32nm HKMG process. The power plan for the design includes using a gated power core, a systematic algorithmic solution for low power that was optimized for clocks, sequential and random logic sizing, and multiple frequency domains with separate PLL power planes. This flow allows the minimum operating voltage for the core and caches to be reduced by over 250mV over previous processes.
Oracle presented a next generation (T4) 64bit SPARC processor SOC that is based on the network server market applications. The 8 SPARC core device features a crossbar switch and unified 16 way L3 Cache for 5x integer and 7x floating point improvement in single thread, since pass performance. The new instruction pipeline allows for data and bus coherency between the cores and the DDR3, PCIe Gen2 interface and the Ethernet interfaces. The single execution throughput design allows for predictive branch processing as well as an ALU and Branching unit that supports the AES security core and a separate Floating Point Graphics Unit on the same die. Power optimization was once again gained by gated logic on the core, gated power for unused cores and a redesign into 32 separate and optimized SRAM blocks.
The other highlight from the session was also an Intel presentation of their new mobile solution. It was a dual Atom core built on a 32nm CMOS process that includes the 802.11 RF Transciever. The low power, in order execution processors, were built using only 3 different types of transistor pairs – logic (HP/LP devices), low power (LP devices) and the HV I/O devices that need to support the 1.8v/3.3v interface. To support the RF section, the device uses a high resistivity substrate which allows for the creation of an RF optimized metal stack for making high Q antennas and inductors. The device Atom cores were scaled designs from the known 45nm cores and added the standard test wrapper for the IP blocks on the port as well as incorporating a new IDSF (Intel on Die Switch Fabric) for the bus. The RF features spread spectrum control and an auto skew management system that holds the system to within 5% of phase shift and can minimize clock injection to the substrate.
The other papers in the session were on the same theme – maximize direct networking throughput (ethernet, internal bus, WiFi) and low power into a known instruction set for code compatibility. The new benchmark is being based on Gbps/W rather than Ghz.


