| | |

ISSCC 2013 – Session 3 new processor tech

February 2013 – This years Solid State Circuit conference featured a processor session that was focusing on increasing cache size and more cores at low power. The rise in data size being handled by these units, has driven the systems to explore new architectures that are much more throughput oriented rather than MIPS/GFLOPs oriented. This big data appears in two forms, and has resulted in a split in the architecture roadmap. One path is for very large numbers of small data length data sets such as transactional data, the other is for small numbers of large data length data sets such as multi-media content.

The session started with the IBM paper on their new zEnterprise EC12 system. This multi-cchip module uses 6 six-core CPU chips (total of 36 cores) and 2 L4 Cache chips. New power optimization includes the use of embedded DRAM in the design, pushing the L4 cache up to 384MB. The process for both chips features 4 logic thresholds with a Low Vt option for some critical paths, 15 layers of metal and is built on a 32nm HKMG process. One of the major changes in the design is the use of multiple clock grids. The cores, L1 and L2 cache feature independent 1:1 clock grids, while the L3 cache uses an independent lower power 2:1 grid. The I/Os for the MCU chip are asynchronous and are on an additional independent grid. This results in about 35ps of clock skew which can be zeroed by delay control logic.

The next paper by Oracle was a 16 core CPU that could be designed into an 8 socket motherboard design. The T5 system, in addition to the expanded 8 DDR3 schedulers, the 7 Coherency handlers at 12.8Gb/s and the 2×8 PCIe 3.0 interfaces at 8GT/s do so with a new DVFS architecture to reduce power.

The balance of the papers from AMD, Fujitsu, and several universities from China and Taiwan all continued the theme of increased I/O, high core counts (up to 64) in the context of Low Power. The Low Power criteria drove the creation of new clocking, and cell designs. The increased I/O caused a systematic doubling of the higher order L3 & L4 caches used in the systems. This increased throughput is directly attributable to the rise in big data and the need to get the latency off the bus to be able to process it and not have the CPUs idle.

Similar Posts