ARM Servers and Interconnect take the Stage
February 6, 2013 – The Linley Data Center conference relocated to the Hyatt in Santa Clara for this year’s event focusing on interconnect/networking and power. The keys were the advances in high speed interconnect as driven by the cloud and software defines networking, as well as optimizing performance/power for the growing data centers & servers.
With a lead from Big Switch Networks on Software Defined Networking (SDN), they overviewed the requirements of different operating systems and their need for different points of optimization. The inclusion of new Operating Systems such as Android and iOS as endpoints in the network, are driving the servers to address flexibility for mixed OS networks. This change in endpoint OS will impact the servers much more as the Internet of Things takes hold, and much of the M2M traffic comes from custom Linux kernels from these devices. As a result, the flexibility of SDN is needed to help optimize the data centers in the cloud to there clients needs and data models.
Following the SDN intro, Aquantia, Inphi and Applied Micro Circuits presented on solutions and the growing application need for 40G and 100G aggregation networks. These networks are still dominated by 10G ports, typically with short run copper and multi-mode fiber solutions. The 40G solutions are 4×10 configurations. The market is still 90%+ 10x10G for the 100G solutions, but 4x25G solutions are starting to appear in the market since the new silicon became available in mid 2012. The software control and gearbox to translate the 10×10 to 4×25 systems are also starting to emerge as they shave shifted to being manufactured in CMOS. The main systems in the data center are shifting to 100G as the cost differential for wholesale conversion to 40G vs 100G is not a major difference as a transition from 10G.
On the server side, the discussion shifted to the ARM vs Intel debate. Calxeda and Applied Micro Circuits showed their new ARM based plug in cards that were featuring multi-core ARM solutions. At this time the solutions are based on A9 cores, and A15’s are planned in the near future. These are running at a lower power per core than Intel designs and are differentiated by the SOC that include them. Calxeda and Applied Micro took two different approaches to the Server on a chip SOC – one uses stock ARM cores with custom interconnect/fabrics and peripheral controllers, and the the other has an Architectural license for ARM and uses a customized core with standard interconnects.
These two methods provide different points of optimization for the end application. One is better for edge of network, web servers and applet processing, the other is optimized for high speed transactional and DBMS applications with up to 1/2TB of DRAM per multi-core server. The challenge with these cores for the mainstream compute side of the servers is the address space. The A9 architecture is a 32 bit design for physical and virtual memory addressing the A15 is half 32 and 64 bit and then the newly created A57 is 64/64 bit. As a result, the standard server virutualization via hypervisor from VMWare and other software do not currently run on the platform at the same memory and I/O mapping as the x86 cores. This will then require a code port for applications operating on these devices. The power per core has to balanced with the power per user when comparing to the x86 with hypervisor applications to truly understand the end operating costs. One of the other key costs is the system – not the CPU. Some of the ARM based SoCs include the SATA and NIC (up to 10G) interface as part of the part, rather than needing a separate chip with the Southbridge for the design.
The overall trend is that power per user per throughput is the key operating point for new data centers and servers. Selection of memory, interfaces from DDR2 vs DDR3 and DDR4 and PCIe Gen2 vs Gen3 area all points to look at for the system. The ARM race is on, but there are no 64 bit products targeted until 2014. These parts are also under a foundry model where allocation of process space for the smaller nodes has to fight against the needs of the mobile devices – hence most of the current parts are on 40nm for 2013. Intel for the x86 architecture, is moving forward is code commonality, and process advancement (14nm for 2014) for power reduction. The race is still up for grabs for the next generation servers.


