|

Low-Latency SAN

April 7, 2015, Data Storage Innovations Conference, Santa Clara, CA—Rubin Mohan from HP and Craig Carlson from QLogic talked about the next-generation of low latency SAN. They described the various components and aspects of storage performance and how those factors affect latency.

Everyone is at least vaguely familiar with the basic performance measurements. IOPS is I/O per second, bandwidth is a data rate of bits per second. Latency in a system is very different from that of a single drive or other component, and is dropping from the high ms to us. The move to higher performance drives and new architectures is the driver for the changes.

Storage performance is driven by the protocol. Within the datacenter, the choices are Infiniband (IB), Fibre Channel (FC), Fiber Channel over Ethernet (FCoE), and iSCSI. The underlying protocol affects performance and reliability. FC is the fabric of choice for the enterprise networks but needs to be a separate network plus some type of bridge function. FCoE has promise to overcome the bridging issues, but still needs to improve to meet the needs for reliability. iSCSI is resurging especially at the 10 Gbps data rates due to its low cost and well known infrastructure.

One would think that getting faster storage and switches would have a significant impact on network performance. At present, switch latency is only a small part of the total delay, which is dominated by the network adapters and interfaces. As the systems become faster, the switching speeds will become more important.

A faster set of drives does have an impact on the overall latency. An all flash system can have under 500 us end-end latency with 16 Gbps infrastructure and is about 2.5 times lower latency that other configurations. This improvement will be even better when the next generation of FC comes out at 32 Gbps or in a multi-lane configuration to 128 Gbps. The first products will be out next year.

The network-level latency of a SCSI drive on FC, FCOE or iSCSI interface will be about 1-2.5 ms. Changing to SSD will decrease latency to 0.5-1 ms. Changing the interface to some type of RDMA (Remote Direct Memory Access) like iSER or SMB can reduce this latency further to 100 to 300 us with SSDs. The advent of NVMe will become significant in ’17. The RDMA bypasses most layers of software and provides a direct attachment to the memory hardware.

RDMA presents a zero-copy capability which reduces CPU context switching performed at the target. The elimination of this software intensive process impacts latency. The main versions of RDMA are iSER which is iSCSI over RDMA and SMB direct which is NFS using direct I/O. both are at least partially supported by VMware. Further work on the stacks is needed to get more performance from these interface. The other underlying infrastructure can involve tradeoffs between latency and system complexity, costs and system congestion.

The NVMe group is in the process of releasing the next generation of their specifications, which will be fabric agnostic. This spec will enable NVMe transport over RDMA or FC, and adapters will allow connections to other fabrics. The new spec defines new functions like mapping, control function control, end-point identification, error recovery, and authentication.

FC-NVMe will handle the flash on existing FCP protocols and fabrics with the FCP semantics for transport. Since FC is dominant in the datacenter, the FC-NVMe spec will permit standard development to add NVMe into the FC ecosystem right now. While considering the addition of NVMe to the storage infrastructure, you should also consider the nature of the cabling and adapters. The wiring and optical fibers are long-life infrastructure, and many more functions will converge beyond the current 10 Gbps data rates. As a result, one should consider adopting the latest version of fibers and cabling (OM4 and CAT7) as a part of any upgrades and additions. Increasing the bandwidth will reduce latencies.
 

Similar Posts