| |

SSD Tips and Tricks

August 14, 2013, Flash Memory Summit, Santa Clara, CA—Swapna Yasarapu from STEC moderated a panel on SSD use cases and implementation. Panel members were George Crump from Storage Switzerland, Bruce Moxon from STEC, and Sean Stead from STEC.

Stead provided an overview of flash in storage, including platforms and form factors, as well as the main interfaces. The flash interfaces of choice for the larger storage arrays include PCIe and NVMe as the direct attached and in-box interfaces and form factors, while SATA (serial advanced technology attachment) is being displaced by SAS (serial-attached SCSI (Small Computer System Interface)) in the enterprise.

The major value for changing from HDDs to SSDs is in speeding up data access. Both consumers and enterprises are seeing the improvements, with the enterprise getting advanced flash management tools, DSP for better data integrity. The enterprise-class storage products have improved endurance, performance, and reliability, as well as error correction built in to the devices.

At the same time, there are still issues with endurance and wear out as geometries continue to shrink. The reduced device geometries exacerbate the reliability and error issues. Nevertheless, expect to see increases in unit volume and capacity across all market segments over time.

Crump agreed on the endurance and reliability issues, and added that security is an issue. Unlike a HDD that can be degaussed and crushed, the flash in the SSDs can retain their data unless you crush the ICs.

Endurance and wear out are getting worse as the industry moves from SLC to MLC and now to TLC architectures. The problems get worse with each new process generation, but may be addressed by better controllers. One solution is to over-provision the system, but this adds significant costs. More intelligent controllers are moving to intelligent write schemes and using larger DRAM buffers for better organized write and coalesce functions. The big issue here is to have enough energy storage to maintain the data in the DRAM until the data are written to the flash.

The reliability issues call for flash-aware RAID architectures and better power management, especially for the DRAM buffers. These requirements may be a significant driver for adoption of MRAM or other high-speed non-volatile memory devices. Moving more flash into RAID forces users to account for in-place chip failures and data redundancy. Power failures are especially bad for flash devices and reliability, so power failure management calls for UPS and NVRAM in larger systems to stage data for gentle shut down.

Security and erase is getting more difficult as the flash drives all generate hidden cells. Bad cells are defined as read-only, so you cannot erase them. If enough cells are deemed bad and comprise a major portion of a block, the data in them can be recovered. The alternative is to use self-encrypted drives that are crypto-erased by deleting the key.

The need for over provisioning is reduced with intelligence, as are issues with reliability and power failures.

Moxon offered app and operations considerations at the system level. The trends in the datacenter are to move physical assets to virtualized ones to reduce CAPEX. The advent of the cloud architectures favors agility and reduced OPEX. The hyperscale datacenter is moving towards new architectures with no SAN. This characteristic also applies to high-performance computing, and other structures that need scale-out capabilities.

All of the functions at the server level are not in hardware, but in a virtual machine. Scaled out systems use SHARD (Software Hazard Analysis and Resolution in Design) techniques. In a private clouds, users are going towards open stack and open software defined networks without SANs as separate racks.

Some options for solid-state storage deployment are as low-latency direct attached storage (DAS), SAN/DAS cache, NAS/SAN, control cache, tiered storage, high-performance storage parks, and all solid-state storage arrays. The various use cases have different benefits, but all share increased performance and density as well as sequential I/O improvements. Operations implications for these changes must consider the scope of the impact against operational impedance. Existing migration, high availability, management tools, and operational procedures are all affected.

In a host-side cache, the SSD can be configured as a read cache on the DAS in a write-through mode. This allows all the data to be written to the HDD SAN after caching. In a host-based primary store, the SSD is direct attached storage, but in high availability mode. There are questions on how to best replicate the data to the SAN, but changes in operations can address the issues. Finally, the SSDs can be arranged as a solid-state array within a software-defined store in SAN, NAS, or high-availability DAS.

As an example of the outcomes of these system-level changes, they tested database performance. With 40 HDDs, they got 2510 transactions per second. Adding 240 GB solid state cache raised the number of transactions to 9844. Increasing the cache to 480 GB doubled the number of transactions versus the smaller cache. The increases were even higher when the additional cache was in PCIe RAID.

The changes for the next generation of systems needs users to change their architectures and apps to take advantage of the higher throughput and IOPS. Memory and object abstraction also need updating. In addition, other non-volatile RAM can be set up as a separate memory pool.

Increase endurance without over-provisioning?
Crump acknowledged that some over-provisioning is necessary, but can be minimized. The system needs to have more control to manage the solid-state storage. For endurance, the user needs to know when the devices will fail, by monitoring performance and using predictive analytics for all drives.
Moxon added that big data apps need to have their write paths optimized for I/O. endurance under random writes needs better garbage collection, as write amplification causes multiple writes to locations. Sequential writes are ok , because this results in larger chunks to store. Random reads are calling for new classes of memory like NVMe front ends.

Always on encryption and key disposal?
Crump noted that the keys can be stored anywhere in the system, and the drives considered as appliances. Pulling the key kills all data in the drive.
Moxon commented that key management is a part of the Trusted Compute Group’s tools.
 

Similar Posts