| |

Active Archive for Analytics

September 19, 2013, SNIA Analytics and Big Data Summit, Santa Clara, CA—Mark Seamans from FileTek described the challenges and solution elements for using an active archive to store massive quantities of data for analytics. The growth of data volume causes the datacenter to become overloaded.

The big data revolution is putting tremendous pressure on IT departments. Data volume is growing at exponential rates, which challenges the storage systems, existing apps and processes, along with the IT staff. The current tiers of tape, disk, and cloud struggle to meet the demands of the data sets. Budgets cannot scale to handle the costs of capacity increases and staffing associated with traditional storage approaches.

The key IT processes exacerbate the problems as they no longer work for the growing data volumes. Data migration projects can induce data loss and corruption, since the error rates and data sizes are of the same order of magnitude.

One solution is to change the storage architecture to an active archive to get more flexibility and to meet the requirements for your organization. The issues of durability, access, availability, and performance have to be adjusted for “right costing” to be able to retrieve the value of those data over time. Backup data is only of value for a little over 2 years, while archive data may have value that exceeds the company’s lifetime.

A complete approach to large-scale “forever” file management calls for analysis, migration, and management strategies. Analysis helps to understand the what, where, and how of the data and includes metadata. The metadata helps to understand the assets and their use patterns, and include such information as creation date, class access, size, type, etc.

Migration has to be policy driven as a process and not as a project. Investments in the migration approach help to keep the migration from becoming a science project. Key capabilities in the migration include manageability, audit trails, reporting and alerting, throttling, and link support. If the data is migrated into an active archive, the users will still have direct file access for data analysis.

Small files create big problems in big data. Storing large numbers of small files causes storage and access issues. Storage solutions should provide capabilities for aggregating sets of small files into larger groups while providing users transparent access to each file. This approach is like a “tar ball” without needing to manage the tar process.

Large files have their own problems. Access requires though to the storage location. It may be more efficient to pre-stage data if large data sets are stored on slow access media like tape. Most active archive solutions provide staging capabilities and dynamic cache management to handle stated data elegantly as long as it is actively being used. A broad range of devices from many vendors supports active archive management. Storage hardware can be in disk, tape, cloud, SSD, or NBT in RAID, SAN or NAS configurations.

The challenge for big data is that the hardware hard error rates may be smaller than the data set. Disk errors represent the number of bits read before a sector failure, and for tape, the number of bits read before a bit failure. A consumer SATA drive has a BER 1014 or 0.01 PB equivalent data move before an error. Enterprise SATA is 1015, SAS/FC is 1016, LTO is 1017or 11.1 PB, and T10000A/B/C is 1019 or 1110.22 PB.

The challenge is that the data volumes are becoming larger than the BER, so the increased rebuild times and I/O increase the likelihood that another hard error will occur during rebuild. The reason that the BER is critical is that the time to an error is directly proportional to the BER. A bank of 100 consumer SATA drives is likely to have an error in 2.3 hours, versus LTO at 96.2 days or Oracle T10000 at 15 years.

An active archive presents a standard user and app access as NAS. The active archive layer then access the actual data on whatever medium it is on, and presents it to the user by moving the data sets through the various media for best access. This infinitely scalable “forever data store” allows the users to have continuous access to the data, while administrator can monitor and tune storage policies through the built-in management functions. The active archive layer handles media management, mirroring, storage interfaces, and other infrastructure functions.

Some case studies of active archive implementations show how this technology provides value. Yale University HPC center is a full service facility for genome analysis. They provide RNA expression profiling, DNA genotyping, and sequencing. The active archive is the core infrastructure for research data storage and provides automated storage management and proactive monitoring for data integrity and availability.

The National Institutes of Health develops new information technologies and infrastructures to help researchers study fundamental, molecular-level biomedical problems. Their sequence read archive is an application that captures, stores, shares and protects all genomic sequencing data funded by the US government. Ingest rates are multi-terabytes a day from many sites in a 24PB system. The storage can handle direct read from tape, virtualized storage, and scalability to trillions of files and many PB of managed data.

XOS Digital provides asset management, facility design, integration services, and digital coaching technologies for collegiate and professional sports organizations. Active archive is integrated into the XOS vault for asset management and is the secondary and tertiary repository for “in season” content and the primary repository for post-season content.

The value of active archive is cost management, scalability, flexibility, and performance. Storage approaches and technologies can evolve over time in a NAS-based approach. Built-in software handles capabilities and policies for management, increasing overall system automation. See www.activearchive.com for more details.
 

Similar Posts