| |

Making Enterprise Data Warehouses More Flexible with Hadoop

September 19, 2013, SNIA Analytics and Big Data Summit, Santa Clara, CA—Rob Rosen from Pentaho described the steps to make enterprise data warehouses more flexible by using Hadoop as a front end. The changing nature of the data calls for better ways to manage the unstructured data types.

Enterprise Data Warehouses (EDW) are facing a crisis, only one in six have architectures that can scale within a reasonable schedule. By next year, there will be more data sources and uses, and more of that data will be unstructured. The current practice of just storing all data collected and hoping that someone figures out later how to use that data, is un-realistic. The data sets are getting more complex and their analysis requires an understanding of the context.

An increasingly common way to adapt to the big data problem is to put a Hadoop cluster on the front-end of the process to provide some interface structure. The problem is that Hadoop is geeky. A Hadoop cluster offers the promise of competitive advantage through finding new ways to increase revenue or improving operational efficiency while reducing costs. As a part of the infrastructure, Hadoop is very hard to use in standard form, has high barriers to implementation for new talent and has the lack of a business sponsor.

In one example, a telco started call volume analysis. The structured data could be exported, transformed, and loaded into a system, but the unstructured data and metadata to go along with them are necessary to get useful results. The EDW was not scalable due to high investments and reduced user responsiveness so query changes created higher complexity, more compute resources, greater I/O volume, and increased latency for the overall system.

Instead, they put a Hadoop cluster on the front end and reduced the costs of hardware changes through the use of commodity hardware. Open source tools and other architectures to change the technologies led to a speed up in interactive use. The result is a combined structured and unstructured data set in ETL with Hadoop. They also created an active archive that reduced costs compared with the EDW storage costs.

But there is no free lunch. The issues related to the transition to Hadoop from their prior EDW included changing technologies, changes in the Hadoop stack, and replacing Hive batch functions with interactive ones. These changes will eventually displace more of the EDW systems and allow users to get to the data in ways that matter to the company.

Future systems will move to a visual mapping function for analysis and visualization to enable a visual map reduce process. SQL over Hadoop will allow the users to extract more from context.
 

Similar Posts