| |

Hadoop application guidelines for big data

September 2012 – At the Dataweek conference in San Francisco, a great deal of the discussion was on Hadoop and the appropriate use of this new storage and data retrieval technology. A key point is understanding what type of data is most appropriate for the Hadoop system. This structure is optimized for sequential data that is addressed as a whole set, not for individual record retrieval. These large data sets are generally unstructured, which requires the analysis to be done on the full Hadoop data block, with adhoc query. Traditional RDBMS environments are best for single record selection from a known structured RDBMS schema that is architect with vertical apriori query in mind.

A challenge for Hadoop is the database solution uses map-reduce as the primary qery search technique. This low level command, however, make for complications and high degrees of coding proficiency to use the solution. To address this need, there are higher level language APIs such as the open source solution Pig and Hive. These are both part of the Apache open source project.

To help facilitate the query, a new command structure called hcatalog is being used to bring SQL like query systems to the Hadoop data. This allows for the selection of smaller data sets from the global data base and then pass the results to visualization tools from third parties. The main marketplace around Hadoop is not the database itself, but the analytic tools and API based query tools to simplify the query task. These additional language tools support higher level languages for programming the query.

One other key application note is to not use a database schema on top of Hadoop. The unstructured nature of the system is not meant for real time use or single data view use – these are SQL concepts. The distributed nature and duplication of the Hadoop environment works best for large numbers of small data objects rather than large data objects such as media files. These media files, audio, video and stills, are extended length objects and not well suited to the replication and division of the Hadoop storage environment. For these files, the creation of a an index fle with a traditional RDBMS is the most highly optimized.

Similar Posts