| |

ACG Looks at Big Data

 October 27, 2011, Association for Corporate Growth, Santa Clara, CA–The Silicon Valley chapter of ACG looked at big data in a panel moderated by John Furrier of Silicon Angle. The trends in big data are being generated by the growing rate of data production and by consumers who are using these data in new ways. The panel included George Symons from Solix, Bassel Ojjeh from nPario, David Lyle from Informatica, Bill Schmarzo from EMC, and Stacy Passeri from KiteTail.

Big data is increasing the rate of changes in data stores. Consumers are finding new ways to use new and older (legacy created) data?
Symons stated that users must plan on data lifecycles, and develop management practices including data tiers, because this is not just a database at a large scale. The challenge really is what to do with all the new data. Users need metadata to enable backend management, and help to manage the data streams. Automatically generated metadata is much more valuable than none at all.

Many data acquisition systems have a “call home” function, which is a log of short-term operations, functions, etc. that gets sent back to the vendor. Over the long term, this data is a source for data mining to improve the customer experience. For example, you could predict a type of drive that’s developing problems and be proactive in working with customers of those drives.

Ojjeh noted that a real challenge is working with real-time data. Worms have evolved to use the collected data from infected computers, such as performance monitoring for data centers. Virtualization makes it harder to have analytics and to detect issues in real-time. The key requirement for most systems is up time.

Lyle suggested that data integration is the real challenge. Now, people can generate over a terabyte of data to populate their data warehouses. Most of the social media companies have decided that their content is not worth the cost or effort to store all the data. They just process the data even though they could do more analysis. But the skills are changing, so the question to ask is what is the business value of all this data. New technologies and programming languages are coming on the scene all the time. Statistics are moving big data into a space to analyze. People are concerned about how to use the data before their competitors figure out how to use the same information.

Schmarzo asked why bother with all the data. A utility company in the East Coast installed smart meters and now knows everything about electrical functions in terms of use, time of day, and duration of use. Now they can offer more services to customers, such as maintenance predictions based on benchmarks of similar appliances. They can create a market channel for high-efficiency products based on current usage modes. And finally, they can become a resource for effective energy conservation measures. This changes the way they operate and forces them to look at value and dynamics in their data.

Passeri offered increased transparency and accountability as the value for big data. Their company is working to access foundational data and use it to help people in the financial markets. One issue is proprietary ownership of the data, because access could ruin spreads between income and expenses. They separate the data from the owner and anonymize it to be a third-party data aggregation platform. Over time, big data will be used as currency greater value than cash, and will become a line item on a balance sheet.

Technology is changing very rapidly. Companies need to develop new business models and track trend data. What enables data to be part of your mindset, and how is that a change from 5 years ago?
Schmarzo offered examples of the book “Money Ball”. The leading companies are creating metrics that matter in changing a business. This metric was dollars per win in baseball. Some of the existing metrics such as RBI and batting average are poor predictors of performance. Although some players try gaming the system, in general selecting appropriate metrics can measure true performance.

Big data requires better application of metrics. Analytics and cloud computing are not the important factors, and neither is new technology. Companies must get to better decisions more quickly, to gain advantages. In addition, companies are scared of what their competitors might be doing. This use of data is increasing very rapidly, as companies need to make decisions based on the velocity and granularity of those data.

Profit? Losers?
Lyle opined that companies who are not moving will get left behind. Each vertical depends on people with business knowledge to develop their key metrics. The trick is to not ask any question but to ask the proper questions. High technology, big data, and mobile apps for the big drivers, because it’s now cost effective to analyze all the data.

How to get to big data?
Passeri reasoned that many institutions, government, and individuals are not up-to-date on the hardware and software. As a result, they sell their services without using the latest technology buzzwords, but instead offer a solution. They use simple language for the various financial concepts. Their solution is a clearing platform for counties. They offer solutions rather than big data.

Data storage and data centers?
Ojjeh replied that storage systems are now able take advantage of efficient computation for in-line, real-time analysis, versus the old method of storing and then analyzing. Both homegrown and off-the-shelf tools are using more intelligent algorithms to solve the problems. Tools like Hadoop offer a solution on how to use the data. For data mining, it is better to analyze on-the-fly, and much cheaper than storage.

Symons observed the industry is going through a transition. Google now uses off-the-shelf systems for their compute and storage nodes. They are not buying enterprise grade hardware because they let the equipment die and discard it, rather than expend time and resources on maintenance. There is always a trade-off between on-the-fly analysis versus direct attached storage when doing data analysis operations.

Winners and losers?
Schmarzo responded that technology exists forever. Some technologies become zombies, like the first generation of business intelligence tools. This technology failed because it was too techie and had very bad user interfaces. Simplification is an art, the best companies focus on ease of use. Now, user interfaces are changing to take advantage of the greater intelligence in computers and networks. These capabilities make the systems easier to use.
Lyle answered virtualized data versus warehouses. It is unclear in the warehouse how much of the data require lots of processing. Instead, the warehouse is becoming a long-term cache with pre-calculated addresses.

Hadoop and similar tools are open source. How to make profits with open source software?
Symons suggested that the user needs to understand the issues of what data do I need, and what data help me make better decisions as a part of business line management.
Schmarzo added the user has to inventory the set of decisions that need attention. For example, Yahoo has 21 questions and actions, and uses data and analytics to support the decisions.
Lyle opined that business intelligence is a bottoms-up process, but it is better to prototype and model first, then implement.
Ojjeh stated that the key is to determine what is the right measure. The user interface is important. Older business intelligence tools focused on the analytics, but didn’t work to determine if the right metrics were in use. Now tools are comparing metrics and validating them in context.

How to determine the right questions?
Lyle stated that users need someone with domain expertise on the development team. Then you can build-in and encapsulate that knowledge.
Schmarzo said that Procter and Gamble uses a small number of people to analyze the business. This business analysis is not the same as analyzing the data.
Passeri considered risk mitigation to be important. Even in large organizations like the government, they can look for patterns and detect fraud rings. It is important to build in an advocate to review how the technology impacts the human resources.

By ’18, the human resources will be short 1.5 million people with skills in big data. How to manage and also incorporate new data?
Symons suggested that predictability is important. Users can look at error rates to predict the need for maintenance or when replacement is necessary.
Schmarzo figured that the human resources is the biggest challenge. The technology will be ok and the educational institutions are starting to move on the issue. The problem is that users will need to have general business people learn about the data and statistics. Everyone will have to learn to use data to improve decisions.

Question: why not just make smart apps?
Lyle responded you have to hide the complexity under the app with some type of smart agents. The technology is available and capable, but the developers are not.
Symons added that you need business and finance expertise to make the app work. Developers need to have a fairly high level of understanding of the data to be able to run a business.

Question: what about the effects of big data on privacy. Cloud and mobile are good developer environments for disrupting existing businesses. Now developers can do new things with the technology, leading to new venture creations?
Schmarzo opined that it may be good, if not private. Targeted ads are more likely to be viewed and get better responses than general ads, because everyone wants a better experience. Aha moments come from knowledge gleaned from existing data and enable improvements in service levels.

Real-time decisions and analytics?
Schmarzo offered velocity and complexity are the key factors. Other factors are a compromise between technology perspectives and design complexity.
Symons noted that a company trying to manage costs needs to understand tier supply chain and revenue details. The analytics are still important for these functions.
Passeri noted that some banks have purpose-build apps for big data feeds through Linked-in, Twitter, etc. to give wealth managers iPads the ability to show clients real-time data. They integrate the data in the cloud. At the same time, the cloud disintegrates the analysis processes. Data is harder to get to and too spread out.

Similar Posts