Showing posts with label analytics. Show all posts
Showing posts with label analytics. Show all posts

Thursday, November 29, 2012

Big Data, Analytics and Hadoop


 Three of the biggest catch terms in technology today are Big Data, Hadoop and Analytics.  Often times, they are used interchangeably, when in reality they are three unique capabilities and categories.  Each solves a different, but related, set of struggles within IT and the businesses that leverage IT for unique advantages.

Start with Big Data.  Big Data is a term we hear more and more often about the developing struggles at many companies to cope with growing volumes of data and changing varieties of data.  There is no mark that says where Big Data begins, but is rather a dramatic change for an organization that requires them to look at new tools, technology, processes and skill sets to address the challenges they face.  Big Data is focused on the efficient storage, movement and organization of these evolving data types.

Analytics is related to Big Data, but has some variations.  Where as Big Data focuses on the data itself and the infrastructure to manage the data, Analytics is about putting that data to use and enhancing the capabilities of an organization.  Analytics is about taking that data and ensuring that actionable information comes from it that the company can execute on.

Finally, Hadoop, one of the most common technologies being talked about today.  Hadoop is a technology that can be used to enable organizations to manage their Big Data, while building a platform for more advanced analytic capabilities.  Hadoop is a rapidly growing open source project that has been adopted by many main stream organizations as a standard method for the storage and processing of complex data sets.  Hadoop is one of many tools that you can use to enable Big Data and begin to implement an Analytics strategy.

Wednesday, July 18, 2012

Data Curation

Big Data is a term that we are hearing more and more to describe the growing challenges in managing the volume, variety and velocity of today’s data sets. Data today is being gathered from more diverse sources, analyzed more often and decisions made more quickly then ever before. But, ultimately, Big Data is not a technology, but rather changes to the methodologies that an organization uses to turn data into information. Big Data challenges can be resolved with a variety of new technologies on the market, but at a fundamental level, Big Data is process.

One growing area of focus within the Big Data space is Data Curation (DC). DC is the organization and tracking of this data, turned information, to ensure that it is accessible in an understandable and reliable fashion over a given period of time. There are three primary components too DC:

Finding the Data – Ensuring that the data is uniformly organized against a documented standard.

Ensuring Data Availability – Data, like any physical asset has a lifecycle. Data access must be planned over the lifecycle of the data to ensure it is available today and in the future against a pre-defined set of criteria for retention and accessibility.

Data Quality – Ensuring that data within an environment meets a minimum level of quality standards to ensure reporting and analysis activities are not tainted by non-compliant data.

If you think of incoming data as a giant set of piles with no labels or patterns, Data Curation is the process to turn that pile into information that is neatly organized in file folders consistent against a set of standards to enable users can quickly locate the information required.

As data volumes grow the lifecycle of the associated information becomes more and more difficult to manage, both the process and the underlying technology. All technologies that store date must ultimately be replaced at the end of their useful life, DC is the process to ensure that data is properly migrated to new platforms, the integrity of that data protected and the accessibility of the data not impacted in a negative way. DC must span the entire life of the data including creation, management, migration and ultimately disposal, per appropriate policies.

Data Curation is about organization of data to create information, while ensuring access to that information. DC is technology agnostic and requires a mindset beyond just acquisition of the data, but also inclusive of how long the data must be made available and retained. Firms that work with Big Data must consider the implications of their growing data sets as they work to understand the data and make better business decisions based on the derived information.