Big data is a broad term that refers to datasets so large and complex that they require specially designed hardware and software tools for processing. Such datasets are usually terabytes or exabytes in size. These datasets are collected from a variety of sources: sensors, climate information, public information such as magazines, newspapers, and articles. Other examples of big data generation include purchase transaction records, web logs, medical records, military surveillance, video and image archives, and large-scale e-commerce.

There is a surge of interest in big data and big data analytics and their impact on businesses. Big data analytics is the process of studying large amounts of data to find patterns, correlations, and other useful information that can help businesses better adapt to changes and make smarter decisions.

1. Hadoop

Hadoop is a software framework that can perform distributed processing of large amounts of data. However, Hadoop processes data in a reliable, efficient, and scalable manner. Hadoop is reliable because it assumes that computing elements and storage will fail, so it maintains multiple copies of working data, ensuring that processing can be redistributed to failed nodes. Hadoop is efficient because it works in a parallel manner, speeding up processing through parallel processing. Hadoop is also scalable and can handle PB-level data. In addition, Hadoop relies on community servers, so its cost is relatively low and anyone can use it.

Hadoop

Hadoop is a distributed computing platform that allows users to easily architect and use. Users can easily develop and run applications that process massive amounts of data on Hadoop. It mainly has the following advantages:

1. High reliability.Hadoop's ability to store and process data is trustworthy.

2. High scalability. Hadoop distributes data among available computer clusters and completes computing tasks. These clusters can be easily expanded to thousands of nodes.

3. High efficiency. Hadoop can dynamically move data between nodes and ensure dynamic balance among nodes, so the processing speed is very fast.

4. High fault tolerance. Hadoop can automatically save multiple copies of data and automatically reassign failed tasks.

Hadoop comes with a framework written in Java, so it is ideal to run on Linux production platforms. Applications on Hadoop can also be written in other languages, such as C++.

2. HPCC

HPCC, abbreviation for High Performance Computing and Communications. In 1993, the Federal Coordinating Council for Science, Engineering, and Technology submitted to Congress the report "Grand Challenges: High Performance Computing and Communications", also known as the HPCC plan, which is the U.S. President's strategic science project. Its purpose is to solve a number of important scientific and technological challenges by strengthening research and development. HPCC is a plan implemented by the United States for the information superhighway. The implementation of this plan will cost tens of billions of dollars. Its main goals are: develop scalable computing systems and related software to support terabit-level network transmission performance, develop gigabit network technology, and expand research and education institutions and network connectivity capabilities.

HPCC

The project mainly consists of five parts:

1. High Performance Computer Systems (HPCS), including research on future generations of computer systems, system design tools, advanced typical systems, and evaluation of existing systems;

2. Advanced Software Technology and Algorithms (ASTA), including software support for grand challenge problems, new algorithm design, software branches and tools, computational science and high-performance computing research centers, etc.;

3. National Research and Education Network (NREN), including research and development of intermediate stations and gigabit-level transmission;

4. Basic Research and Human Resources (BRHR), including basic research, training, education, and course materials, designed to increase the flow of innovative ideas by rewarding investigator-initiated, long-term research in scalable high-performance computing; to expand the pool of skilled and trained personnel by enhancing education and high-performance computing training and communications; and to provide the necessary infrastructure to support these research and investigation activities;

5. Information Infrastructure Technology and Applications (IITA), with the aim of ensuring U.S. leadership in advanced information technology development.

3. Storm

Storm

Storm is free open-source software, a distributed, fault-tolerant real-time computing system. Storm can very reliably process huge data streams and is used to handle Hadoop batch data. Storm is simple, supports many programming languages, and is fun to use. Storm was open-sourced by Twitter. Other well-known application enterprises include Groupon, Taobao, Alipay, Alibaba, Happy Elements, Admaster, etc.

Storm has many application areas: real-time analytics, online machine learning, continuous computation, distributed RPC (Remote Procedure Call protocol, a way to request services from a remote computer program over a network), ETL (Extraction-Transformation-Loading, i.e., data extraction, transformation, and loading), and so on. Storm's processing speed is amazing: tested to process 1 million data tuples per second per node. Storm is scalable, fault-tolerant, and easy to set up and operate.

4. Apache Drill

In order to help enterprise users find more effective ways to speed up Hadoop data queries,the Apache Software Foundationrecently launched an open-source project called "Drill". Apache Drill implements Google's Dremel.

According to Hadoop vendorMapR Tomer Shiran, product manager at Technologies, said that "Drill" is already operating as an Apache incubator project and will be continuously promoted to software engineers worldwide.

Apache Drill

The project will create an open-source version of Google's Dremel Hadoop tool (Google uses this tool to speed up internet applications of Hadoop data analysis tools). And "Drill" will help Hadoop users achieve the goal of querying massive datasets faster.

The "Drill" project is actually inspired by Google's Dremel project: that project helps Google analyze and process massive datasets, including analyzing crawled web documents, tracking application data installed on Android Market, analyzing spam, analyzing test results on Google's distributed build system, and so on.

By developing the "Drill" Apache open-source project, organizations will be able to establish the API interfaces and flexible and powerful architecture that Drill belongs to, thereby helping to support a wide range of data sources, data formats, and query languages.

5. RapidMiner

RapidMiner is a world-leading data mining solution with advanced technology to a very large extent. It covers a wide range of data mining tasks, including various data arts, and can simplify the design and evaluation of data mining processes.

RapidMiner

Features and characteristics

  • Free provision of data mining techniques and libraries
  • 100% written in Java code (can run on operating systems)
  • Data mining process is simple, powerful, and intuitive
  • Internal XML guarantees a standardized format to represent and exchange data mining processes
  • Large-scale processes can be automated with simple scripting languages
  • Multi-level data views ensure effective and transparent data
  • Interactive prototyping with graphical user interface
  • Command line (batch mode) for automated large-scale applications
  • Java API (Application Programming Interface)
  • Simple plugin and promotion mechanisms
  • Powerful visualization engine, many cutting-edge visualization modeling for high-dimensional data
  • Support for more than 400 data mining operators

Yale University has successfully applied it in many different application fields, including text mining, multimedia mining, feature design, data stream mining, integrated development methods, and distributed data mining.

6. Pentaho BI

The Pentaho BI platform is different from traditional BI products. It is a process-centric, solution-oriented framework. Its purpose is to integrate a series of enterprise-level BI products, open-source software, APIs, and other components to facilitate the development of business intelligence applications. Its emergence enables a series of independent BI-oriented products such as Jfree, Quartz, etc. to be integrated together to form complex, complete business intelligence solutions.

Pentaho BI

The Pentaho BI platform, the core architecture and foundation of the Pentaho Open BI suite, is process-centric because its central controller is a workflow engine. The workflow engine uses process definitions to define business intelligence processes executed on the BI platform. Processes can be easily customized, and new processes can be added. The BI platform includes components and reports to analyze the performance of these processes. Currently, the main components of Pentaho include report generation, analysis, data mining, workflow management, and so on. These components are integrated into the Pentaho platform through technologies such as J2EE, WebService, SOAP, HTTP, Java, JavaScript, and Portals. Pentaho is primarily distributed in the form of the Pentaho SDK.

The Pentaho SDK contains five parts: the Pentaho platform, the Pentaho sample database, the standalone-running Pentaho platform, the Pentaho solution examples, and a preconfigured Pentaho web server. Among them, the Pentaho platform is the most important part of the Pentaho platform, encompassing the main body of the Pentaho platform source code; the Pentaho database provides data services for the normal operation of the Pentaho platform, including configuration information, Solution-related information, etc. It is not necessary for the Pentaho platform and can be replaced by other database services through configuration; the standalone-running Pentaho platform is an example of the standalone running mode of the Pentaho platform, demonstrating how to run the Pentaho platform independently without application server support.

The Pentaho solution example is an Eclipse project used to demonstrate how to develop related business intelligence solutions for the Pentaho platform.

The Pentaho BI platform is built on servers, engines, and components. These provide the system's J2EE server, security, portal, workflow, rule engine, charts, collaboration, content management, data integration, analysis, and modeling functions. Most of these components are standards-based and can be replaced with other products.

Author: Jingwei Fanglue
Source: http://www.36dsj.com/archives/22617