NoSQL Introduction
NoSQL (NoSQL = Not Only SQL) means "not only SQL".
In modern computing systems, a huge amount of data is generated on the network every day.
A large part of this data is handled by relational database management systems (RDBMS). In 1970, E.F. Codd's paper on the relational model, "A relational model of data for large shared data banks," made data modeling and application programming simpler.
As proven in practice, the relational model is very suitable for client-server programming and has far exceeded expected benefits. Today it is the dominant technology for structured data storage in network and business applications.
NoSQL is a revolutionary new database movement. It was proposed early on and gained increasing momentum by 2009. Proponents of NoSQL advocate the use of non-relational data storage. Compared with the overwhelming use of relational databases, this concept is undoubtedly an injection of entirely new thinking.
. Eric Evans from Rackspace once again proposed the concept of NoSQL. At that time, NoSQL mainly referred to non-relational, distributed database design patterns that do not provide ACID.
A transaction is called "transaction" in English, similar to transactions in the real world. It has the following four characteristics:
1. A (Atomicity)
Atomicity is easy to understand: all operations in a transaction are either completely completed or not done at all. The condition for a transaction to succeed is that all operations succeed. If any one operation fails, the entire transaction fails and needs to be rolled back.
For example, in a bank transfer of 100 yuan from account A to account B, there are two steps: 1) take 100 yuan from account A; 2) deposit 100 yuan into account B. These two steps must be completed together or not at all. If only the first step is completed and the second fails, the money will inexplicably be 100 yuan short.
2. C (Consistency)
Consistency is also easy to understand: the database must always be in a consistent state, and the execution of a transaction will not change the database's original consistency constraints.
For example, given an integrity constraint a+b=10, if a transaction changes a, then b must be changed so that a+b=10 still holds after the transaction, otherwise the transaction fails.
3. I (Isolation)
So-called independence means that concurrent transactions do not affect each other. If a transaction wants to access data that is being modified by another transaction, as long as the other transaction has not committed, the data it accesses will not be affected by the uncommitted transaction.
For example, if a transaction transfers 100 yuan from account A to account B, and before this transaction is completed, if B queries its account at this time, it will not see the newly added 100 yuan.
4. D (Durability)
Durability means that once a transaction is committed, its changes will be permanently saved in the database, and will not be lost even if a crash occurs.
Distributed Systems
A distributed system consists of multiple computers and software components for communication, connected via a computer network (local network or wide area network).
A distributed system is a software system built on top of a network. Because of the characteristics of software, distributed systems have a high degree of cohesion and transparency.
Therefore, the difference between a network and a distributed system lies more in the high-level software (especially the operating system) than in the hardware.
Distributed systems can be applied on different platforms such as PCs, workstations, local area networks, and wide area networks.
Advantages of Distributed Computing
Reliability (fault tolerance):
An important advantage of distributed computing systems is reliability. A system crash on one server does not affect the other servers.
Scalability:
In a distributed computing system, more machines can be added as needed.
Resource sharing:
Sharing data is an essential application, such as banking and reservation systems.
Flexibility:
Since the system is very flexible, it is easy to install, implement, and debug new services.
Faster speed:
A distributed computing system can have the computing power of multiple computers, giving it faster processing speed than other systems.
Open system:
Because it is an open system, the service can be accessed locally or remotely.
Higher performance:
Compared with centralized computer networks, clusters can provide higher performance (and a better cost-performance ratio).
Disadvantages of Distributed Computing
Troubleshooting:
Troubleshooting and diagnosing problems.
Software:
Less software support is a major disadvantage of distributed computing systems.
Network:
Network infrastructure problems, including transmission issues, high load, information loss, etc.
Security:
The characteristics of open systems cause distributed computing systems to have problems such as data security and sharing risks.
What is NoSQL?
NoSQL refers to non-relational databases. NoSQL is sometimes also called an abbreviation for Not Only SQL. It is a general term for database management systems that are different from traditional relational databases.
NoSQL is used for the storage of extremely large-scale data (for example, Google or Facebook collect trillions of bits of data for their users every day). These types of data storage do not require a fixed schema and can be horizontally scaled without extra operations.
Why use NoSQL?
Today, we can easily access and scrape data through third-party platforms (such as Google, Facebook, etc.). Users' personal information, social networks, geographic locations, user-generated data, and user operation logs have multiplied. If we want to mine this user data, SQL databases are no longer suitable for these applications, but the development of NoSQL databases can handle these large data well.

Examples
Social networks:
Separate records: UserID, first_name,last_name, age, gender,...
Task: Find all friends of friends of friends of ... friends of a given user.
Wikipedia pages:
Combination of structured and unstructured data
Task: Retrieve all pages regarding athletics of Summer Olympic before 1950.
RDBMS vs NoSQL
RDBMS
- Highly organized structured data
- Structured Query Language (SQL)
- Data and relationships are stored in separate tables.
- Data Manipulation Language, Data Definition Language
- Strict consistency
- Basic transactions
NoSQL
- Represents "Not Only SQL"
- No declarative query language
- No predefined schema
- Key-value pair storage, column storage, document storage, graph databases
- Eventual consistency, not ACID properties
- Unstructured and unpredictable data
- CAP theorem
- High performance, high availability, and scalability

A Brief History of NoSQL
The term NoSQL first appeared in 1998. It was a lightweight, open-source relational database developed by Carlo Strozzi that did not provide SQL functionality.
In 2009, Johan Oskarsson of Last.fm initiated a discussion about distributed open-source databases
The "no:sql(east)" conference held in Atlanta in 2009 was a milestone, with the slogan "select fun, profit from real_world where relational=false;". Therefore, the most common interpretation of NoSQL is "non-relational", emphasizing the advantages of Key-Value Stores and document databases, rather than simply opposing RDBMS.
CAP theorem
In computer science, the CAP theorem, also known as Brewer's theorem, states that for a distributed computing system, it is impossible to simultaneously satisfy the following three points:
- Consistency(All nodes have the same data at the same time)
- Availability(Guarantees that every request receives a response, whether successful or failed)
- Partition tolerance(The loss or failure of any information in the system will not affect the continued operation of the system)
The core of the CAP theorem is that a distributed system cannot simultaneously well satisfy the three requirements of consistency, availability, and partition tolerance; at most, it can only well satisfy two at the same time.
Therefore, according to the CAP theorem, NoSQL databases are divided into three categories: systems that satisfy the CA principle, systems that satisfy the CP principle, and systems that satisfy the AP principle:
- CA - single-point cluster, systems that satisfy consistency and availability, usually not very powerful in scalability.
- CP - systems that satisfy consistency and partition tolerance, usually not particularly high in performance.
- AP - systems that satisfy availability and partition tolerance, usually may have lower requirements for consistency.

Advantages/Disadvantages of NoSQL
Advantages:
- - High scalability
- - Distributed computing
- - Low cost
- - Flexible architecture, semi-structured data
- - No complex relationships
Disadvantages:
- - No standardization
- - Limited query functionality (so far)
- - Eventual consistency is unintuitive programming
BASE
BASE: Basically Available, Soft-state, Eventually Consistent. Defined by Eric Brewer.
The core of the CAP theorem is that a distributed system cannot simultaneously satisfy consistency, availability, and partition tolerance well; at most, it can only satisfy two of them well at the same time.
BASE is the weak requirement principle that NoSQL databases usually have for availability and consistency:
- Basically Available -- Basically available
- Soft-state -- Soft state/flexible transaction. "Soft state" can be understood as "connectionless", while "Hard state" is "connection-oriented"
- Eventually Consistency -- Eventual consistency, which is also the ultimate goal of ACID.
ACID vs BASE
| ACID | BASE |
|---|---|
| Atomicity (Atomicity) | Basically available (Basically Available) |
| Consistency (Consistency) | Soft state/flexible transaction (Soft state) |
| Isolation (Isolation) | Eventual consistency (Eventual consistency) |
| Durability (Durable) |
NoSQL Database Classification
| Type | Some representatives |
Hbase
Cassandra
Hypertable
As the name implies, data is stored by column. The biggest feature is that it is convenient for storing structured and semi-structured data, convenient for data compression, and has a great IO advantage for queries on a certain column or a few columns.
Document store
MongoDB
CouchDB
Document storage generally uses a JSON-like format to store content, and the stored content is document-based. This makes it possible to create indexes on certain fields and implement some functions of relational databases.
Key-value store
Tokyo Cabinet / Tyrant
Berkeley DB
MemcacheDB
Redis
The value can be quickly queried by key. Generally speaking, storage does not care about the format of the value and accepts it as is. (Redis includes other functions)
Graph store
Neo4J
FlockDB
The best storage for graph relationships. Using traditional relational databases to solve this would result in low performance and inconvenient design and use.
Object store
db4o
Versant
Operate the database through syntax similar to object-oriented languages, and access data through objects.
XML database
Berkeley DB XML
BaseX
Efficiently store XML data and support XML internal query syntax such as XQuery, XPath.
Who is using it
Many companies are already using NoSQL:- Mozilla
- Adobe
- Foursquare
- Digg
- McGraw-Hill Education
- Vermont Public Radio