ZooKeeper is a software project of the Apache Software Foundation. It provides open-source distributed configuration services, synchronization services, and naming registration for large-scale distributed computing.

ZooKeeper's architecture achieves high availability through redundant services.

Zookeeper's design goal is to encapsulate those complex and error-prone distributed consistency services into an efficient and reliable set of primitives, and provide them to users through a series of simple and easy-to-use interfaces.

It is a typical solution for distributed data consistency. Distributed applications can implement functions such as data publish/subscribe, load balancing, naming services, distributed coordination/notification, cluster management, Master election, distributed locks, and distributed queues based on it.

Who should read this tutorial?

This tutorial is intended for professional programmers. Through this tutorial, you can learn about Zookeeper applications step by step.

Zookeeper data structure

The namespace provided by Zookeeper is very similar to a standard file system, stored in key-value form. The name key is a series of path elements separated by slashes./Each node in the Zookeeper namespace is identified by a path.

Related CAP theory

CAP theory states that for a distributed computing system, it is impossible to simultaneously satisfy the following three points:
  • Consistency: In a distributed environment, consistency refers to whether data can remain consistent across multiple replicas, which is equivalent to all nodes accessing the same latest data copy. Under the requirement of consistency, when a system performs an update operation in a consistent data state, it should ensure that the system's data remains in a consistent state.
  • Availability:Every request can get a correct response, but it does not guarantee that the data obtained is the latest data.

  • Partition tolerance:A distributed system must still be able to provide services that satisfy consistency and availability in the event of any network partition failure, unless the entire network environment fails.

A distributed system can satisfy at most two of the three requirements: Consistency, Availability, and Partition tolerance.

Among these three basic requirements, at most two can be satisfied at the same time. P is mandatory, so you can only choose between CP and AP. Zookeeper guarantees CP, while Eureka, the registry in the Spring Cloud system, implements AP.

BASE theory

BASE is the abbreviation of Basically Available, Soft-state, and Eventually Consistent.

  • Basically Available:When a distributed system fails, it is allowed to lose some availability (service degradation, page degradation).

  • Soft state:The distributed system is allowed to have an intermediate state. Moreover, the intermediate state does not affect the availability of the system. The intermediate state here means that data updates between different data replicas (data backup nodes) may be delayed, eventually reaching consistency.

  • Eventually consistent:Data replicas achieve consistency after a period of time.

BASE theory is the result of a trade-off between consistency and availability in CAP. The core idea of the theory is: we cannot achieve strong consistency, but each application can use appropriate methods based on its own business characteristics to make the system eventually consistent.

Related resources

Zookeeper official website:https://zookeeper.apache.org/