Zookeeper's leader election has two phases: one is leader election during server startup, and the other is when the leader server goes down during operation. Before analyzing the election principle, let's first introduce a few important parameters.

  • Server ID (myid): the larger the number, the greater the weight in the election algorithm.
  • Transaction ID (zxid): the larger the value, the newer the data, and the greater the weight.
  • Logical clock (epoch-logicalclock): the logical clock values are the same in the same round of voting, and the value increases after each vote is completed.

Election status:

  • LOOKING: Campaigning state
  • FOLLOWING: Follower state, synchronizes leader state, participates in voting.
  • OBSERVING: Observer state, synchronizes leader state, does not participate in voting.
  • LEADING: Leader state

1. Leader election during server startup

When each node starts, it is in the LOOKING (observing) state, and then the main election process begins. Here we take a cluster composed of three machines as an example. When the first server, server1, starts, leader election cannot be performed. When the second server, server2, starts, the two machines can communicate with each other and enter the leader election process.

  • (1) Each server sends out a vote. Since it is the initial situation, server1 and server2 both vote for themselves as the leader server. Each vote contains the recommended server's myid, zxid, and epoch, represented by (myid, zxid). At this time, server1's vote is (1,0), server2's vote is (2,0), and then each sends its own vote to other machines in the cluster.

  • (2) Receive votes from each server. After each server in the cluster receives a vote, it first checks the validity of the vote, such as checking whether it is this round's vote (epoch) and whether it comes from a server in the LOOKING state.

  • (3) Process votes separately. For each vote, the server needs to compare the votes of other servers with its own vote. The comparison rules are as follows:

    • a. Compare epoch first.
    • b. Check zxid. The server with the larger zxid is preferred as the leader.
    • c. If the zxid is the same, then compare myid. The server with the larger myid becomes the leader server.
  • (4) Count votes. After each vote, the server tallies the voting information and determines whether more than half of the machines have received the same voting information. Both server1 and server2 determine that two machines in the cluster have accepted the vote (2,0), so server2 has been selected as the leader node.

  • (5) Change server status. Once the leader is determined, each server updates its own status accordingly. If it is a follower, it changes to FOLLOWING; if it is the leader, it changes to LEADING. At this time, server3 starts up and directly joins, changing itself to FOLLOWING.

2. Leader election during operation

When the leader server in the cluster crashes or becomes unavailable, the entire cluster cannot provide services to the outside and enters a new round of leader election.

  • (1) Change status. After the leader crashes, other non-Observer servers change their own server status to LOOKING.
  • (2) Each server sends out a vote. During operation, the zxid on each server may be different.
  • (3) Process votes. Rules are the same as in the startup process.
  • (4) Count votes. Same as the startup process.
  • (5) Change server status. Same as the startup process.