Back to Articles
EngineeringRaftConsensusTechnology

Raft Consensus: How Distributed Servers Agree

CIBIVISHNU A C
CIBIVISHNU A C
7 min1 view

Imagine you have three servers working together:

1Server 1 Server 2 Server 3
2 │ │ │
3 └────────────┼───────────┘
4 │
5 Shared application

All three servers need to agree on the same information.

For example, they might need to agree on:

"The current leader is Server 1."

But what happens if Server 1 suddenly crashes?

Who becomes the new leader? How do the other servers agree on the change?

This is where consensus algorithms come in.

One of the most popular consensus algorithms is Raft.

What Is Raft?

Raft is a consensus algorithm that allows a group of servers to agree on the same state, even when some servers fail.

It is commonly used in distributed systems where consistency is important.

The basic idea is simple:

Multiple servers work together, but they behave as if there is one reliable source of truth.

Raft does this using three main concepts:

  • Leader
  • Followers
  • Log replication

Let's understand them with a simple example.

Leader and Followers

Imagine we have three Raft servers:

1 Raft Cluster
2 
3 ┌───────────┐
4 │ Server 1 │
5 │ Leader │
6 └─────┬─────┘
7 │
8 ┌─────┴─────┐
9 ▼ ▼
10 ┌──────────┐ ┌──────────┐
11 │ Server 2 │ │ Server 3 │
12 │ Follower │ │ Follower │
13 └──────────┘ └──────────┘

At any given time, one server acts as the leader.

The other servers are followers.

The leader is responsible for handling changes and making sure those changes are replicated to the followers.

For example, suppose an application wants to change:

1/config/port = 8080

The request goes to the leader.

The leader then replicates that change to the followers.

How Is the Leader Chosen?

When a Raft cluster starts, there is initially no leader.

The servers wait for a short, randomized amount of time.

Eventually, one server's timer expires first.

That server becomes a candidate and asks the other servers to vote for it.

For example:

1Server 1 → "I want to become leader. Will you vote for me?"
2 
3Server 2 → "Yes"
4 
5Server 3 → "Yes"

Server 1 now has a majority of votes, so it becomes the leader.

1Server 1 → Leader
2Server 2 → Follower
3Server 3 → Follower

This process is called a leader election.

What Happens When the Leader Crashes?

This is where Raft becomes really useful.

Suppose Server 1 is the leader:

1 Server 1
2 Leader
3 X
4 CRASHED
5 
6Server 2 Server 3
7Follower Follower

The followers stop receiving messages from the leader.

After waiting for their election timeout, one of them starts an election.

For example:

1Server 2 → "I want to become leader."
2 
3Server 3 → "I vote for Server 2."

Server 2 receives a majority of the votes.

Now:

1Server 1 → DOWN
2 
3Server 2 → Leader
4Server 3 → Follower

The system has automatically recovered from the leader failure.

What Is a Majority?

Raft relies heavily on the idea of a majority, also called a quorum.

For three servers:

13 servers
2Majority = 2

For five servers:

15 servers
2Majority = 3

The cluster can continue making progress as long as a majority of servers are available.

For example, with three servers:

1Server 1 → DOWN
2Server 2 → UP
3Server 3 → UP
4 
52 out of 3 are available
6→ Majority exists
7→ Cluster can continue

But if two servers fail:

1Server 1 → DOWN
2Server 2 → DOWN
3Server 3 → UP
4 
51 out of 3 are available
6→ No majority
7→ Cluster cannot safely commit new changes

This is important because Raft would rather stop making changes than allow different parts of the cluster to disagree.

How Does Raft Replicate Data?

Now let's say the leader receives a request:

1SET /app/version = 2

The leader first adds the operation to its log.

1Leader
2 
3Log:
41. SET /app/version = 1
52. SET /app/version = 2

It then sends the new log entry to the followers.

1 Leader
2 │
3 SET /app/version = 2
4 ┌────┴────┐
5 ▼ ▼
6 Server 2 Server 3

The followers add the same entry to their logs.

Once a majority of servers have stored the entry, the leader can commit it.

1Server 1 → Entry stored
2Server 2 → Entry stored
3Server 3 → Entry stored
4 
5Majority reached
6 ↓
7 COMMITTED

The committed change becomes part of the cluster's agreed state.

Why Use a Log?

The log is one of the most important parts of Raft.

Instead of simply saying:

1Current state = X

Raft keeps a history of operations:

11. Create user
22. Update configuration
33. Start service
44. Change leader information
55. Update configuration

Each server tries to maintain the same sequence of operations.

If the logs are consistent, the servers can arrive at the same state by applying those operations in the same order.

Think of it like a recipe.

If three people follow the same recipe, in the same order, they should end up with the same result.

What If a Follower Misses an Update?

Suppose the leader sends an update to Server 2, but Server 2 temporarily loses its network connection.

1Leader
2 │
3 ├──────────X────────── Server 2
4 │
5 └───────────────────── Server 3

Server 3 receives the update, but Server 2 doesn't.

Later, Server 2 reconnects.

The leader can send the missing log entries to Server 2.

1Server 2:
2 
3Before:
41. A
52. B
6 
7After synchronization:
81. A
92. B
103. C
114. D

This allows the follower to catch up with the rest of the cluster.

What Happens During a Network Failure?

Now imagine the network splits the cluster.

1 Network failure
2 
3 Server 1 │ Server 2 Server 3
4 │
5 X

Server 1 is separated from Servers 2 and 3.

Servers 2 and 3 still have a majority:

1Server 2 + Server 3 = 2/3

They can continue operating and elect a leader.

Server 1, however, cannot safely commit new changes because it does not have a majority.

This prevents two isolated parts of the cluster from independently making decisions that could later conflict.

When the network connection comes back, the servers synchronize again.

Why Is Raft Called a Consensus Algorithm?

The word consensus simply means agreement.

Imagine three friends trying to decide where to eat.

If everyone chooses a different restaurant, there is no consensus.

But if two out of three agree on one restaurant, they have a majority decision.

Raft applies a similar idea to computers, but with strict rules for elections, logs, replication, and failures.

The goal is to make sure that the cluster agrees on a consistent sequence of operations.

Raft in the Real World

Raft isn't usually something you interact with directly.

Instead, it is used inside distributed systems.

For example, etcd uses Raft to replicate its data between nodes.

A simplified view looks like this:

1 Application
2 │
3 ▼
4 etcd
5 │
6 ┌───────┴───────┐
7 ▼ ▼
8 Raft Leader Followers
9 │ │
10 └───────┬───────┘
11 ▼
12 Replicated state

When an etcd cluster needs to agree on a change, Raft handles the leader election and replication underneath.

This is one reason understanding Raft makes it much easier to understand how systems like etcd work.

Conclusion

Raft solves a very important problem in distributed systems:

How can multiple servers agree on the same state when servers and networks can fail?

It does this using a relatively simple set of ideas:

  • Elect a leader.
  • Let the leader coordinate changes.
  • Replicate changes through a log.
  • Commit changes only after reaching a majority.
  • Elect a new leader when the current leader fails.
  • Synchronize followers that fall behind.

Once you understand these ideas, distributed systems like etcd become much easier to visualize.

Behind what looks like a simple key-value operation, there can be multiple servers communicating, voting, replicating logs, and making sure everyone agrees on what happened.