Back to Articles
EngineeringTechnologyDistributed systemsetcd

etcd: The Small Database That Keeps Distributed Systems in Sync

CIBIVISHNU A C
CIBIVISHNU A C
6 min2 views

Imagine you have a system running on three different servers.

One server knows that a service is running. Another server needs to know the same thing. A third server also needs that information. Now imagine one of the servers suddenly crashes.

How do all the servers agree on what is true?

This is the kind of problem etcd is designed to solve.

etcd is a distributed key-value store designed to hold important data used for coordination in distributed systems. It is strongly consistent and uses the Raft consensus algorithm to keep multiple servers in agreement.

What Is etcd?

At its simplest, etcd stores data as key-value pairs.

For example:

1/service/api/port → 8080
2/service/api/status → running
3/database/primary → server-2

You can think of it like a small, highly reliable dictionary shared between multiple machines.

But etcd is different from a normal key-value database.

Its main purpose is not to store millions of user records or large files. Instead, it stores small but important pieces of information that distributed systems need to agree on. The etcd documentation describes it as a coordination service designed for relatively small amounts of data.

Why Do We Need etcd?

Consider a simple application with three servers:

1​ ┌─────────────┐
2 │ Application │
3 └──────┬──────┘
4 │
5 ┌─────────┼─────────┐
6 ▼ ▼ ▼
7 Server 1 Server 2 Server 3

Suppose Server 1 is the leader of your application.

The other servers need to know: "Who is the current leader?"

You could store this information in a normal database.

But now another problem appears.

What happens if the database itself goes down?

You have simply moved the problem somewhere else.

etcd solves this by running as a cluster of multiple nodes.

1​ etcd cluster
2 ┌─────────────┐
3 │ Leader 1 │
4 └─────┬───────┘
5 │
6 ┌────────────────┴──────────────┐
7 ▼ ▼
8 ┌──────────────────┐ ┌─────────────────┐
9 │ Node 2 Follower │ │ Node 3 Follower │
10 └──────────────────┘ └─────────────────┘

The nodes communicate with each other and maintain the same state.

If one node fails, the remaining nodes can continue operating as long as they still have a majority, or quorum.

The Important Part: Raft

This is where etcd becomes interesting.

etcd uses an algorithm called Raft to keep its nodes synchronized. Raft provides leader election and replicated logs so that multiple machines can maintain the same state.

Imagine three etcd nodes. When a change happens, the leader receives the request and replicates that change to the other nodes.

For example:

1Client
2 │
3 │ PUT /database/primary = server-2
4 ▼
5Node 1 Leader
6 │
7 ├──────────► Node 2
8 │ Follower
9 │
10 └──────────► Node 3
11 Follower

Once enough members have accepted the change, it can be committed.

This means the cluster doesn't simply trust one machine. Multiple machines participate in maintaining the state.

What Happens When the Leader Dies?

This is one of the most useful parts of etcd.

Suppose we start with:

1Node 1 → Leader
2Node 2 → Follower
3Node 3 → Follower

Now Node 1 suddenly crashes.

The remaining nodes notice that the leader is no longer responding.

1Node 1 → DOWN
2Node 2 → ?
3Node 3 → ?

Node 2 and Node 3 can hold an election.

One of them becomes the new leader:

1Node 1 → DOWN
2Node 2 → Leader
3Node 3 → Follower

The application does not need to manually choose the new leader.

This is one of the reasons etcd is useful for distributed systems.

What Can You Store in etcd?

You can store simple values such as:

1/config/database/host → db.example.com
2/config/database/port → 5432
3/service/api/instance-1 → 10.0.0.10
4/service/api/instance-2 → 10.0.0.11
5/leader → instance-2

The keys can be organized using paths, making the data feel somewhat like a directory structure.

For example:

1/config
2 ├── database
3 │
4 ├── host
5 │
6 └── port
7 │
8 └── application
9 ├── timeout
10 └── workers

This makes etcd useful for storing configuration and service-related information.

Watching for Changes

One particularly useful feature is watching.

Imagine an application wants to know whenever its configuration changes.

Instead of constantly asking:

1"Did the configuration change?"
2"Did the configuration change?"
3"Did the configuration change?"

the application can watch a key.

For example:

1/config/database/host

If the value changes:

1db-old.example.com
2 ↓
3db-new.example.com

etcd can notify the application that something changed.

This allows applications to react to configuration or state changes without continuously polling the database.

etcd and Kubernetes

One of the most well-known users of etcd is Kubernetes.

A Kubernetes cluster has a huge amount of information that needs to be coordinated.

For example:

  • Which pods exist?
  • Which nodes are available?
  • What deployments should be running?
  • What configuration has been applied?
  • What is the current state of the cluster?

Kubernetes uses etcd as the backing store for its cluster state.

You can think of it roughly like this:

1​ Kubernetes
2 │
3 ▼
4 ┌─────────┐
5 │ etcd │
6 └────┬────┘
7 │
8 ┌──────────┼──────────┐
9 ▼ ▼ ▼
10 Nodes Pods Services

The important idea is that etcd provides Kubernetes with a consistent source of truth for cluster state.

How Does an Application Talk to etcd?

etcd provides a gRPC-based API, and the project also provides etcdctl, a command-line client. The standard ports are 2379 for client communication and 2380 for communication between etcd members.

For example, you can store a value:

1etcdctl put mykey "hello"

And retrieve it:

1etcdctl get mykey

The result is simply:

1mykeyhello

It looks simple because the basic operation really is simple.

The complexity is hidden underneath, in the distributed consensus and replication system.

Is etcd a Normal Database?

Not really.

You can think of etcd as a database because it stores persistent data, but it is designed for a different job.

A traditional application database might contain:

1Users
2Orders
3Products
4Payments
5Messages

etcd is better suited for things like:

1Who is the leader?
2What configuration should I use?
3Which service instances are available?
4What is the current cluster state?

It is designed for coordination rather than large-scale application data storage.

Why Not Just Use Redis?

Redis is excellent for caching, fast data access, queues, and many other workloads.

But etcd has a different focus.

The key difference is strong consistency and distributed consensus.

When several machines need to agree on an important piece of state, etcd's Raft-based design makes it particularly useful.

This doesn't mean etcd is a replacement for Redis or PostgreSQL. They solve different problems.

Conclusion

etcd may look like a simple key-value store from the outside, but underneath it is solving one of the hardest problems in distributed systems:

How can multiple machines agree on the same information even when some machines fail?

By combining a simple key-value interface with Raft consensus, replication, leader election, and watch functionality, etcd provides a reliable coordination layer for distributed applications.

And that's why, despite its simple API, etcd is such an important piece of infrastructure behind systems like Kubernetes.