Ceph data protection: replication vs erasure coding
Once you’ve decided to run Ceph, one of the first real decisions per storage pool is how to protect the data: replication or erasure coding. They reach the same goal — surviving disk and node failures — but the trade-offs are opposite, and picking the wrong one for a workload is expensive to undo later.
Replication — keep full copies
Replication stores N complete copies of every object across different
failure domains. The common default is size = 3: three copies, typically on
three different hosts, with min_size = 2 (the pool keeps serving as long as two
copies are available).
Pros
- Simple and robust. The logic is trivial; there’s nothing to decode.
- Fast recovery. To heal a lost copy, Ceph just copies an existing one.
- Good performance, especially for small and random I/O — no encode/decode on the data path.
The cost
- Storage overhead.
size = 3means you store 3× the data: only ~33% of raw capacity is usable. That’s a lot of disk to buy.
Erasure coding — data plus parity
Erasure coding (EC) splits each object into k data chunks plus m parity
chunks — written as “k+m”. A 4+2 profile, for example, writes 6 chunks and
can lose any 2 of them without losing data.
Pros
- Space-efficient. With
4+2, overhead is 1.5× instead of 3× — roughly ~67% usable capacity. For large clusters that’s a dramatic saving.
The cost
- CPU. Every write encodes, and reads during degradation decode — that work isn’t free.
- Latency, especially for small and random writes, and partial reads.
- Slower, heavier recovery — rebuilding a chunk means reading several others.
- Needs enough failure domains. A
k+mprofile wants at leastk+mdistinct hosts (or whatever your failure domain is) to place chunks safely.
Head to head
| Replication (3×) | Erasure coding (4+2) | |
|---|---|---|
| Usable capacity | ~33% | ~67% |
| Failures tolerated | 2 copies lost | any 2 chunks lost |
| Small / random I/O | fast | slower |
| CPU cost | low | higher |
| Recovery | fast, simple | slower, heavier |
| Best for | hot, random, block | cold/warm, large, sequential |
How to choose
A workable rule of thumb:
- Replication for hot, latency-sensitive, random I/O — block storage (RBD) backing VMs and databases, and anything where tail latency matters more than disk cost.
- Erasure coding for capacity-driven, large, sequential data — object storage (RGW), backups, archives, media — where space efficiency wins and the extra latency is acceptable.
And you don’t have to pick once: Ceph lets you set the protection per pool, so a single cluster can run replicated pools for VMs and EC pools for object storage side by side.
The honest part
Erasure coding looks like free capacity, but it isn’t free — you pay in CPU, latency, and recovery time, and those costs show up exactly when the cluster is already under stress (a failure). The space saving is real and often worth it for the right data, but match it to the workload. As always with Ceph: decide by what the workload actually does, and verify under realistic load — not by the headline capacity number.
Designing pools, choosing a profile, or rebalancing a cluster that’s running out of space? Linux administration & Ceph is part of what we do.
