Redis is an in-memory store. You'd think that if the process crashes, the data is gone. In reality it saves to disk, survives a server failure and scales across machines. Let's look at how that works and which settings to pick.
The key space is cut into 16384 slots, and every slot is owned by one node. The slot number comes from the key name, so the client knows the address up front and talks to a single node instead of polling all of them. When a node stops responding, its slots move to the replica as a whole — the range stays the same, only the machine behind it changes.
Persistence: RDB and AOF
By default Redis keeps data only in memory: a restart loses everything. Two mechanisms prevent that.
RDB (Redis Database) is a snapshot: Redis writes a full binary image of the data to disk (dump.rdb), on a schedule or on demand with SAVE / BGSAVE. BGSAVE spawns a child process and doesn't block Redis.
# Save a snapshot manually (asynchronous, non-blocking)
BGSAVE
# Check when the last save happened
LASTSAVE
Configuration in redis.conf:
# Save a snapshot if at least 1000 keys changed within 60 seconds
save 60 1000
# Snapshot file
dbfilename dump.rdb
dir /var/lib/redis
Upside of RDB: a compact file and a fast restart. Downside: a crash between snapshots loses everything written since the last one.
AOF (Append Only File) is a command log: every write command is appended to a file (appendonly.aof), and on restart Redis replays it.
# Enable AOF
appendonly yes
appendfilename "appendonly.aof"
# fsync policy (flush to disk):
# always — after every command (maximum durability, slower)
# everysec — once per second (a good balance, at most 1 s of data lost)
# no — the OS decides (fast, unreliable)
appendfsync everysec
Upside of AOF: you lose at most one second of data. Downside: the file grows and startup is slower. Redis periodically compacts the AOF (BGREWRITEAOF), dropping redundant commands. Since Redis 7.0 the log is not one file: a base snapshot plus incremental files live in appendonlydir, and appendfilename only sets the base of their names.
Short formula: for important data enable both RDB and AOF — Redis works with both, and on startup AOF takes priority.
Memory and eviction policies
Redis lives in RAM, so it needs a limit — otherwise it eats all the server's memory.
# Memory limit (for example, 512 megabytes)
maxmemory 512mb
# What to do when the limit is reached — the eviction policy
maxmemory-policy allkeys-lru
Eviction policies — what Redis does when there is no memory left (the default is noeviction):
| Policy | Behavior |
|---|---|
noeviction | Reject writes with an error (doesn't touch the data) |
allkeys-lru | Evict any keys on a "least recently used" basis |
volatile-lru | Evict only keys with a TTL, by LRU |
allkeys-lfu | Evict any keys on a "least frequently used" basis |
volatile-lfu | The same, but only among keys with a TTL |
allkeys-random | Evict any keys at random |
volatile-ttl | Evict keys with the smallest remaining TTL |
volatile-random | Evict keys with a TTL at random |
When to choose what:
- Redis as a cache (data can be lost) →
allkeys-lru. The most common choice. - The cache has clearly "hot" keys →
allkeys-lfu: it looks at how often a key is used, not at when. - Redis as a cache, but only some keys have a TTL →
volatile-lru. - Redis as a primary store (data must not be lost) →
noevictionplus memory control in the application.
Next to a value Redis keeps two marks: when the key was last used and how often. LRU looks at the first, LFU at the second:
live example
import java.util.LinkedHashMap;
import java.util.Map;
public class EvictionDemo {
record Stat(long lastUsed, long hits) {}
public static void main(String[] args) {
Map<String, Stat> keys = new LinkedHashMap<>();
String[] calls = {"user:42", "user:42", "user:42", "user:42", "user:42",
"report:jan", "report:feb", "report:jan"};
for (int tick = 1; tick <= calls.length; tick++) {
Stat was = keys.getOrDefault(calls[tick - 1], new Stat(0, 0));
keys.put(calls[tick - 1], new Stat(tick, was.hits() + 1));
}
String lru = null, lfu = null;
for (String key : keys.keySet()) {
if (lru == null || keys.get(key).lastUsed() < keys.get(lru).lastUsed()) lru = key;
if (lfu == null || keys.get(key).hits() < keys.get(lfu).hits()) lfu = key;
}
keys.forEach((key, stat) -> System.out.println(
key + ": hits " + stat.hits() + ", last used at tick " + stat.lastUsed()));
System.out.println("allkeys-lru evicts " + lru);
System.out.println("allkeys-lfu evicts " + lfu);
}
}
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
LRU evicts user:42 — used long ago, though used the most; LFU evicts report:feb and keeps the hot key. Redis counts frequency approximately and decays it over time (lfu-log-factor, lfu-decay-time), otherwise a key popular a year ago would never be evicted.
In Spring Boot @Cacheable knows nothing about eviction: the next access to an evicted key simply recomputes the value.
Replication: primary and replica
Replication keeps several copies of the data on different servers. One server is the primary (accepts writes), the rest are replicas (read-only, synchronized with the primary).
# On the replica in redis.conf (or via the CLI):
REPLICAOF 192.168.1.10 6379
The replica continuously receives a stream of changes from the primary. On the first connection the primary runs BGSAVE, sends the snapshot and then streams the accumulated commands — that's how the replica catches up.
What replication gives you:
- Horizontal read scaling — the application reads from replicas, sparing the primary.
- Backup copy — if the primary goes down, the replica holds the data.
On its own a replica does not become primary on failure — for that you need Sentinel.
Sentinel: fault tolerance
Redis Sentinel is a set of watcher processes that monitor the primary and the replicas. If the primary becomes unavailable, Sentinel votes on a new primary among the replicas and notifies the applications.
# sentinel.conf — minimal configuration
sentinel monitor mymaster 192.168.1.10 6379 2
# "mymaster" — the cluster name
# 2 — the quorum: how many Sentinels must agree that the primary is unavailable
sentinel down-after-milliseconds mymaster 5000
sentinel failover-timeout mymaster 60000
A typical layout: 3 Sentinel nodes (an odd number for the quorum), 1 primary, 1-2 replicas. The application connects to Sentinel and gets the current primary address from it.
In Spring Boot, Sentinel is configured via application.yml:
spring:
data:
redis:
sentinel:
master: mymaster
nodes:
- sentinel1:26379
- sentinel2:26379
- sentinel3:26379
Cluster: sharding
Redis Cluster is needed when the data outgrows one server or the write load is too high. Cluster spreads keys across several primary nodes — this is sharding.
Cluster divides the key space into 16384 hash slots, each primary owning a range.
The slot number is CRC16 of the key name modulo 16384, and the client computes it before sending the command. If the name contains curly braces, only what is inside them is hashed: {user:42}:cart and {user:42}:orders always land in the same slot. That matters because an operation over several keys at once (MGET, a transaction, a script) only works inside one slot.
# Minimal cluster: 3 primaries + 3 replicas (6 nodes)
# Launch via redis-cli:
redis-cli --cluster create \
192.168.1.10:6379 192.168.1.11:6379 192.168.1.12:6379 \
192.168.1.13:6379 192.168.1.14:6379 192.168.1.15:6379 \
--cluster-replicas 1
In Spring Boot, a cluster is configured similarly to Sentinel:
spring:
data:
redis:
cluster:
nodes:
- 192.168.1.10:6379
- 192.168.1.11:6379
- 192.168.1.12:6379
Lettuce (the default driver in Spring Data Redis) handles a cluster transparently: it redirects each request to the right node.
The choice is simple: one server with non-critical restarts is fine with RDB; need fault tolerance — add Sentinel; data doesn't fit one server — go to Cluster, which runs its own failover and is not paired with Sentinel.
Monitoring and common problems
The INFO command
INFO returns detailed statistics about the server:
# General information
INFO
# Memory section only
INFO memory
# Replication statistics only
INFO replication
Key metrics:
used_memory— how much memory is usedconnected_clients— the number of connectionskeyspace_hits/keyspace_misses— cache hits and missesrdb_last_bgsave_status— the status of the last snapshotrole— primary or replica
Slow commands
Redis runs commands in a single thread, so a command whose duration grows with the data size pauses every client at once. The classic one is KEYS *: it walks the whole key space in one indivisible call, seconds of downtime on millions of keys. The replacement is SCAN: a cursor walk in batches, and between calls the server serves everyone else; its result may contain duplicates and has to be repeated until the cursor returns zero. The same trick covers the relatives: SSCAN instead of SMEMBERS over a huge set, FLUSHALL with ASYNC. A long Lua script blocks the server just the same: script atomicity isn't free.
To avoid hunting for such commands by hand, Redis keeps a log of the slow ones:
# Set the threshold in redis.conf (in microseconds, 10000 = 10 ms)
slowlog-log-slower-than 10000
slowlog-max-len 128
# View slow commands
SLOWLOG GET 10
# Iterate over keys without blocking (instead of KEYS)
SCAN 0 MATCH user:* COUNT 100
Big keys
A huge value in a single key (say, a list of a million elements) is a common problem. Such a key:
- takes a long time to serialize into RDB/AOF;
- blocks the server on operations like
DEL(deleting a large list is synchronous).
Delete such keys with UNLINK instead of DEL — freeing the memory moves to the background.
# Find big keys (the built-in scanner)
redis-cli --bigkeys
# Asynchronously delete a big key
UNLINK huge_list_key
In short
RDB— a snapshot to disk, fast startup, possible data loss between snapshots.AOF— a command log, you lose at most a second, the file needs periodic rewriting.maxmemory+allkeys-lru— the standard choice for a cache;allkeys-lfuprotects hot keys,noevictionis for a store where data must not be lost.- Replication gives a backup copy and read scaling; Sentinel adds automatic failover.
- Cluster is needed when a single node runs out of memory or write capacity: 16384 slots, the slot number is CRC16 of the key.
KEYS *blocks the server — useSCAN;DELof a big key also blocks — useUNLINK.