Starting Elasticsearch is not hard. The harder part is making sure it doesn't drown in its own data a month later. This article is about running a cluster under real conditions: controlling index growth, making backups, choosing the right hardware, and knowing when something goes wrong.
The application only knows the alias logs, not the index names. Once the largest primary shard reaches 50 GB or the index turns seven days old, rollover creates the next index and moves the is_write_index flag onto it. From that moment the previous index is read-only and travels down by age: in warm it is merged into a single segment, in cold its replicas are dropped, on day ninety it is deleted.
The problem of growing indices
Imagine you're writing logs into Elasticsearch. At first everything is fine: search is fast, there's plenty of space. Six months later the disk is full, the cluster is slow, three-year-old logs take up space, but no one reads them anymore.
The solution is Index Lifecycle Management (ILM). You describe a policy: how long an index lives on fast hardware, when it moves to slow and cheap storage, when it gets deleted. Elasticsearch runs it automatically.
The policy describes four phases in the life of an index: hot — active writes on fast disks, warm — read-only on SSD, cold — rare reads without replicas, delete — removal. Transitions are driven by index age: below, that is seven, thirty, and ninety days.
There's also a fifth phase — Frozen: data is stored as a snapshot on S3, searchable but read 10-100 times slower. It fits auditing or regulatory compliance.
How to create a policy
PUT /_ilm/policy/logs-policy
{
"policy": {
"phases": {
"hot": {
"actions": {
"rollover": { "max_primary_shard_size": "50gb", "max_age": "7d" }
}
},
"warm": {
"min_age": "7d",
"actions": {
"forcemerge": { "max_num_segments": 1 },
"shrink": { "number_of_shards": 1 }
}
},
"cold": {
"min_age": "30d",
"actions": {
"allocate": { "number_of_replicas": 0 }
}
},
"delete": {
"min_age": "90d",
"actions": { "delete": {} }
}
}
}
}
The max_primary_shard_size threshold is measured on the largest primary shard, not on the whole index: search speed depends on the size of a shard.
ILM works with rollover indices: the application writes to a single alias (logs), and Elasticsearch itself creates logs-000001, logs-000002 once a threshold is reached. This requires a template and a starting index:
PUT /_index_template/logs-template
{
"index_patterns": ["logs-*"],
"template": {
"settings": {
"index.lifecycle.name": "logs-policy",
"index.lifecycle.rollover_alias": "logs"
}
}
}
PUT /logs-000001
{
"aliases": {
"logs": { "is_write_index": true }
}
}
The rollover conditions and phase boundaries are easier to see in action. The program runs a hundred days of writes: while data is plentiful the index closes by size, later by age, and prints which tier each index ended up in.
live example
import java.util.ArrayList;
import java.util.List;
public class IlmRollover {
record Index(String name, int bornDay, double sizeGb, String reason) {}
static String phase(int ageDays) {
if (ageDays >= 90) return "delete";
if (ageDays >= 30) return "cold";
if (ageDays >= 7) return "warm";
return "hot";
}
public static void main(String[] args) {
List<Index> indices = new ArrayList<>();
indices.add(new Index("logs-000001", 0, 0, "writing"));
for (int day = 1; day <= 100; day++) {
int last = indices.size() - 1;
Index writeIndex = indices.get(last);
double size = writeIndex.sizeGb() + (day <= 40 ? 12 : 2);
int age = day - writeIndex.bornDay();
String reason = size >= 50 ? "size" : age >= 7 ? "age" : "writing";
indices.set(last, new Index(writeIndex.name(), writeIndex.bornDay(), size, reason));
if (!reason.equals("writing")) {
String next = String.format("logs-%06d", indices.size() + 1);
indices.add(new Index(next, day, 0, "writing"));
}
}
for (Index index : indices) {
int age = 100 - index.bornDay();
System.out.printf("%s age %3d d %4.0f GB rollover: %-7s tier: %s%n",
index.name(), age, index.sizeGb(), index.reason(), phase(age));
}
}
}
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
Backups through snapshots
Elasticsearch has no equivalent of pg_dump. Instead there's the snapshot: a copy of the data in remote storage (S3, GCS, Azure Blob, or NFS).
A snapshot is incremental: the first time everything is copied, after that only new data. Under the hood Elasticsearch copies the immutable Lucene segment files that aren't already in the repository.
Step 1: register a repository
PUT /_snapshot/s3-backup
{
"type": "s3",
"settings": {
"bucket": "my-es-backups",
"region": "eu-west-1",
"compress": true,
"base_path": "es-cluster-1"
}
}
S3 support is built into the distribution since Elasticsearch 8; on 7.x it was a separate plugin. The nodes need an IAM role with access to the bucket.
Step 2: create a snapshot
PUT /_snapshot/s3-backup/snapshot-2026-06-27?wait_for_completion=false
{
"indices": "products,orders,logs-*",
"include_global_state": false
}
Step 3: restore
POST /_snapshot/s3-backup/snapshot-2026-06-27/_restore
{
"indices": "products",
"rename_pattern": "products",
"rename_replacement": "products-restored",
"include_global_state": false
}
You can't restore on top of an open index — that's why we use renaming. After restoring, we switch the alias.
Automation through SLM
To avoid creating snapshots by hand, there's Snapshot Lifecycle Management (SLM):
PUT /_slm/policy/daily-snapshots
{
"schedule": "0 30 1 * * ?",
"name": "<daily-snap-{now/d}>",
"repository": "s3-backup",
"config": {
"indices": ["products", "orders"],
"include_global_state": false
},
"retention": {
"expire_after": "30d",
"min_count": 5,
"max_count": 50
}
}
A snapshot every day at 01:30, kept for 30 days, at least 5 and at most 50.
Tiered storage: hot / warm / cold / frozen
In large clusters, nodes are split by role: hot data lives on fast expensive hardware, old data on cheap hardware. ILM moves indices between tiers automatically.
| Tier | Hardware | What to store |
|---|---|---|
| Hot | NVMe, lots of RAM | Active writes, the last 1-7 days |
| Warm | SSD | Read-only, 7-30 days |
| Cold | HDD, little RAM, 0 replicas | Rare reads, 30-90 days |
| Frozen | Snapshot on S3, disk as cache | Auditing, very rare reads |
A node's role is set in elasticsearch.yml:
node.roles: [data_hot, data_content]
# or
node.roles: [data_warm]
For small clusters (up to 10 nodes, up to 10 TB) tiering isn't needed — it adds complexity without benefit. It becomes relevant at volumes from 10-20 TB or 50+ nodes.
Force merge: why and when
Each Elasticsearch shard consists of several segments — files on disk. New data is written into new segments. If a hundred of them pile up, ES reads a hundred files for every query, which is slower.
Once an index stops receiving new writes (a rollover has happened), it can be "compacted" into a single segment:
POST /logs-000001/_forcemerge?max_num_segments=1
The effect: reads speed up by 10-30%, and metadata takes less memory.
Important: force merge is a heavy operation — it loads the disk and CPU and can take hours. Don't run it on an actively writing index. ILM does force merge automatically in the warm phase at the right moment.
How to choose the cluster size
JVM heap
Set it to half of the node's RAM. The hard limit is 31 GB: above that Java switches to a different addressing mode and the benefit of a large heap disappears. If there's more data — better to take two nodes with a 31 GB heap each than one with a 64 GB heap.
The ideal node: 64 GB RAM — 31 GB for the heap and the rest for the filesystem cache Lucene relies on.
Shards per node
Rule of thumb: no more than 600-800 shards per node with a 30 GB heap. Each shard adds metadata overhead. A typical first-deployment mistake: creating thousands of indices with five shards each and ending up with 20,000 shards on 10 nodes.
If the cluster is slow while the data is small, the shard count is almost always the reason.
Size of a single shard
The optimal range: 10-50 GB. Smaller means unnecessary overhead, larger means slow search and long merge operations.
For a 1 TB index you need 20-100 primary shards. With two replicas — 60-300 shards in total.
Write throughput
A single node handles roughly 5-20 thousand documents per second (depending on document size and settings). For 100K documents per second you need 5-20 nodes.
Tuning tip: by default refresh_interval=1s, which creates many small segments under a heavy write load. If there are many writes and data freshness isn't critical, you can raise it to 30s — this gives 2-3 times more throughput.
Cluster monitoring
Prometheus exporter
The standard tool is elasticsearch_exporter. It runs as a container next to Elasticsearch, polls _nodes/stats, and exposes metrics in Prometheus format.
Key metrics
| Metric | What it means | When to alert |
|---|---|---|
elasticsearch_cluster_health_status | Cluster status: green / yellow / red | red — immediately, yellow — investigate |
elasticsearch_jvm_memory_used_bytes / max_bytes | Heap usage | Consistently > 85% |
elasticsearch_jvm_gc_collection_seconds_count | Garbage collection frequency | Old GC more than once a minute |
elasticsearch_indices_indexing_index_time_seconds | Indexing time | Rising — the load is building up |
elasticsearch_indices_search_query_time_seconds | Search time | Rising — problems with queries or mapping |
elasticsearch_thread_pool_rejected_count | Rejected tasks | Any value > 0 |
elasticsearch_filesystem_data_available_bytes | Free disk space | Below 15% |
Disk fill thresholds
Elasticsearch reacts to disk fill automatically:
- 85% — stops allocating new shards on this node.
- 90% — starts moving shards to other nodes.
- 95% (flood stage) — all indices on this node are switched to read-only. Writes stop.
Flood stage is an emergency mode. Elasticsearch clears the read_only_allow_delete block itself once node usage drops below the high watermark of 90%, but freeing space is still manual: expand the disk or delete old indices.
Common mistakes
Too many shards. One 100 GB index is better than 100 indices of 1 GB each.
Dynamic fields without limits. If you write JSON with thousands of different keys (attr_color, attr_size, attr_brand_...), Elasticsearch creates a separate field for each. With a million unique fields — memory overflow. The fix: dynamic: false in the mapping and the flattened type for arbitrary attributes.
Large aggregations. A terms aggregation with size: 10000 over billions of documents can kill a node. Use a composite aggregation with paged loading instead.
Disabled _source. You can save space by removing _source from the index. But then you can't update a document or reindex — only a full rebuild from the source. It fits only logs and metrics, where the original data is stored somewhere else.
In short
- ILM describes an index's life policy: hot → warm → cold → delete. For logs, metrics, and events — a must.
- Rollover closes an index by primary shard size or by age; the application always writes to a single alias.
- Snapshots are the only way to back up: incremental, stored in S3 / GCS / Azure, created and expired on a schedule by SLM.
- Hot/warm/cold is needed from about 10 TB of data; in small clusters it's overkill.
- Heap — no more than 31 GB, shards — no more than 600-800 per node, shard size — 10-50 GB.
- Flood stage at 95% disk stops writes; the block clears itself below 90%, but freeing space is still on you.
What to read next
- Elasticsearch fundamentals — how the cluster, shards, and replicas are structured.
- Query DSL and relevance — how to write efficient queries.
- Elasticsearch client code — integration on the application side.
- Search: PostgreSQL FTS or Elasticsearch — when a separate cluster is not needed at all.