A real application almost never keeps its data in one place. Alongside the main database there usually live a cache, a search index, an analytics store, denormalized tables, read models. It's easy to drown in this zoo and lose track of who owns what and which copy to believe when they disagree.
There's a simple frame that instantly brings order: split every store into two kinds — the system of record and derived data.
The application writes to one place — the system of record. From there the change fans out to the derived sets as a single stream, so the cache, the index, and the read model share one order of updates. Losing a derived set loses no data: you throw it away and derive it from the source again. Losing the source itself is losing data.
The source of truth and its reflections
The system of record (the source of truth) is the store where a fact lives exactly once and in authoritative form. Usually that's your normalized database: a user placed an order — the order is written here. And if the data in the source ever diverges from the data in any other store, the source is by definition right.
Derived data is the result of transforming data from the source, which can always be recomputed. Almost everything else falls here:
- a cache — a fast copy of part of the data;
- a search index (Elasticsearch, GIN) — the same data repacked for text search;
- a materialized view or a denormalized field — a precomputed result;
- a read model in CQRS — data laid out exactly for a specific screen;
- an analytics store (ClickHouse) — the same events, but shaped for scans.
The defining trait of derived data is redundancy: it duplicates what already exists in the source, for the sake of read speed.
The key property: derived data is disposable
From this definition follows a conclusion that changes how you treat data: derived data is disposable. Lost the cache, the index, or the read model? No disaster: rebuild it from the source, and the system is good as new. But lose the system of record itself — now that's a real disaster, there's nothing to restore from.
The rest follows: back up and protect the source of truth first. Every derived set must have a from-scratch rebuild procedure (reindex, warm the cache, rebuild the projection).
The three ways to produce derived data
How does data flow from the source into derived sets? There are three processing modes, and it pays to tell them apart:
- Online (a service). Waits for a request and answers as fast as possible; what matters is response time and availability. This is how the application itself works. But online is no good for producing large derived sets.
- Batch. Takes a large set of data as input, grinds on it for minutes or hours, produces a result; what matters here is throughput, and it runs on a schedule (a nightly reindex, a mart recompute). The big-data classic is MapReduce/Spark; in an ordinary backend it's simpler — a queue-based background worker.
- Stream. In between: reacts to events as they arrive, with low latency, but processes a stream of changes rather than the whole set at once. This is how derived data is kept fresh in near-real time.
The discipline: derive, don't dual-write
The main mistake with derived data is writing to it directly, bypassing the source. The classic example is a dual write: the application writes to both the database and the cache (or the search index) as two separate operations. Sooner or later one operation succeeds and the other fails — and the data silently diverges.
The right principle: write only to the system of record, and derive everything else from it — as a single stream of changes. Then all the derived sets share one order of updates, they're consistent with each other, and they rebuild easily. Technically this is done with events and the outbox pattern, or via CDC (change data capture — reading the database's change log). The same principle underlies event sourcing: the source of truth is the stream of events, and everything else (projections, read models) is derived from it.
The difference fits in a dozen lines: the source is a map of orders, the change stream is a list of records. First the cache is written directly, then the same cache is thrown away and derived from the stream again.
live example
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
public class DerivedData {
record Change(String orderId, String status) {}
static final Map<String, String> source = new LinkedHashMap<>();
static final List<Change> changes = new ArrayList<>();
public static void main(String[] args) {
write("order-1", "NEW");
write("order-2", "NEW");
write("order-1", "PAID");
Map<String, String> cache = derive();
cache.put("order-1", "SHIPPED");
System.out.println("source: " + source);
System.out.println("cache: " + cache);
System.out.println("match: " + cache.equals(source));
Map<String, String> rebuilt = derive();
System.out.println("rebuilt: " + rebuilt);
System.out.println("match: " + rebuilt.equals(source));
}
static void write(String orderId, String status) {
source.put(orderId, status);
changes.add(new Change(orderId, status));
}
static Map<String, String> derive() {
Map<String, String> view = new LinkedHashMap<>();
for (Change c : changes) {
view.put(c.orderId(), c.status());
}
return view;
}
}
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
The write that bypassed the source failed nowhere — the divergence shows up only in the output: the cache says SHIPPED, the source says PAID. The copy derived from the stream always matches the source.
Where this applies
The moment a cache, a search index, or a separate mart appears next to the database — you're already managing derived data, whether you called it that or not. Here's where people stumble:
- Dual-writing to the database and the cache/index — two operations without a shared transaction will sooner or later diverge.
- Having no rebuild procedure for a derived set — and "reindex from scratch" suddenly turns out to be impossible, while losing the index becomes an incident out of nowhere.
- Treating a read model or a denormalized field as a "second truth" — and on a discrepancy starting to guess who's right.
In short
- The system of record is the store where a fact lives exactly once; on a discrepancy it is the one that's right.
- Derived data (cache, index, materialized view, read model, mart) is redundant: it is derived from the source.
- A derivative is disposable: throw it away and rebuild it. Only the source is irreplaceable — back that up first.
- Every derived set needs a from-scratch rebuild procedure, and one that has been run at least once.
- Three modes: online answers a request, batch rebuilds the whole set on a schedule (reliable but lagging), stream keeps it fresh from the change stream.
- Write to the source only; update derivatives from a single stream of changes (events/outbox/CDC), not with a second operation alongside.
What to read next
- The read model in CQRS — derived data shaped for a specific screen.
- Stream processing — CDC and time in streams: how the change stream itself works.
- Cache invalidation — keeping a derived copy fresh.
- Event sourcing — a stream of events as the source of truth.