The model and Cypher cover how to store and how to ask. What is left is what matters in practice: how to design a graph so that queries stay fast, and how it lives in production.
An index on a label finds only the starting node — the entry point. Without it the query scans every node of the label; with it the right node is found at once. After that, in both cases, the traversal follows edges from node to neighbours, and the index makes no difference there.
Design from the queries, not from the entities
A relational schema is usually drawn from the entities: tables, columns, relationships later. In a graph it works the other way round — queries first. The key question is "which traversals must the database answer quickly". The structure grows out of the answers: what becomes a node, what stays a property, where the relationships point.
Node or property. The rule is simple: if you need to walk the relationships by a value, it is a node; if the value is only read together with the entity, it is a property. The city a person lives in is a city property until you need "find everyone in this city": then :City becomes a node with relationships of its own. A property cannot be traversed, a node can.
Relationship direction follows the meaning: BOUGHT goes from person to product. A relationship can be traversed either way — direction does not get in the way of reading, but it makes the model clearer and helps the planner.
Properties on relationships are what the relational model does not give you for free. "Employed since" or "transfer amount" live on the relationship itself, not in a separate join table.
Indexes and uniqueness constraints
Index-free adjacency speeds up the traversal, but every traversal has an entry point — the node it starts from. Finding that starting node by a property (say, Person {email: 'ivan@shop.ru'}) without an index means a full scan of every node with that label. Hence:
- An index on a label property (
CREATE INDEX person_email FOR (p:Person) ON (p.email)) is needed for every property a query uses to anchor itself to a concrete node. It is the first thing to check when a query is slow. - A uniqueness constraint (
CREATE CONSTRAINT person_email_unique FOR (p:Person) REQUIRE p.email IS UNIQUE) not only guarantees uniqueness, it also creates an index. It is critical forMERGE: without a unique keyMERGEis slow and risks creating duplicates under concurrent writes.
Indexes in Neo4j speed up the lookup of the starting node, not the traversal: that reverses the habit from SQL, where an index is needed behind every JOIN.
The difference is visible in plain Java: the label's nodes are a list, the index is a map, and the neighbours are direct references.
live example
import java.util.ArrayList;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
public class GraphEntry {
record Person(String email, List<String> bought) {}
public static void main(String[] args) {
List<Person> label = new ArrayList<>();
Map<String, Person> index = new HashMap<>();
for (int i = 1; i <= 200_000; i++) {
Person person = new Person("user" + i + "@shop.ru",
i == 150_000 ? List.of("book", "mug") : List.of());
label.add(person);
index.put(person.email(), person);
}
int compared = 0;
for (Person person : label) {
compared++;
if (person.email().equals("user150000@shop.ru")) {
break;
}
}
System.out.println("label scan :Person — comparisons: " + compared);
Person start = index.get("user150000@shop.ru");
System.out.println("index lookup — comparisons: 1, neighbours: " + start.bought());
}
}
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
Supernodes: the main trap
The most common performance problem in a graph is a supernode: a node with an enormous number of relationships. Say, a node for the country "Russia" with tens of millions of people attached to it. A traversal through such a node degenerates into scanning all of its relationships, and that kills the speed.
There are several cures: move frequent filters into relationship properties, split the supernode into sub-categories (not "city" but "city district"), or model things so that the traversal never goes through the hot centre.
Operations: memory, backups, scale
- Memory. Neo4j keeps the hot part of the graph in the page cache; the more of the graph fits in memory, the faster the traversal. It is the first resource to watch: the working set should fit into the page cache.
- Backups. The licence has direct consequences here: a hot backup of a live database (
neo4j-admin database backup) and clustering exist only in the paid edition. The free one leaves you a dump of a stopped database (neo4j-admin database dump) — that is a window of downtime you have to plan for. - Scale. Neo4j scales first and foremost vertically (more memory and cores on one machine) and through secondary copies of the database in a cluster — they are read from, not written to. Sharding a graph — cutting a connected graph across machines — is hard by nature (a relationship can cross a shard boundary) and is needed less often than it seems.
Where people stumble
- No index on the starting property — every query then begins with a full node scan. The traversal is fast, the entry is not.
MERGEwithout a uniqueness constraint — slow, and a source of duplicates under concurrent writes.- Pulling in a graph database for a single hierarchy — a category tree or an org chart is solved by recursive SQL in the database you already run; a second database costs more to operate.
- Thinking about sharding right away — almost always too early; vertical growth and secondary copies cover most workloads.
In short
- A graph is designed from the queries: traversals first, then the decision — node, property or relationship.
- The index belongs on the entry point: it finds the starting node and has no effect on the traversal along relationships.
- A uniqueness constraint creates an index and makes
MERGEfast and safe under concurrent writes. - A supernode with a huge number of relationships is the main source of slow traversals; spot it while designing.
- Neo4j grows vertically and through read-only secondary copies; hot backups and clustering are paid-edition only.
What to read next
- Graph data in plain words: recursive SQL or a graph database — whether you need a graph at all.
- Neo4j: where a graph database is used, from recommendations to GraphRAG — recommendations, fraud detection, knowledge graphs.
- Graphs — how breadth-first and depth-first traversals work underneath.