← Back to the section

In languages like C, you have to free memory manually: forget to call free and you get a leak, free it twice and you get a crash. In Java this is handled by the garbage collector (GC): it finds objects that no one needs anymore and reclaims their memory on its own. Let's look at how this works and which settings let you influence it.

new Order() — every new object lands in Young Young Generation Old Generation roots stack, static tmp tmp tmp tmp tmp session alive objects are created, the young generation fills up minor GC: stop-the-world pause, walk of references from the roots 5 of 6 are unreachable — memory is free, the survivor moves to Old

An object is alive while there is a path of references to it from the roots. The young generation is cleaned often and quickly — almost everything in it is already garbage, and the rare survivor moves to the old one.

Why automatic garbage collection is needed

When you write new Order(), the object is created in a memory area called the heap. As long as there is at least one live reference to the object, it is needed. Once no references remain (for example, the variable went out of scope), the object becomes garbage — it can no longer be reached from running code.

The garbage collector periodically walks through live objects, starting from the "roots" (local variables on the stack, static fields), and everything it cannot reach is considered garbage and freed.

Hence the benefit: an entire class of errors — leaks, double frees, access to already freed memory — simply does not happen, and you can think about the logic instead of who deletes an object and when.

You can check this rule right in the code. WeakReference is a "weak" reference: it lets you look at an object but does not protect it from collection. If no ordinary references are left, after a collection it returns null.

live example

import java.lang.ref.WeakReference;

public class GcDemo {
    record Order(String id) {}

    public static void main(String[] args) throws InterruptedException {
        Order kept = new Order("ord-1");
        Order dropped = new Order("ord-2");

        WeakReference<Order> lookKept = new WeakReference<>(kept);
        WeakReference<Order> lookDropped = new WeakReference<>(dropped);

        dropped = null;                 // no reference from the roots anymore

        System.gc();                    // a request to collect right now
        Thread.sleep(100);

        System.out.println("kept is still referenced: " + lookKept.get());
        System.out.println("dropped is not:           " + lookDropped.get());
    }
}
Run

Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →

kept is referenced from the stack, so the object is alive. The reference to dropped was set to null and the collector took it: the weak reference returns null. System.gc() is here only for the demonstration — you do not write it in normal code, the JVM picks the moment itself. That is the price of automation: sometimes a collection pauses the application, and the timing is not yours to choose.

The heap and generations

Most objects live very briefly: they are created inside a method, do their job and become garbage. The generational model of the heap is built on this observation. The heap is divided into two main parts:

  • Young Generation — all new objects land here. It fills up quickly.
  • Old Generation — objects that survived several collections in the young generation "move" here, that is, those that live a long time.

Correspondingly, there are two kinds of collections:

  • minor GC — collection only in the young generation. It happens often and runs quickly, because most young objects are already garbage and there are few live ones to check.
  • major GC — a collection that reaches the old generation. It happens less often but takes longer.

That collections run on their own, and often, is visible without any logs: the collector's counters are available from the program itself.

live example

import java.lang.management.GarbageCollectorMXBean;
import java.lang.management.ManagementFactory;

public class MinorGcDemo {
    public static void main(String[] args) {
        GarbageCollectorMXBean young = ManagementFactory.getGarbageCollectorMXBeans().get(0);
        System.out.println("young generation collector: " + young.getName());
        System.out.println("collections before the loop: " + young.getCollectionCount());

        byte[] last = null;
        for (int i = 0; i < 200_000; i++) {
            last = new byte[1024];      // garbage: lives for one iteration
        }

        System.out.println("allocated 200,000 arrays of " + last.length + " bytes");
        System.out.println("collections after the loop: " + young.getCollectionCount());
    }
}
Run

Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →

Two hundred thousand one-kilobyte arrays is 200 MB of garbage, but the heap never held that much for a second: each array died on the next iteration, and the young generation was cleaned several times in a row. The name in the first line of the output is the collector that runs for you by default.

Apart from those there is full GC — a collection of the entire heap. For the modern G1 collector this is not a scheduled phase but a sign that the normal mode did not cope: memory ran out, and it stopped the application to sort the whole heap out. A full GC in the logs is always a reason to investigate, not a fact of life.

Stop-the-world pauses

To safely recount references and move objects, the collector sometimes needs to briefly stop all application code. This is the stop-the-world pause: the application freezes and responds to nothing while the GC does its work, then continues.

In plain terms: a librarian cannot count the shelves while readers walk around rearranging books — the entrance is closed for a short time, order is restored, and it opens again. The more books (objects) there are and the longer the cleanup, the more noticeable the delay. The whole difference between collectors is how short and how rare they can make those pauses.

The main collectors and when to use each

Java 21 ships several collectors in a single JVM (HotSpot). The choice is always a trade-off between two quantities:

  • throughput — what fraction of time the machine spends on useful work rather than on collection;
  • latency — how short the stop-the-world pauses are.
CollectorStrengthWhen it fits
Serialminimal overheadsmall applications, little memory, a single core
Parallelmaximum throughputbatch processing where pauses are not critical
G1balance of pauses and throughputthe default, suits most services
ZGC / Shenandoahvery short pauseslarge heaps, requirements for consistently low latency

Serial GC is the simplest: all the work in a single thread and always with a stop-the-world pause. It is good where there is not much data.

Parallel GC (also called the throughput collector) collects garbage using several threads. There are pauses, but in total the least time goes into collection: the maximum CPU share is left to the application. It suits tasks where overall processing speed matters and short stalls are tolerable.

G1 GC (Garbage-First) is the default collector since Java 9. It divides the heap into many small regions and collects first those with the most garbage (hence the name). It tries to keep pauses within a given budget. This is a reasonable balance that suits most server applications with no tuning at all.

ZGC and Shenandoah are collectors focused on minimal pauses. They do most of their work concurrently with the application, barely stopping it, so pauses stay very short even on heaps of tens and hundreds of gigabytes. The price is slightly higher CPU and memory usage. They are needed where noticeable stalls are unacceptable. Shenandoah is not present in every JDK build — check that your build accepts the flag.

A short rule of thumb: by default — G1; need maximum throughput — Parallel; need consistently tiny pauses on a large heap — ZGC.

Key startup parameters

The collector's behavior and the heap size are set by flags when launching java.

Heap size:

# initial heap size 512 MB, maximum 2 GB
java -Xms512m -Xmx2g -jar app.jar
  • -Xms — the initial heap size;
  • -Xmx — the maximum heap size.

A common trick for servers is to set -Xms equal to -Xmx. Then the JVM reserves the entire heap right away and does not spend time gradually expanding it under load.

Choosing a collector:

java -XX:+UseG1GC   -jar app.jar   # G1 (the default anyway)
java -XX:+UseZGC    -jar app.jar   # ZGC, short pauses
java -XX:+UseParallelGC -jar app.jar   # Parallel, maximum throughput

A pause-time goal hint (for G1):

# ask G1 to try to keep pauses around 100 ms
java -XX:MaxGCPauseMillis=100 -jar app.jar

-XX:MaxGCPauseMillis is a goal, not a guarantee. The JVM will try to honor it by balancing region sizes and collection frequency, but it cannot promise an exact value.

Garbage collection logs:

# write GC events to the console
java -Xlog:gc -jar app.jar

# more detail and write to a file
java -Xlog:gc*:file=gc.log:time,uptime -jar app.jar

-Xlog:gc enables the collection event log: you can see when and which collector ran, how long the pause was, and how much memory was freed.

How to choose and not overdo it

The main advice is to start with the default settings. G1 in Java 21 suits most applications well, and a flag tweaked "by eye" makes things worse more often than better.

A sensible order of actions:

  1. Run the application as is (G1 by default).
  2. Set -Xmx to match the actual available memory — this is the most influential parameter.
  3. If there is a performance problem, measure first: enable -Xlog:gc and check whether the issue really is garbage collection rather than your code or the database.
  4. Only if pauses genuinely get in the way should you try a different collector (ZGC for short pauses) or the -XX:MaxGCPauseMillis goal, one change at a time and with before/after measurements.

Long pauses: first look at what lives in the heap

A typical picture: the service answers in a second or two, the GC log shows rare but long full GCs, and after every collection the heap stays almost full. The key clue is exactly the size after a collection: if 3.5 GB out of 4 are occupied, those are live objects the collector cannot free. It walks and repacks almost the whole heap — and gains almost nothing.

Raising -Xmx in this situation only postpones the problem and makes the pauses longer: the number of live objects does not shrink, and there is more to walk. The right first step is a heap dump and an analysis of what holds the memory:

jcmd <pid> GC.heap_dump /tmp/heap.hprof      # take a dump from a running process
# or upfront: -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/dumps

The dump is opened in Eclipse MAT or VisualVM, where you look at the "dominators" — the objects through which the most memory is reachable. There are two usual answers: a leak (a growing static list, unclosed resources, a cache without eviction) is fixed in the code; or an honestly large cache that lacks memory is limited by size and time to live. Only after that, if the live size really is large, do you think about the heap size and a collector with short pauses.

In short

  • The garbage collector frees the memory of objects that no longer have a path of references from the roots on its own — no manual free is needed.
  • The heap is divided into young and old generations; minor GC cleans the young one (often and quickly), major GC cleans the old one (rarely and slowly); a full GC in G1 is already an emergency mode.
  • A stop-the-world pause is a brief stop of the application during collection; the collectors' job is to make pauses shorter and rarer.
  • Serial — for small applications, Parallel — for maximum throughput, G1 — the balance and the default, ZGC and Shenandoah — tiny pauses on large heaps.
  • -Xms/-Xmx set the initial and maximum heap size; -XX:+UseG1GC/-XX:+UseZGC choose the collector; -XX:MaxGCPauseMillis is the pause goal; -Xlog:gc enables logs.
  • Start with the default settings and don't tune the GC until the logs prove the bottleneck is really there.