When several threads touch the same counter, a plain counter++ breaks — not because the operation is wrong, but because it is not atomic. The AtomicInteger class and its relatives solve this without locks — through the hardware instruction CAS.
Both threads read 7. The first CAS goes through, the second is refused: the cell no longer holds the value it expected. The thread does not go to sleep — it re-reads the fresh value and tries again.
Why counter++ is dangerous in multithreaded code
Writing counter++ looks like a single operation, but the processor splits it into three steps:
- Read the current value from memory.
- Add one.
- Write the result back.
If two threads run these three steps interleaved, both can read the same value, increment it, and write it back — the counter ends up one lower than it should be. This is the classic race condition.
One way to guard against it is synchronized. But a lock is an expensive operation: a thread that fails to acquire the monitor goes to sleep and waits to be woken up. For simple operations on a number this is overkill.
CAS — the "compare and swap" instruction
CAS (compare-and-swap) is a single indivisible processor instruction:
"If the current value at this address equals the expected one, replace it with the new one. Otherwise — do nothing, report failure."
Formula: CAS(address, expected, new) → true/false
Because the instruction is atomic in hardware, no other thread can wedge itself between the check and the write. No monitor is needed — no putting threads to sleep and waking them up. The whole java.util.concurrent.atomic package is built on this instruction.
AtomicInteger, AtomicLong, AtomicReference
AtomicInteger and AtomicLong are wrappers around int and long with atomic operations. AtomicReference<V> does the same for a reference: replace a node in a data structure or swap a configuration without locking.
live example
import java.util.concurrent.atomic.AtomicInteger;
import java.util.concurrent.atomic.AtomicReference;
public class AtomicBasics {
public static void main(String[] args) {
AtomicInteger counter = new AtomicInteger(0);
System.out.println("incrementAndGet -> " + counter.incrementAndGet());
System.out.println("getAndIncrement -> " + counter.getAndIncrement());
System.out.println("addAndGet(5) -> " + counter.addAndGet(5));
System.out.println("compareAndSet(7, 20) -> " + counter.compareAndSet(7, 20));
System.out.println("compareAndSet(7, 99) -> " + counter.compareAndSet(7, 99));
AtomicReference<String> config = new AtomicReference<>("v1");
System.out.println("v1 -> v2: " + config.compareAndSet("v1", "v2"));
System.out.println("v1 -> v3: " + config.compareAndSet("v1", "v3"));
}
}
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
incrementAndGet and getAndIncrement differ in what they return: the new value or the previous one. And compareAndSet is a manual CAS: the first call goes through, the counter really does hold 7; the second answers false, it already holds 20. The reference behaves the same way: the second config swap fails — the thread expected v1, and the cell already holds v2. That is protection against overwriting someone else's write.
In methods that take your own function — updateAndGet, accumulateAndGet — CAS runs in a loop: if the swap fails because another thread got there first, it reads the fresh value and tries again. For a plain increment such as incrementAndGet the loop is usually not needed: on x86 the JVM substitutes a single machine instruction for atomic addition.
What a lock-free increment looks like inside
Let us assemble an increment by hand — the very loop the JDK hides inside updateAndGet — counting wasted attempts:
live example
import java.util.concurrent.atomic.AtomicInteger;
import java.util.stream.IntStream;
public class CasLoop {
static final AtomicInteger counter = new AtomicInteger();
static final AtomicInteger retries = new AtomicInteger();
static void increment() {
while (true) {
int current = counter.get();
if (counter.compareAndSet(current, current + 1)) return;
retries.incrementAndGet();
}
}
public static void main(String[] args) {
IntStream.range(0, 4).parallel().forEach(t -> {
for (int i = 0; i < 50_000; i++) increment();
});
System.out.println("counter: " + counter.get());
System.out.println("wasted loop passes: " + retries.get());
}
}
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
The counter always comes out as exactly 200000 — not a single update is lost. The number of wasted passes jumps from run to run: that is contention for the cell. Under low contention CAS succeeds on the first try; under high contention threads spin in this loop without going to sleep. That is exactly what lock-free means: progress is guaranteed for at least one of the threads at any given moment.
The ABA problem
CAS has a subtle weakness: it compares only the value, not the "history" of changes.
Imagine: thread A read the value "X". While it is thinking, thread B changed "X" → "Y" → back to "X". When thread A runs a CAS with the expected "X", the check passes successfully — even though the value has already changed twice in the meantime.
The solution is AtomicStampedReference: it holds a pair (reference + numeric stamp), and CAS checks both. Both reactions to the same swap:
live example
import java.util.concurrent.atomic.AtomicReference;
import java.util.concurrent.atomic.AtomicStampedReference;
public class AbaDemo {
public static void main(String[] args) {
AtomicReference<String> plain = new AtomicReference<>("X");
AtomicStampedReference<String> stamped = new AtomicStampedReference<>("X", 0);
int[] holder = new int[1];
String seen = stamped.get(holder);
plain.set("Y"); plain.set("X");
stamped.set("Y", 1); stamped.set("X", 2);
System.out.println("CAS on value: " + plain.compareAndSet("X", "Z"));
System.out.println("CAS on stamp: " + stamped.compareAndSet(seen, "Z", holder[0], holder[0] + 1));
}
}
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
The plain CAS answers true and notices nothing. The stamped one answers false: the value is the same, but the stamp is already 2, not 0. In application code ABA rarely comes up, but in lock-free queues, where nodes go back to a pool and get reused, it has to be kept in mind.
LongAdder — when contention is high
AtomicLong works great under moderate load. But if dozens of threads continuously increment the same object, they start fighting over it — CAS loops get longer, and cores burn cycles for nothing.
LongAdder solves this differently: it holds an array of cells instead of a single value. Each thread usually works with its own cell, rarely colliding with others. The total is the sum of all cells, returned by sum(). The difference shows up across sixteen parallel tasks:
live example
import java.util.concurrent.atomic.AtomicLong;
import java.util.concurrent.atomic.LongAdder;
import java.util.stream.IntStream;
public class AdderVsAtomic {
static long millis(Runnable step) {
long start = System.nanoTime();
IntStream.range(0, 16).parallel().forEach(t -> {
for (int i = 0; i < 1_000_000; i++) step.run();
});
return (System.nanoTime() - start) / 1_000_000;
}
public static void main(String[] args) {
AtomicLong atomic = new AtomicLong();
LongAdder adder = new LongAdder();
System.out.println("AtomicLong: " + millis(atomic::incrementAndGet) + " ms");
System.out.println("LongAdder: " + millis(adder::increment) + " ms");
System.out.println("totals: " + atomic.get() + " and " + adder.sum());
}
}
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
The total is the same for both — 16 million, nothing lost. The time differs by an order of magnitude: hundreds of milliseconds against tens. The price is that sum() is not exact at the moment of the call: while it walks the cells, some change. So LongAdder fits counters and statistics, but not logic that needs a precise atomic read point.
When to choose atomics, and when to choose locks
Atomic variables win on a simple operation — an increment, a reference swap, a flag — under moderate contention and on a hot path where every nanosecond counts. Locks (synchronized, ReentrantLock) take over where a compound operation of several steps has to be protected, or where several variables have to be updated consistently: a single CAS cannot express that, and code with synchronized is easier to read.
Rule of thumb: start with synchronized or ReentrantLock if you are unsure. Introduce atomics where you have measured a bottleneck and know that a lock is unnecessary there.
In short
counter++is not atomic — two threads can lose an update.- CAS is a hardware "replace if the value matches" instruction; the whole
java.util.concurrent.atomicpackage is built on it. - The CAS loop retries under contention; the thread does not sleep, but it spends processor cycles.
- The ABA problem: CAS does not see intermediate changes;
AtomicStampedReferenceguards against it. LongAdderis faster thanAtomicLongunder high contention — by spreading work across cells; the price is an approximatesum().
What to read next
- Race conditions in multithreaded programs — where the problem atomic variables close begins.
- The Java memory model (happens-before) — why the visibility of changes is a separate problem and how it relates to atomicity.
- Locks: ReentrantLock and conditions — when you need explicit locks instead of atomic variables.
- Concurrent collections — ready-made thread-safe data structures, built partly on CAS as well.