← Back to the section

Hibernate does a lot for you — and that is exactly why it easily hides problems until your application starts slowing down or corrupting data. Here are the ones you will run into most often. We start with the classic one: an entity put into a HashSet goes missing there.

HashSet: the bucket is picked by hash — eight buckets, index = hash & 7 1. the object is not saved yet: id = null, hash 0 — it goes into bucket 0 0 Product 1 · 2 · 3 · 4 · 5 · 6 · 7 · 2. the database assigned id = 42 on save: the hash is now 42, we look in bucket 2 0 Product 1 · 2 ? 3 · 4 · 5 · 6 · 7 · bucket 2 is empty — contains() = false, while the object still sits in bucket 0 3. hash from the business key sku: bucket 7 both before and after saving 0 · 1 · 2 · 3 · 4 · 5 · 6 · 7 Product contains() = true — the id appeared, the hash did not change, same bucket

A HashSet picks the bucket once — at insertion time, from the current hash. While id is null the hash is zero and the object lands in bucket 0; after the save the database assigns id = 42, the hash changes, and the lookup goes to bucket 2, which is empty. The object never left the set — contains simply stops finding it. A hash from the business key sku does not change on save, so the bucket stays the same.

equals and hashCode on an entity

When an entity goes into a HashSet or HashMap, Java uses equals and hashCode. Without an explicit override you get the one from Object — comparison by reference. That is almost always wrong.

The first instinct is to generate them from id. But there is a trap: until the entity is saved, id is null. Two new objects with id == null will turn out to be "equal", and an object already put into a set changes its hash the moment the database assigns an id.

@Entity
public class Product {

    @Id
    @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    // Workable minimum: equals by id, guarded against null and against proxies
    @Override
    public boolean equals(Object o) {
        if (this == o) return true;
        if (!(o instanceof Product p)) return false;
        return id != null && id.equals(p.getId());   // getId(), not p.id
    }

    @Override
    public int hashCode() {
        return 31; // constant: the hash will not change once the database assigns an id
    }
}

That is the workable minimum: equals is guarded against null, and a constant hashCode does not change when the object is saved.

One detail matters: p.getId(), not p.id. A lazy association hands you a proxy — a subclass whose own fields stay empty until a getter loads them, so reading the field directly returns null and the comparison silently fails. A proxy also has a different class, which is why hashCode here is a constant and not getClass().hashCode().

The better option is a business key: a field that is unique and stable even before saving (a SKU, an email, a UUID generated in code rather than by the database).

@Entity
public class Product {

    @Id
    @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    @Column(unique = true, nullable = false)
    private String sku; // business key

    @Override
    public boolean equals(Object o) {
        if (this == o) return true;
        if (!(o instanceof Product p)) return false;
        return sku != null && sku.equals(p.sku);
    }

    @Override
    public int hashCode() {
        return Objects.hashCode(sku);
    }
}

The trap shows up without a database — the id is assigned in code:

live example

import java.util.HashSet;
import java.util.Objects;
import java.util.Set;

public class EntityHashDemo {

    static class ById {
        Long id;

        @Override
        public boolean equals(Object o) {
            return o instanceof ById other && Objects.equals(id, other.id);
        }

        @Override
        public int hashCode() {
            return Objects.hashCode(id);
        }
    }

    static class BySku {
        Long id;
        final String sku;

        BySku(String sku) {
            this.sku = sku;
        }

        @Override
        public boolean equals(Object o) {
            return o instanceof BySku other && sku.equals(other.sku);
        }

        @Override
        public int hashCode() {
            return sku.hashCode();
        }
    }

    public static void main(String[] args) {
        Set<ById> byId = new HashSet<>();
        ById first = new ById();
        byId.add(first);
        System.out.println("hash by id, id still null: contains = " + byId.contains(first));
        first.id = 42L;
        System.out.println("same object after the save:  contains = " + byId.contains(first));
        System.out.println("yet it is in the set:       size = " + byId.size());

        Set<BySku> bySku = new HashSet<>();
        BySku second = new BySku("SKU-7");
        bySku.add(second);
        second.id = 42L;
        System.out.println("hash by business key, after the save: contains = " + bySku.contains(second));
    }
}
Run

Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →

A short formula: the hash must not change when the object moves from transient to managed — which means you cannot compute it from id if the database assigns the id.

Open Session in View

Open Session in View (OSIV) is a pattern where the Hibernate session stays open for the entire duration of the HTTP request, including rendering the response. In Spring Boot it is enabled by default (spring.jpa.open-in-view=true).

At first glance it looks convenient: lazy collections load in the template, no LazyInitializationException. In practice there are two hidden harms:

  1. A database connection is held longer than needed — in a typical configuration until the end of the request, including slow template rendering. Under load the connection pool runs out quickly.
  2. It masks N+1: queries fire off to the database from the presentation layer, where nobody expects or controls them.
# application.yml — disable OSIV
spring:
  jpa:
    open-in-view: false

Once disabled, LazyInitializationException becomes explicit — right where data was not loaded inside a transaction, and that is a good thing: the problem is visible. The fix is to load what you need in the service layer (via JOIN FETCH or EntityGraph) and return a DTO, not an entity.

Returning an entity from a controller instead of a DTO

An entity is not a DTO. If you return an @Entity directly from a controller to Jackson, several unpleasant things happen:

  • Leaking the internal structure: the client sees every field, technical ones included.
  • Recursion during serialization: bidirectional relationships (@OneToMany + @ManyToOne) loop forever — the cure is @JsonIgnore or @JsonManagedReference, which clutter the domain code.
  • Unexpected loading: Jackson walks every field, lazy collections included — with the session still open (OSIV) Hibernate runs extra queries.
// Bad: returning the entity straight from the controller
@GetMapping("/products/{id}")
public Product getProduct(@PathVariable Long id) {
    return productRepository.findById(id).orElseThrow();
}

// Good: convert to a DTO in the service layer
@GetMapping("/products/{id}")
public ProductDto getProduct(@PathVariable Long id) {
    return productService.getById(id); // maps to a DTO inside
}

A DTO per response is not bureaucracy — it is the boundary between the internal model and the public contract.

CascadeType.ALL and orphanRemoval — a dangerous combination

CascadeType.ALL propagates all operations (PERSIST, MERGE, REMOVE, REFRESH, DETACH) from the parent to child entities. In most cases that is excessive.

The combination CascadeType.ALL + orphanRemoval = true is especially dangerous: if you remove a child object from the parent's collection, Hibernate deletes it from the database.

@Entity
public class Order {

    @OneToMany(mappedBy = "order",
               cascade = CascadeType.ALL,  // includes REMOVE
               orphanRemoval = true)        // deletes the row from items when removed from the collection
    private List<OrderItem> items = new ArrayList<>();
}

// In code:
order.getItems().clear(); // <-- this will delete ALL rows in order_item for this order!

Use only the cascade types you actually need: usually CascadeType.PERSIST and CascadeType.MERGE. Add REMOVE only where children make no sense without the parent — receipt line items, say.

// Explicit and safe
@OneToMany(mappedBy = "order", cascade = {CascadeType.PERSIST, CascadeType.MERGE})
private List<OrderItem> items = new ArrayList<>();

Bulk operations done one object at a time

The usual approach through a JPA repository: load the entities, change them in a loop, save. With thousands of rows that becomes thousands of separate UPDATEs:

// Bad: N UPDATE queries to the database
List<Product> products = productRepository.findAll();
for (Product p : products) {
    p.setPrice(p.getPrice().multiply(BigDecimal.valueOf(1.1)));
    productRepository.save(p); // in fact redundant: the object is already managed
}

Note that save() does not write on every iteration, as is often assumed: all the UPDATEs go out when the transaction commits. The cost is elsewhere: the context keeps every loaded product plus a snapshot of each, and the database still gets one UPDATE per row.

Instead, use a bulk query through JPQL or native SQL:

// Good: a single UPDATE for every row
@Modifying(clearAutomatically = true, flushAutomatically = true)
@Query("UPDATE Product p SET p.price = p.price * 1.1 WHERE p.category = :category")
int increasePricesByCategory(@Param("category") String category);

Such a query goes past the persistence context: the first-level cache does not know about the changes, so entities loaded earlier are stale. clearAutomatically = true clears the context after the query, flushAutomatically = true writes pending changes out before it; without them the next read returns the old prices.

For more on the persistence context and caching, see the article Persistence Context.

merge vs save: what happens

In Spring Data JPA save() does one of two things depending on the state of the object:

  • If id == null (or the object is new according to Persistable) — it calls entityManager.persist().
  • If id is set — it calls entityManager.merge().

merge behaves in a non-obvious way: it does not update the object you passed in but returns a managed copy from the persistence context. The original stays detached.

Product detached = new Product();
detached.setId(42L);
detached.setName("New name");

Product managed = productRepository.save(detached);
// detached is still detached, its changes are not tracked!
// managed is the managed copy, and that is what you should keep working with

managed.setPrice(BigDecimal.valueOf(999)); // this will hit the database on flush
detached.setPrice(BigDecimal.valueOf(0));  // this is saved NOWHERE

To update specific fields, load the entity inside a transaction and change it there — then merge is not needed at all.

In short

  • equals/hashCode by id are dangerous: before saving, id == null. Use a constant hashCode or a business key.
  • OSIV holds a connection to the end of the request and hides N+1 — disable it and load data explicitly inside a transaction.
  • Do not return an @Entity from a controller — use a DTO so you do not leak the internal model or get unexpected queries during serialization.
  • CascadeType.ALL + orphanRemoval delete rows on clear() of the collection — use them only deliberately, and prefer an explicit set of types.
  • Do bulk changes with bulk queries (@Modifying + @Query), not in a loop; afterwards clear the first-level cache.
  • save() with a non-null id calls merge — the returned object is managed, the original stays detached.
  • Persistence Context — how Hibernate tracks changes and when they are flushed to the database.
  • Lazy vs Eager loading — why LazyInitializationException happens and how to avoid it properly.
  • The N+1 problem — diagnosing and solving the most common cause of slow queries.
  • Spring Data JPA — repositories on top of Hibernate: methods, derived queries, @Query.