Hibernate does a lot for you — and that is exactly why it easily hides problems until your application starts slowing down or corrupting data. Here are the ones you will run into most often. We start with the classic one: an entity put into a HashSet goes missing there.
A HashSet picks the bucket once — at insertion time, from the current hash. While id is null the hash is zero and the object lands in bucket 0; after the save the database assigns id = 42, the hash changes, and the lookup goes to bucket 2, which is empty. The object never left the set — contains simply stops finding it. A hash from the business key sku does not change on save, so the bucket stays the same.
equals and hashCode on an entity
When an entity goes into a HashSet or HashMap, Java uses equals and hashCode. Without an explicit override you get the one from Object — comparison by reference. That is almost always wrong.
The first instinct is to generate them from id. But there is a trap: until the entity is saved, id is null. Two new objects with id == null will turn out to be "equal", and an object already put into a set changes its hash the moment the database assigns an id.
@Entity
public class Product {
@Id
@GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;
// Workable minimum: equals by id, guarded against null and against proxies
@Override
public boolean equals(Object o) {
if (this == o) return true;
if (!(o instanceof Product p)) return false;
return id != null && id.equals(p.getId()); // getId(), not p.id
}
@Override
public int hashCode() {
return 31; // constant: the hash will not change once the database assigns an id
}
}
That is the workable minimum: equals is guarded against null, and a constant hashCode does not change when the object is saved.
One detail matters: p.getId(), not p.id. A lazy association hands you a proxy — a subclass whose own fields stay empty until a getter loads them, so reading the field directly returns null and the comparison silently fails. A proxy also has a different class, which is why hashCode here is a constant and not getClass().hashCode().
The better option is a business key: a field that is unique and stable even before saving (a SKU, an email, a UUID generated in code rather than by the database).
@Entity
public class Product {
@Id
@GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;
@Column(unique = true, nullable = false)
private String sku; // business key
@Override
public boolean equals(Object o) {
if (this == o) return true;
if (!(o instanceof Product p)) return false;
return sku != null && sku.equals(p.sku);
}
@Override
public int hashCode() {
return Objects.hashCode(sku);
}
}
The trap shows up without a database — the id is assigned in code:
live example
import java.util.HashSet;
import java.util.Objects;
import java.util.Set;
public class EntityHashDemo {
static class ById {
Long id;
@Override
public boolean equals(Object o) {
return o instanceof ById other && Objects.equals(id, other.id);
}
@Override
public int hashCode() {
return Objects.hashCode(id);
}
}
static class BySku {
Long id;
final String sku;
BySku(String sku) {
this.sku = sku;
}
@Override
public boolean equals(Object o) {
return o instanceof BySku other && sku.equals(other.sku);
}
@Override
public int hashCode() {
return sku.hashCode();
}
}
public static void main(String[] args) {
Set<ById> byId = new HashSet<>();
ById first = new ById();
byId.add(first);
System.out.println("hash by id, id still null: contains = " + byId.contains(first));
first.id = 42L;
System.out.println("same object after the save: contains = " + byId.contains(first));
System.out.println("yet it is in the set: size = " + byId.size());
Set<BySku> bySku = new HashSet<>();
BySku second = new BySku("SKU-7");
bySku.add(second);
second.id = 42L;
System.out.println("hash by business key, after the save: contains = " + bySku.contains(second));
}
}
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
A short formula: the hash must not change when the object moves from transient to managed — which means you cannot compute it from id if the database assigns the id.
Open Session in View
Open Session in View (OSIV) is a pattern where the Hibernate session stays open for the entire duration of the HTTP request, including rendering the response. In Spring Boot it is enabled by default (spring.jpa.open-in-view=true).
At first glance it looks convenient: lazy collections load in the template, no LazyInitializationException. In practice there are two hidden harms:
- A database connection is held longer than needed — in a typical configuration until the end of the request, including slow template rendering. Under load the connection pool runs out quickly.
- It masks N+1: queries fire off to the database from the presentation layer, where nobody expects or controls them.
# application.yml — disable OSIV
spring:
jpa:
open-in-view: false
Once disabled, LazyInitializationException becomes explicit — right where data was not loaded inside a transaction, and that is a good thing: the problem is visible. The fix is to load what you need in the service layer (via JOIN FETCH or EntityGraph) and return a DTO, not an entity.
Returning an entity from a controller instead of a DTO
An entity is not a DTO. If you return an @Entity directly from a controller to Jackson, several unpleasant things happen:
- Leaking the internal structure: the client sees every field, technical ones included.
- Recursion during serialization: bidirectional relationships (
@OneToMany+@ManyToOne) loop forever — the cure is@JsonIgnoreor@JsonManagedReference, which clutter the domain code. - Unexpected loading: Jackson walks every field, lazy collections included — with the session still open (OSIV) Hibernate runs extra queries.
// Bad: returning the entity straight from the controller
@GetMapping("/products/{id}")
public Product getProduct(@PathVariable Long id) {
return productRepository.findById(id).orElseThrow();
}
// Good: convert to a DTO in the service layer
@GetMapping("/products/{id}")
public ProductDto getProduct(@PathVariable Long id) {
return productService.getById(id); // maps to a DTO inside
}
A DTO per response is not bureaucracy — it is the boundary between the internal model and the public contract.
CascadeType.ALL and orphanRemoval — a dangerous combination
CascadeType.ALL propagates all operations (PERSIST, MERGE, REMOVE, REFRESH, DETACH) from the parent to child entities. In most cases that is excessive.
The combination CascadeType.ALL + orphanRemoval = true is especially dangerous: if you remove a child object from the parent's collection, Hibernate deletes it from the database.
@Entity
public class Order {
@OneToMany(mappedBy = "order",
cascade = CascadeType.ALL, // includes REMOVE
orphanRemoval = true) // deletes the row from items when removed from the collection
private List<OrderItem> items = new ArrayList<>();
}
// In code:
order.getItems().clear(); // <-- this will delete ALL rows in order_item for this order!
Use only the cascade types you actually need: usually CascadeType.PERSIST and CascadeType.MERGE. Add REMOVE only where children make no sense without the parent — receipt line items, say.
// Explicit and safe
@OneToMany(mappedBy = "order", cascade = {CascadeType.PERSIST, CascadeType.MERGE})
private List<OrderItem> items = new ArrayList<>();
Bulk operations done one object at a time
The usual approach through a JPA repository: load the entities, change them in a loop, save. With thousands of rows that becomes thousands of separate UPDATEs:
// Bad: N UPDATE queries to the database
List<Product> products = productRepository.findAll();
for (Product p : products) {
p.setPrice(p.getPrice().multiply(BigDecimal.valueOf(1.1)));
productRepository.save(p); // in fact redundant: the object is already managed
}
Note that save() does not write on every iteration, as is often assumed: all the UPDATEs go out when the transaction commits. The cost is elsewhere: the context keeps every loaded product plus a snapshot of each, and the database still gets one UPDATE per row.
Instead, use a bulk query through JPQL or native SQL:
// Good: a single UPDATE for every row
@Modifying(clearAutomatically = true, flushAutomatically = true)
@Query("UPDATE Product p SET p.price = p.price * 1.1 WHERE p.category = :category")
int increasePricesByCategory(@Param("category") String category);
Such a query goes past the persistence context: the first-level cache does not know about the changes, so entities loaded earlier are stale. clearAutomatically = true clears the context after the query, flushAutomatically = true writes pending changes out before it; without them the next read returns the old prices.
For more on the persistence context and caching, see the article Persistence Context.
merge vs save: what happens
In Spring Data JPA save() does one of two things depending on the state of the object:
- If
id == null(or the object is new according toPersistable) — it callsentityManager.persist(). - If
idis set — it callsentityManager.merge().
merge behaves in a non-obvious way: it does not update the object you passed in but returns a managed copy from the persistence context. The original stays detached.
Product detached = new Product();
detached.setId(42L);
detached.setName("New name");
Product managed = productRepository.save(detached);
// detached is still detached, its changes are not tracked!
// managed is the managed copy, and that is what you should keep working with
managed.setPrice(BigDecimal.valueOf(999)); // this will hit the database on flush
detached.setPrice(BigDecimal.valueOf(0)); // this is saved NOWHERE
To update specific fields, load the entity inside a transaction and change it there — then merge is not needed at all.
In short
equals/hashCodebyidare dangerous: before saving,id == null. Use a constanthashCodeor a business key.- OSIV holds a connection to the end of the request and hides N+1 — disable it and load data explicitly inside a transaction.
- Do not return an
@Entityfrom a controller — use a DTO so you do not leak the internal model or get unexpected queries during serialization. CascadeType.ALL + orphanRemovaldelete rows onclear()of the collection — use them only deliberately, and prefer an explicit set of types.- Do bulk changes with bulk queries (
@Modifying + @Query), not in a loop; afterwards clear the first-level cache. save()with a non-nullidcallsmerge— the returned object ismanaged, the original staysdetached.
What to read next
- Persistence Context — how Hibernate tracks changes and when they are flushed to the database.
- Lazy vs Eager loading — why
LazyInitializationExceptionhappens and how to avoid it properly. - The N+1 problem — diagnosing and solving the most common cause of slow queries.
- Spring Data JPA — repositories on top of Hibernate: methods, derived queries,
@Query.