Elasticsearch is a powerful search engine, but writing requests to its JSON API by hand means a lot of boilerplate. Spring Data Elasticsearch takes that work over: you describe the document with a Java class and do search and indexing with the usual Spring methods.
Spring Data does nothing magical: it builds JSON out of the object following the annotations and hands it to the index. Elasticsearch takes over from there — the document first sits in the write buffer, where search cannot see it. Refresh moves it into a segment, once a second by default; hence the rule "saved does not mean found".
Connecting
Add the dependency to build.gradle.kts:
dependencies {
implementation("org.springframework.boot:spring-boot-starter-data-elasticsearch")
}
And configure the cluster address in application.properties:
spring.elasticsearch.uris=http://elasticsearch:9200
spring.elasticsearch.username=elastic
spring.elasticsearch.password=${ES_PASSWORD}
spring.elasticsearch.connection-timeout=2s
spring.elasticsearch.socket-timeout=10s
Spring Boot creates a RestClient — the HTTP client to Elasticsearch — puts ElasticsearchOperations on top of it, and implements the ElasticsearchRepository interfaces it finds in your packages on the fly.
How to describe a document with @Document
In JPA you describe a table with @Entity; in Elasticsearch @Document plays that role and links a Java class to an index.
@Document(indexName = "products")
public class ProductDoc {
@Id
private String id;
@Field(type = FieldType.Text, analyzer = "english")
private String name;
@MultiField(
mainField = @Field(type = FieldType.Text, analyzer = "english"),
otherFields = {
@InnerField(suffix = "raw", type = FieldType.Keyword)
}
)
private String description;
@Field(type = FieldType.Long)
private Long categoryId;
@Field(type = FieldType.ScaledFloat, scalingFactor = 100)
private BigDecimal price;
@Field(type = FieldType.Boolean)
private Boolean inStock;
@Field(type = FieldType.Date, format = DateFormat.date_time)
private Instant createdAt;
// getters and setters
}
A few important details:
@Field(type = FieldType.Text, analyzer = "english")— a full-text search field with a language analyzer (stemming, stop words).@MultiField— one field in two variants:Textfor meaning-based search andKeyword(suffix.raw) for exact filters and sorting.FieldType.ScaledFloat— the recommended type for prices: it stores the number as an integer with a scale (100 → cents), so rounding never eats them.- This is not JPA: there is no
@Entity,@Table, or Hibernate transactions here. The class exists only to map the ES document.
The difference between Text and Keyword is not about storage but about what goes into the index: Text runs through an analyzer and falls apart into words, Keyword goes in whole:
live example
import java.util.Arrays;
import java.util.List;
public class MappingDemo {
static List<String> asText(String value) {
return Arrays.stream(value.toLowerCase().split("[^\\p{L}\\p{N}]+"))
.filter(word -> !word.isEmpty())
.toList();
}
public static void main(String[] args) {
String title = "Hand Cream Velvet";
System.out.println("Text: " + asText(title));
System.out.println("Keyword: [" + title + "]");
System.out.println("found by the word \"cream\" in Text: " + asText(title).contains("cream"));
System.out.println("found by the word \"cream\" in Keyword: " + title.equals("cream"));
}
}
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
So description is searched by words, and description.raw filtered and sorted.
For fine-grained control over the mapping, point to a ready-made JSON file: @Mapping(mappingPath = "elasticsearch/product-mapping.json").
Simple queries with ElasticsearchRepository
The fastest start is a repository interface:
public interface ProductRepository extends ElasticsearchRepository<ProductDoc, String> {
Page<ProductDoc> findByCategoryId(Long categoryId, Pageable pageable);
List<ProductDoc> findByNameContainingAndInStockTrue(String namePart);
long countByPriceBetween(BigDecimal min, BigDecimal max);
}
Spring parses the method names and generates the Query DSL. That covers simple cases: filters by exact values, sorting, pagination.
Where the repository stops being enough:
- Method names grow to 8–10 words and stop being readable.
- There's no control over fuzzy search (
fuzziness), field boosts, or aggregations. - No complex
boolquery with several conditions.
That's what ElasticsearchOperations is for.
Flexible queries with ElasticsearchOperations
ElasticsearchOperations is a lower-level API with full control over the query:
@Service
@RequiredArgsConstructor
public class ProductSearchService {
private final ElasticsearchOperations elasticsearch;
public SearchHits<ProductDoc> search(String text, Set<Long> categories,
BigDecimal minPrice, Pageable pageable) {
var criteria = new Criteria("name").matches(text)
.and(new Criteria("inStock").is(true));
if (!categories.isEmpty()) {
criteria = criteria.and(new Criteria("categoryId").in(categories));
}
if (minPrice != null) {
criteria = criteria.and(new Criteria("price").greaterThanEqual(minPrice));
}
var query = new CriteriaQuery(criteria, pageable);
return elasticsearch.search(query, ProductDoc.class);
}
}
When even CriteriaQuery isn't enough — aggregations, or a score that decays with age — you take NativeQuery: almost the same JSON, but with type checking at compile time:
public SearchHits<ProductDoc> searchWithFunctionScore(String text) {
var query = NativeQuery.builder()
.withQuery(q -> q
.functionScore(fs -> fs
.query(qq -> qq.match(m -> m.field("name").query(text)))
.functions(f -> f
.gauss(g -> g.date(d -> d
.field("createdAt")
.placement(p -> p.origin("now")
.scale(Time.of(t -> t.time("30d"))).decay(0.5)))))
))
.build();
return elasticsearch.search(query, ProductDoc.class);
}
Bulk indexing
Indexing documents one at a time is slow: every call is a separate HTTP request with a wait for confirmation. Not how you load a large catalog.
The Bulk API sends thousands of documents in one request:
public void reindexAll(List<ProductDoc> docs) {
var queries = docs.stream()
.map(doc -> new IndexQueryBuilder()
.withId(doc.getId())
.withObject(doc)
.build())
.toList();
elasticsearch.bulkIndex(queries, ProductDoc.class);
}
For very large volumes (millions of documents), documents go in batches of 500–5000 and automatic refresh is turned off for the load. Spring Data has no knob for that — the setting is changed in Elasticsearch itself:
PUT /products/_settings
{ "index": { "refresh_interval": "-1" } }
Afterwards you put "1s" back and call elasticsearch.indexOps(ProductDoc.class).refresh(), or the documents you just loaded won't be found. The speedup over one-at-a-time indexing is 10–50x.
How to keep the index up to date: four approaches
In most applications, Elasticsearch is not the main database but a search index sitting next to PostgreSQL: data appears in PostgreSQL, and you need to reflect it in ES promptly. There are four ways.
Dual write — simple, but unreliable
The obvious option: write to PostgreSQL and to ES in one method.
@Transactional
public void save(Product product) {
productRepo.save(product); // PostgreSQL
elasticsearch.save(toDoc(product)); // Elasticsearch
}
The problem: a PostgreSQL transaction and an HTTP request to ES are two different operations. PostgreSQL commits, ES returns a network error — the data diverges and nothing brings it back. Prototypes only, not production.
Transactional Outbox — reliable, but needs infrastructure
Inside the PostgreSQL transaction, together with the business data, we write an event to an outbox table. A separate reader takes events from it and sends them to ES.
@Transactional
public void save(Product product) {
productRepo.save(product);
outboxRepo.save(new OutboxEvent(
UUID.randomUUID(),
"product.updated",
toJson(product)
));
}
The reader is a plain @Scheduled method: take a batch of unpublished events, send them to ES, mark them published.
Pros: the data won't diverge — the event is written atomically with the business data, and if ES fails it stays unread and gets retried. Cons: an outbox table, reader logic, and lag monitoring.
The difference shows without Elasticsearch — what matters is that the database write and the search write are not in one transaction:
live example
import java.util.ArrayList;
import java.util.Iterator;
import java.util.List;
public class SyncDemo {
static final List<String> postgres = new ArrayList<>();
static final List<String> outbox = new ArrayList<>();
static final List<String> index = new ArrayList<>();
static boolean esAvailable = false;
static void sendToEs(String doc) {
if (!esAvailable) {
throw new IllegalStateException("ES is unavailable");
}
index.add(doc);
}
static void dualWrite(String doc) {
postgres.add(doc);
try {
sendToEs(doc);
} catch (RuntimeException e) {
// the transaction is already committed — nothing to roll back
}
}
static void writeWithOutbox(String doc) {
postgres.add(doc);
outbox.add(doc);
}
static void relay() {
for (Iterator<String> it = outbox.iterator(); it.hasNext(); ) {
try {
sendToEs(it.next());
it.remove();
} catch (RuntimeException e) {
return;
}
}
}
public static void main(String[] args) {
dualWrite("product-1");
writeWithOutbox("product-2");
relay();
System.out.println("ES is down. database: " + postgres + ", index: " + index + ", waiting: " + outbox);
esAvailable = true;
relay();
System.out.println("ES is back. index: " + index + ", waiting: " + outbox);
System.out.println("product-1 lost for good: " + !index.contains("product-1"));
}
}
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
The dual write loses product-1 silently; the event in outbox waited and reached the index once ES came back.
CDC via Debezium → Kafka → ES — the industrial option
With Change Data Capture (CDC) your service knows nothing about Elasticsearch at all: it simply writes to PostgreSQL.
The index is updated from the database log rather than by the application, so it receives everything that was actually written — including changes made around the service. The price is one more pipeline to operate.
Debezium reads the PostgreSQL transaction log (WAL) and publishes each change as an event in Kafka. The Elasticsearch Sink connector reads these events and inserts/updates/deletes documents in ES.
Advantages: the service isn't tied to ES; every change is captured, including direct edits in the database; after a failure the connector resumes from its last position. Indexing latency is usually 100–500 milliseconds. The downside: you deploy and maintain Debezium, Kafka, and Kafka Connect.
Full reindex on a schedule
If the data changes rarely (reference data, catalogs), just rebuild the whole index at night:
@Scheduled(cron = "0 0 3 * * *", zone = "Europe/Moscow")
public void reindexCatalog() {
var newIndex = "products-v" + System.currentTimeMillis();
elasticsearch.indexOps(IndexCoordinates.of(newIndex)).create();
// in batches, not one document at a time — see the bulk section above
var docs = productRepo.findAll().stream().map(this::toDoc).toList();
for (int from = 0; from < docs.size(); from += 1000) {
elasticsearch.save(docs.subList(from, Math.min(from + 1000, docs.size())),
IndexCoordinates.of(newIndex));
}
// switch the alias in one operation: remove from the old index, add to the new one
elasticsearch.indexOps(IndexCoordinates.of(newIndex)).alias(
new AliasActions()
.add(new AliasAction.Remove(
AliasActionParameters.builder()
.withIndices("products-*").withAliases("products").build()))
.add(new AliasAction.Add(
AliasActionParameters.builder()
.withIndices(newIndex).withAliases("products").build()))
);
}
One thing is essential: the alias must be removed from the old index in the same operation — otherwise it points at both and search returns every product twice. The products-* pattern in Remove takes it off every previous version of the index; on the very first run, while there is no alias yet, that action is skipped.
Simple to build and maintain. Not for updates that must show in search in real time.
Aliases — switching indices without downtime
An Elasticsearch schema can't be changed in place: renaming a field or changing a type means a new index and a copy of the data. Aliases keep the application running while that happens.
An alias is a pointer to an index: the application refers to products without knowing which index sits behind it.
# the schema changed: products-v2 is created and filled, now switch the alias
POST /_aliases
{
"actions": [
{ "remove": { "index": "products-v1", "alias": "products" } },
{ "add": { "index": "products-v2", "alias": "products" } }
]
}
Switching is atomic: all queries move to the new index at once, with no errors on the application side.
Common pitfalls
Elasticsearch is not transactional
elasticsearch.save(doc) saves the document immediately and does not roll back together with the PostgreSQL transaction. All the synchronization approaches above are built around this limitation.
A document is not visible immediately after saving
By default, ES refreshes the search index once per second. Search for a document right after saving it in an integration test, and it may not be there. The fix for tests:
elasticsearch.withRefreshPolicy(RefreshPolicy.WAIT_UNTIL).save(doc);
withRefreshPolicy returns a copy of ElasticsearchOperations with another policy: save has no such parameter.
In production it's expensive; better to accept the latency.
The client version and the cluster version must match
The elasticsearch-java client works with a cluster of its own major version: 8.x with 8.x, 9.x with 9.x. When you upgrade the cluster, upgrade the client too.
Field types must be set correctly from the start
If you describe the categoryId field as Keyword instead of Long, you'll have to pass it as a string. You can't change the type later without recreating the index.
Document size
A large document is expensive beyond the upload: change one field and Elasticsearch reindexes the whole thing. The request size is capped at 100 MB by default (http.max_content_length), but it gets heavy long before that — keep large texts elsewhere and leave only searchable fields in ES.
In short
spring-boot-starter-data-elasticsearchsets up the client and repositories;@Documentdescribes the index,@Fieldthe field types, and@MultiFieldgives you a field in two variants at once: words for search, whole for filters.ElasticsearchRepositoryis good for simple method-name queries;ElasticsearchOperationsfor complex ones with full control.- The Bulk API is 10–50x faster than one-at-a-time indexing; for large volumes, temporarily turn off
refresh_interval. - Dual write is unreliable; for production, use Transactional Outbox or CDC via Debezium.
- Use aliases instead of index names — that's how the schema changes without downtime.
- A saved document becomes visible in search about a second later — account for it in tests.
What to read next
- Fundamentals — how the index works, shards, replication.
- Query DSL and relevance — what Spring Data Elasticsearch generates under the hood.
- Operations — ILM, snapshots, cluster tuning.
- Distributed patterns — more on Outbox and CDC.