← Back to the section

To work with S3 from Java people used to reach for AWS SDK v1 — bulky, with blocking I/O and an awkward API. The standard today is AWS SDK v2 (released in 2018): slightly different dependencies, an async HTTP client, immutable builders. Let's walk through how to wire it into Spring without stepping on the usual rakes.

PUT goes straight to the store — bytes never touch the app client backend S3 / MinIO metadata URL, 10 minrow is written first confirm HeadObject users/42doc.pdf documents tablekey = users/42status = PENDING UPLOADED

The backend signs a link and writes a row with status PENDING, while the bytes travel from the client straight into the store, never touching the application's memory or network. The confirmation flips the row to UPLOADED — but only after HeadObject proves the object is really there. Rows whose file never arrived are swept by a background job.

What to add to your dependencies

dependencies {
    implementation(platform("software.amazon.awssdk:bom:2.x.x"))
    implementation("software.amazon.awssdk:s3")
    implementation("software.amazon.awssdk:s3-transfer-manager")  // for large files
    implementation("software.amazon.awssdk:netty-nio-client")      // async HTTP
}

The bom aligns the versions of all AWS modules — you don't have to pin a version on each dependency.

An alternative is Spring Cloud AWS: auto-configuration and properties from application.yml. Convenient for standard AWS; for MinIO and other S3-compatible servers configuring the SDK by hand, as below, is simpler.

How to configure S3Client in Spring

The main object is S3Client. Register it as a Spring bean:

@Configuration
public class S3Config {

    @Bean
    public S3Client s3Client(
            @Value("${aws.s3.region}") String region,
            @Value("${aws.s3.endpoint:#{null}}") String endpoint,
            @Value("${aws.s3.access-key}") String accessKey,
            @Value("${aws.s3.secret-key}") String secretKey) {

        var builder = S3Client.builder()
            .region(Region.of(region))
            .credentialsProvider(StaticCredentialsProvider.create(
                AwsBasicCredentials.create(accessKey, secretKey)));

        if (endpoint != null) {
            builder.endpointOverride(URI.create(endpoint))
                   .forcePathStyle(true);
        }

        return builder.build();
    }

    @Bean
    public S3Presigner s3Presigner(/* the same parameters */) {
        // region and credentialsProvider — same as in s3Client
        var builder = S3Presigner.builder();

        if (endpoint != null) {
            // for MinIO the bucket goes into the path, not into a subdomain —
            // otherwise the signed link does not open
            builder.endpointOverride(URI.create(endpoint))
                   .serviceConfiguration(S3Configuration.builder()
                       .pathStyleAccessEnabled(true).build());
        }

        return builder.build();
    }
}

Settings in application.properties:

aws.s3.region=eu-west-1
aws.s3.access-key=${S3_ACCESS_KEY}
aws.s3.secret-key=${S3_SECRET_KEY}

# For AWS S3 — no endpoint needed, leave it empty.
# For MinIO:
# aws.s3.endpoint=http://minio:9000
# For Yandex Object Storage:
# aws.s3.endpoint=https://storage.yandexcloud.net
# aws.s3.region=ru-central1
# For Cloudflare R2:
# aws.s3.endpoint=https://<account-id>.r2.cloudflarestorage.com
# aws.s3.region=auto

forcePathStyle(true) is required for MinIO: it expects URLs of the form endpoint/bucket/key, whereas AWS by default uses bucket.endpoint/key.

In production the access-key/secret-key pair is usually not set: a role grants the rights and the SDK finds it through DefaultCredentialsProvider — just don't pass a credentialsProvider. On EC2 that role is an IAM Instance Profile, on EKS a service account role (IRSA).

Uploading a file

A plain PUT

For small files putObject is enough:

@Service
@RequiredArgsConstructor
public class AvatarService {

    private final S3Client s3;
    private final String bucket = "user-avatars";

    public void upload(UUID userId, MultipartFile file) throws IOException {
        String key = "users/%s/avatar.jpg".formatted(userId);

        s3.putObject(
            PutObjectRequest.builder()
                .bucket(bucket)
                .key(key)
                .contentType(file.getContentType())
                .contentLength(file.getSize())
                .build(),
            RequestBody.fromInputStream(file.getInputStream(), file.getSize())
        );
    }
}

The SDK doesn't guess the length: RequestBody.fromInputStream takes it as a parameter; the variant without a length reads the whole stream into memory. RequestBody accepts an InputStream, File, Path, byte[] or String.

Large files via TransferManager

From several megabytes up, S3TransferManager is more convenient — it splits the file into parts and uploads them in parallel (multipart upload):

@Bean
public S3AsyncClient s3AsyncClient(/* the same parameters */) {
    // same as S3Client, but via S3AsyncClient.builder()
    // and always .multipartEnabled(true)
}

@Bean
public S3TransferManager transferManager(S3AsyncClient s3Async) {
    return S3TransferManager.builder().s3Client(s3Async).build();
}

@Service
@RequiredArgsConstructor
public class VideoService {

    private final S3TransferManager tm;

    public void upload(Path videoFile, UUID userId) {
        String key = "users/%s/video-%s.mp4".formatted(userId, UUID.randomUUID());

        FileUpload upload = tm.uploadFile(b -> b
            .source(videoFile)
            .putObjectRequest(p -> p.bucket("videos").key(key)));

        upload.completionFuture().join();
    }
}

Multipart is switched on by the client itself: multipartEnabled(true) on a plain S3AsyncClient, or the CRT client S3AsyncClient.crtBuilder(). Without it TransferManager sends the file in one request. Default threshold — 8 MiB; network errors are retried for you.

The mechanics are visible without S3: the file is cut into parts, each part is hashed, and the ETag is assembled from those hashes.

live example

import java.security.MessageDigest;
import java.util.ArrayList;
import java.util.HexFormat;
import java.util.List;
import java.util.Random;

public class MultipartEtag {

    public static void main(String[] args) throws Exception {
        byte[] file = new byte[20 * 1024 * 1024];
        new Random(42).nextBytes(file);
        int partSize = 8 * 1024 * 1024;

        List<byte[]> parts = new ArrayList<>();
        for (int offset = 0; offset < file.length; offset += partSize) {
            int length = Math.min(partSize, file.length - offset);
            MessageDigest part = MessageDigest.getInstance("MD5");
            part.update(file, offset, length);
            parts.add(part.digest());
            System.out.println("part " + parts.size() + ": " + length + " bytes");
        }

        MessageDigest all = MessageDigest.getInstance("MD5");
        parts.forEach(all::update);
        System.out.println("object ETag:   "
            + HexFormat.of().formatHex(all.digest()) + "-" + parts.size());
        System.out.println("md5 of file:   "
            + HexFormat.of().formatHex(MessageDigest.getInstance("MD5").digest(file)));
    }
}
Run

Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →

The -3 tail is the part count, and the value is not the md5 of the file — after a multipart upload ETag cannot prove integrity. For that ask S3 for an explicit checksum: checksumAlgorithm(ChecksumAlgorithm.SHA256).

Downloading a file

// A small file entirely into memory
public byte[] download(String key) {
    return s3.getObjectAsBytes(
        GetObjectRequest.builder().bucket(bucket).key(key).build()
    ).asByteArray();
}

// Save straight into a file on disk
public void downloadToFile(String key, Path target) {
    s3.getObject(
        GetObjectRequest.builder().bucket(bucket).key(key).build(),
        target);
}

// Streaming read — important to close the InputStream
public void processStream(String key) {
    try (InputStream stream = s3.getObject(
            GetObjectRequest.builder().bucket(bucket).key(key).build())) {
        // read the stream in chunks
    }
}

The SDK doesn't close the InputStream for you, and an unclosed connection stays "busy" in the pool — after a few such cases new requests start hanging.

For files over 10 MB getObjectAsBytes is the wrong tool — it keeps the whole content in memory.

Presigned URLs

Sometimes a file has to land in S3 straight from the client, bypassing the backend. The server generates a presigned URL — a temporary signed link the browser or mobile app uses to PUT into S3.

@Service
@RequiredArgsConstructor
public class UploadUrlService {

    private final S3Presigner presigner;

    public String generateUploadUrl(UUID userId) {
        String key = "users/%s/avatar.jpg".formatted(userId);

        PutObjectRequest objectRequest = PutObjectRequest.builder()
            .bucket("user-avatars")
            .key(key)
            .contentType("image/jpeg")
            .contentLength(5 * 1024 * 1024L)
            .build();

        PresignedPutObjectRequest presigned = presigner.presignPutObject(p -> p
            .signatureDuration(Duration.ofMinutes(10))
            .putObjectRequest(objectRequest));

        return presigned.url().toString();
    }
}

The client receives the URL and does a PUT with Content-Type: image/jpeg directly to S3. Add other headers or change the parameters, and the signature won't match: S3 returns an error. A signed contentLength means "exactly this many bytes" — it is not a "no more than 5 MB" limit.

presignGetObject works the same way, for downloading a private file.

How to test without a real S3

In tests MinIO is handy — an S3-compatible server that runs in Docker and behaves like real S3. Testcontainers launches it from the test itself:

@SpringBootTest
@Testcontainers
class AvatarServiceTest {

    @Container
    static MinIOContainer minio = new MinIOContainer("minio/minio:RELEASE.2025-04-22T22-12-26Z")
        .withUserName("test")
        .withPassword("testtest");

    @DynamicPropertySource
    static void s3Props(DynamicPropertyRegistry registry) {
        registry.add("aws.s3.endpoint", () -> minio.getS3URL());
        registry.add("aws.s3.access-key", minio::getUserName);
        registry.add("aws.s3.secret-key", minio::getPassword);
        registry.add("aws.s3.region", () -> "us-east-1");
    }

    @Autowired private AvatarService avatarService;
    @Autowired private S3Client s3;

    @BeforeEach
    void createBucket() {
        // a repeated createBucket fails with BucketAlreadyOwnedByYou
        if (s3.listBuckets().buckets().stream().noneMatch(b -> b.name().equals("user-avatars"))) {
            s3.createBucket(b -> b.bucket("user-avatars"));
        }
    }

    @Test
    void uploads_avatar() throws Exception {
        UUID userId = UUID.randomUUID();
        avatarService.upload(userId, new MockMultipartFile(
            "file", "avatar.jpg", "image/jpeg", "binary".getBytes()));

        var meta = s3.headObject(b -> b
            .bucket("user-avatars")
            .key("users/" + userId + "/avatar.jpg"));
        assertThat(meta.contentLength()).isEqualTo(6);
    }
}

The test runs without an AWS account, in CI, in isolation; MinIOContainer lives in org.testcontainers:minio.

Locally MinIO starts via Docker Compose:

services:
  minio:
    image: minio/minio:RELEASE.2025-04-22T22-12-26Z   # pin the tag: latest breaks builds without warning
    command: server /data --console-address ":9001"
    environment:
      MINIO_ROOT_USER: minio
      MINIO_ROOT_PASSWORD: minio12345
    ports:
      - "9000:9000"   # S3 API
      - "9001:9001"   # Web UI
    volumes:
      - minio-data:/data

volumes:
  minio-data:

After starting, set aws.s3.endpoint=http://localhost:9000 — and the service works with local MinIO.

The problem: "DB + S3" without guarantees

A typical situation: a user uploads a document, and you have to put the file in S3 and create a row in the database. S3 has no transactions, so a failure between the PUT and the INSERT leaves either a file with no row, or a row with no file.

@Transactional won't help here — S3 doesn't take part in a Spring transaction.

The client uploads the file directly to S3, the backend only coordinates:

1. The client sends metadata: { name, size, contentType }
2. The backend creates Document(s3Key, status="PENDING"),
   generates a presigned URL and hands it to the client.
3. The client PUTs to that link directly to S3.
4. The client calls POST /api/docs/{id}/confirm.
5. The backend checks the object via HeadObject, sets status="UPLOADED".

Each step is atomic, and rows left in PENDING with no file are swept by a background job every N minutes.

Pattern: Outbox for deletion

When a file has to be deleted along with its record, an outbox is used: in a single transaction you delete the record and put a job into outbox_events, and a background job performs the S3 deletion:

@Transactional
public void deleteDocument(UUID id) {
    var doc = docRepo.findById(id).orElseThrow();
    docRepo.delete(doc);
    outboxRepo.save(new OutboxEvent("s3.delete", doc.getS3Key()));
}

@Scheduled(fixedDelay = 5000)
public void processS3Outbox() {
    var events = outboxRepo.fetchUnpublished("s3.delete", 100);
    for (var event : events) {
        s3.deleteObject(b -> b.bucket("docs").key(event.payload()));
        outboxRepo.markPublished(event.id());
    }
}

A failed deletion is retried on the next run of the job.

In short

  • AWS SDK v2 is the standard for S3 in Java: immutable builders, retries and timeouts out of the box, an async client when needed.
  • S3Client for ordinary operations, S3TransferManager for large files; multipart needs multipartEnabled(true) or the CRT client, and such an ETag is no longer the md5 of the file.
  • forcePathStyle(true) is mandatory for MinIO, and S3Presigner needs pathStyleAccessEnabled(true).
  • In production the rights come from a role — IAM Instance Profile on EC2, IRSA on EKS — not from a key pair.
  • A presigned URL lets the client write straight into the store; close the InputStream after getObject yourself.
  • S3 is non-transactional: "DB + S3" atomicity comes from a two-phase upload with a status or from an outbox, and MinIOContainer checks it without an AWS account.