Avatars, documents, receipts, exports — files show up in every service sooner or later. Putting them "in the database for now" is the simplest option: transactions, backups, and access come out of the box. It is even the right choice — but only up to a certain point. Let's look at where that point is and what changes beyond it.
A bytea value over two kilobytes is compressed and moved into TOAST — a side table that belongs to the same table: every dump drags it along, and reading the column reassembles the chunks in memory. With object storage the row keeps the bucket and the key, a few dozen bytes: the database dump gets slim, and the client picks the bytes up over a temporary link straight from the storage. For a file of a couple of kilobytes the trade-off flips — one transaction and one backup are cheaper than two systems.
Storing in the database
In PostgreSQL a file goes into a column of type bytea — right next to the rest of the row's data. The obvious upside: the file lives inside the transaction. You save a record together with its file — either both are fully there or neither is. Delete the record and the file goes away automatically. A database backup includes the files too.
CREATE TABLE contracts (
id bigserial PRIMARY KEY,
title text NOT NULL,
signed_at timestamptz,
document bytea -- the file itself lives here
);
This fits when files are small (tens to hundreds of kilobytes) and you need strict atomicity — for example, a signature or a certificate without which the record makes no sense.
Trouble starts with size. A whole row has to fit into roughly two kilobytes, so PostgreSQL first compresses a long value and, if it still does not fit, splits it into chunks and moves them into TOAST, a side table belonging to this same table. The chunks take up space there, they end up in the dump there, and reading the column pulls them all back into memory. More than a gigabyte does not go into bytea at all. A row with a 50-megabyte PDF makes the table unmanageable: every dump drags along gigabytes of immutable files, restores stretch into hours, and parallel downloads lead to memory exhaustion.
Storing in object storage
Object storage (Amazon S3, MinIO, Yandex Object Storage) is an "endless shelf" for files. Each file is a separate object under a key. PostgreSQL stores only the row with metadata and the object key; the file itself lives in the storage.
Serving a file to a client looks different: the service hands the client a presigned URL — a temporary signed link the client uses to download the file directly from the storage, bypassing the service. The connection pool, the application memory, and the service's network stack take no part in transferring the bytes.
The price: the pair "row in the database + object in the storage" spans two different systems. The database transaction knows nothing about the file; if something goes wrong, you have to reconcile them by hand.
How to choose
Six questions. Every "yes" is an argument in favor of object storage:
- Is a typical file larger than ~1 MB? Megabytes in
byteabloat the table and the backups. - Will the files take up more than a quarter of the database volume? A database backup should not drag along immutable PDFs.
- Are files served to users regularly? Every serve through the service is load on the connection pool and memory.
- Do you need a lifecycle: archive, automatically delete? In object storage this is a declarative policy; in the database it is hand-written jobs.
- Are files processed: resize, conversion, antivirus? Processing pipelines build around object storage naturally.
- Do other services or external partners reach the files? Presigned URLs and bucket policies express this out of the box, without proxying.
0–1 "yes" answers — bytea next to the metadata, and don't overcomplicate.
2 or more — object storage.
The first two questions can be answered with arithmetic up front. Take the orders from the sandbox and see what the table turns into if every order carries a 2 MB receipt — and what the same table weighs when it holds nothing but an object key:
live example
WITH receipts AS (
SELECT count(*) AS n FROM orders
)
SELECT n AS orders,
pg_size_pretty(n * 2 * 1024 * 1024) AS bytea_column,
pg_size_pretty(n * 80) AS object_keys
FROM receipts;
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
The gap between the two columns of the answer is exactly the gap in dump size and in restore time.
Two-phase file upload
When object storage is used, the upload is organized in two steps to avoid the "the file exists but the record doesn't" situation (or the other way around).
Step 1. The service creates a record in the database with the status PENDING and hands the client a presigned URL for uploading to the storage.
Step 2. The client uploads the file directly to the storage.
Step 3. The client tells the service "uploaded". The service checks that the object actually exists (the expected size, type) and moves the record to the status ACTIVE.
Records stuck in PENDING and objects without records — "orphaned objects" — are cleaned up by a background job. Inconsistency here is not fully ruled out but is bounded and removed periodically.
Deletion works as a mirror image: the record is marked deleted inside the transaction, and an asynchronous handler deletes the object (following the outbox pattern), because "delete from S3" cannot be rolled back together with the database transaction.
Common mistakes
Large files straight into bytea. At first it goes unnoticed, then — a backup that takes several hours and memory exhaustion during parallel downloads.
Serving files through the service while S3 is right there. A controller that reads the file from the storage and streams it to the client adds unnecessary latency and load on the pool. The right way: hand out a presigned URL and let the bytes go directly.
Uploading the file inside the database transaction. If the transaction rolls back, the file in S3 cannot be taken back — an orphaned object appears. File operations do not take part in the database transaction; that is why the two-phase protocol looks exactly the way it does.
A "bare" URL in the database. If you store only the URL, there is no way to tell whose object it is and whether it can be deleted. Store the object key, the bucket name, and the status — the serving URL is generated from them every time.
A public bucket "to keep it simple". Permanent direct links instead of presigned URLs — until the first scanner brute-forces its way to someone else's documents.
When combining both options makes sense
You can almost always combine them. Small utility files — a signature, a key, a thumbnail of a few kilobytes — live in the database. User files live in object storage. Draw the line by file type and size, not with a single decision for the whole service.
In short
byteain PostgreSQL is a reasonable choice for small files (up to hundreds of kilobytes) where strict transactionality is needed.- A value over 2 KB moves into TOAST next to the same table: the dump, the restore, and every read of the column all drag it along.
- Object storage fits when files are megabytes and larger, are served to users regularly, or require a lifecycle policy.
- Presigned URLs let you serve files directly from the storage, bypassing the service.
- The two-phase upload protocol (PENDING → ACTIVE) protects against a mismatch between metadata and object, and deletion is asynchronous — the database transaction knows nothing about S3.
- Store the object key and the status rather than a ready-made URL; in one service both approaches usually coexist — small utility files in the database, user files in the storage.
What to read next
- What object storage is: bucket, object, key, storage classes and presigned URL — how S3 works from scratch.
- Spring + AWS SDK v2 for S3: integration, patterns, MinIO — upload, confirmation, MinIO in tests.
- Distributed patterns: data consistency across services — outbox and consistency between two systems in general terms.
- S3 in production: backups, replication, cost, and monitoring — what it costs to keep and serve files.