Object storage looks like "drop it and forget it" — but without operational discipline, buckets turn into dozens of terabytes of orphaned data with an unpredictable bill. Let's look at how to set up backups correctly, protect yourself against data loss, and avoid overpaying.
Versioning changes what the operations mean: a PUT stacks a new version on top of the old one, and a DELETE erases nothing at all — it adds a marker that makes GET answer 404. So "deleted" data keeps occupying space and showing up on the bill until lifecycle collects it, while recovery comes down to removing the marker. Replication only picks up what is written after you turn it on.
S3 as a place for backups
S3 is a convenient target for backups of any system: PostgreSQL dumps, MongoDB and Elasticsearch snapshots, configuration files, encrypted secrets.
Why S3 fits this role well:
- 99.999999999% durability (11 nines) — data is stored across multiple availability zones within a single region.
- Cold storage classes — Glacier Deep Archive costs about $1 per terabyte per month; storing year-old daily backups is nearly free.
- Versioning protects against accidental deletion and ransomware.
A typical backup bucket structure looks like this:
s3://backup-bucket/
├── pg/ ← PostgreSQL dumps
├── mongo/ ← MongoDB dumps
├── es/ ← Elasticsearch snapshots
└── app/ ← application exports
On the bucket you enable: versioning, a lifecycle policy (transition to Glacier after 30 days, deletion of old versions after 90), and replication to another region.
Cross-region replication
AWS gives you 11 nines of durability, but the main enemy of data is human error: accidentally deleting a bucket, a compromised account.
Cross-Region Replication (CRR) solves both problems. Every object placed into the source bucket is asynchronously copied to a destination bucket in another region (or account). If something happens to the production account, the backup is out of the attacker's reach.
Configuration via AWS CLI:
{
"Role": "arn:aws:iam::ACCOUNT:role/replication-role",
"Rules": [{
"Priority": 1,
"Filter": {},
"Status": "Enabled",
"Destination": {
"Bucket": "arn:aws:s3:::backup-dr-bucket",
"StorageClass": "STANDARD_IA"
},
"DeleteMarkerReplication": { "Status": "Enabled" }
}]
}
An important point: replication works only for new objects — objects already in the bucket need aws s3 sync or S3 Batch Replication.
Versioning and MFA Delete
Versioning is enabled at the bucket level. Once enabled, every PUT creates a new object version, and a DELETE adds a delete marker — physically the object stays.
Hence the surprising effect: a GET after a DELETE answers 404 even though both versions are still there and still billed, and recovery comes down to removing the marker. The mechanics are visible without S3 at all:
live example
import java.util.ArrayDeque;
import java.util.Deque;
public class VersionedBucket {
record Version(String id, String body, boolean deleteMarker) {}
private static final Deque<Version> versions = new ArrayDeque<>();
private static int counter = 0;
static void put(String body) {
versions.push(new Version("v" + (++counter), body, false));
}
static void delete() {
versions.push(new Version("v" + (++counter), null, true));
}
static String get() {
Version top = versions.peek();
return top == null || top.deleteMarker() ? "404 NoSuchKey" : top.body();
}
public static void main(String[] args) {
put("dump-09-15");
System.out.println("PUT -> GET: " + get());
put("dump-09-16");
System.out.println("PUT -> GET: " + get());
delete();
System.out.println("DELETE -> GET: " + get() + ", versions kept: " + versions.size());
versions.pop();
System.out.println("marker gone -> GET: " + get() + ", versions kept: " + versions.size());
}
}
Run
Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →
In a real bucket the stack of versions is shown by aws s3api list-object-versions, and the marker is removed by delete-object --version-id with the marker's own id.
For critical backup buckets, add MFA Delete: permanently deleting a version or turning versioning off requires a one-time code from the root account's MFA device. A compromised account can add a delete marker, but not erase the data.
aws s3api put-bucket-versioning \
--bucket backup-bucket \
--versioning-configuration Status=Enabled,MFADelete=Enabled \
--mfa "arn:aws:iam::ACCOUNT:mfa/user 123456"
Object Lock is even stricter: WORM mode (write-once-read-many) makes deletion physically impossible for a set period, even for an administrator. It is used where a regulator demands it: financial and medical data.
Lifecycle policies for different cases
Lifecycle manages transitions between storage classes and object deletion. A few typical sets:
User content (avatars, uploads):
- prefix: "tmp/"
Expiration: 7 days
- prefix: "users/"
NoncurrentVersionExpiration: 30 days
AbortIncompleteMultipartUpload: 7 days
Logs and audit:
- prefix: "audit/"
Transitions:
- 30 days → STANDARD_IA
- 90 days → GLACIER
- 365 days → DEEP_ARCHIVE
Expiration: 2555 days # 7 years as the regulator requires
PostgreSQL backups:
- prefix: "pg/"
Transitions:
- 7 days → STANDARD_IA # fast recovery for the first week
- 30 days → GLACIER_IR # millisecond access, like Standard
- 90 days → DEEP_ARCHIVE # long-term storage
Expiration: 365 days
NoncurrentVersionExpiration: 30 days
AbortIncompleteMultipartUpload: 1 day
The AbortIncompleteMultipartUpload rule is worth adding everywhere: incomplete multipart uploads accumulate unnoticed and get billed too.
What makes up the cost of S3
AWS S3 bills you across three line items.
Storage
You pay per gigabyte per month. The storage class determines the price:
| Class | $/GB/month (us-east-1) | Use |
|---|---|---|
| Standard | $0.023 | Active data |
| Standard-IA | $0.0125 | Backups, accessed once a month |
| Glacier Instant | $0.004 | Archive with fast-access capability |
| Glacier Flexible | $0.0036 | Long-term archive, access within hours |
| Deep Archive | $0.00099 | Cold archive, access within a day |
100 TB on Standard is about $2300 per month. On Deep Archive, about $100 per month.
Requests
PUT/COPY/POST/LIST: ~$0.005 per 1000 requests. GET: ~$0.0004 per 1000 requests.
For most applications these are negligible amounts. A million LIST operations per day is already $5/day, or $150 per month.
A typical mistake is writing a million 1 KB files instead of one 1 GB file. On Standard the storage costs the same, but you pay for a million times more requests. In cold classes it gets worse: STANDARD_IA and Glacier Instant bill every object as at least 128 KB, Glacier Flexible and Deep Archive as at least 40 KB. A million 1 KB files in STANDARD_IA are billed as 128 GB.
Egress traffic
The most insidious line item. Traffic from S3 to the internet: ~$0.09 per GB.
100 TB per month = $9000. For public buckets with media content, this line item often exceeds the cost of the storage itself.
Ways to reduce it:
- CloudFront in front of S3 — AWS charges nothing at all for traffic from S3 to CloudFront, so you only pay for delivery to the user: $0.085/GB for the first 10 TB a month and cheaper beyond that.
- VPC Endpoint for traffic from EC2 in the same region — free. For backend services running in AWS, this should always be enabled.
Alternatives to AWS S3
| Provider | Egress traffic | Notable feature |
|---|---|---|
| AWS S3 | $0.09/GB | The standard |
| Backblaze B2 | $0.01/GB | S3-compatible API, cheap traffic |
| Cloudflare R2 | free | S3-compatible API, no traffic charges |
| Wasabi | free (with limits) | Cheap storage |
| Yandex Object Storage | ~$0.005–0.01/GB | Russian jurisdiction |
Cloudflare R2 is especially interesting for public content with high traffic: its S3-compatible API lets you switch via endpointOverride in the SDK without rewriting code.
Monitoring
CloudWatch metrics
Bucket size and object count arrive in CloudWatch on their own — once a day and for free. Request metrics are enabled per bucket separately, and they cost money:
| Metric | What it means | Reason to alert |
|---|---|---|
BucketSizeBytes | Bucket size (once a day) | A sudden spike |
NumberOfObjects | Object count (once a day) | Growth of thousands per day |
AllRequests / 4xxErrors / 5xxErrors | Requests and errors | Errors > 0.1% |
BytesDownloaded | Egress traffic | Basis for estimating cost |
Access logs
Server access logging writes each request to a text file in another bucket: IP, time, action, object key, status, size, latency, user-agent.
aws s3api put-bucket-logging --bucket production-bucket \
--bucket-logging-status '{
"LoggingEnabled": {
"TargetBucket": "audit-bucket",
"TargetPrefix": "logs/production-bucket/"
}
}'
An alternative is CloudTrail data events: JSON format, slightly higher latency, but it integrates with the rest of your CloudTrail events.
S3 Inventory
S3 Inventory generates a daily (or weekly) report on all objects in a bucket in CSV or Parquet format. Columns: key, size, storage class, last-modified date, encryption status.
It's useful for verifying that a lifecycle policy works as expected, and for finding "orphans" — objects that have no records in the database.
Common problems and protection
Account compromise and data deletion
Scenario: an attacker obtained an account with s3:DeleteObject permissions and deleted everything. The combination from the sections above saves you here: versioning keeps the data under markers, MFA Delete and Object Lock make physical erasure impossible, and replication to another account keeps a copy out of that account's reach.
Objects moved to Deep Archive by mistake
A lifecycle policy with an imprecise filter got applied to hot data. Reading from Deep Archive takes up to 12 hours and costs money.
Protection: always test a lifecycle policy on a small test prefix before applying it to the whole bucket.
Recovery: a restore-object request, waiting (up to 12 hours), then copying to Standard.
A bucket accidentally became public
Erroneous edits to the bucket policy made the whole bucket accessible from outside. This risks data leakage and a huge egress bill.
Protection: Block Public Access at the account level — a global switch that prevents any bucket from becoming public, even if someone changes a policy. Always enable it on the root account.
Millions of small files — an unexpected bill
Every event is logged as a separate PUT. After a month — millions of files and noticeable request costs.
Solution: accumulate events locally and flush them to S3 in batches every few minutes or upon reaching a size threshold. Amazon Kinesis Data Firehose does this automatically.
Production bucket checklist
The minimum set of settings that every production bucket should have:
- Block Public Access enabled (at the account level too)
- Versioning enabled
- SSE-S3 or SSE-KMS (since January 2023 SSE-S3 applies to every new object in every bucket; SSE-KMS is enabled separately)
- Lifecycle policy: aborting incomplete multipart uploads after 7 days, deleting old versions, transitioning to cold classes
- CloudWatch alert on a sharp increase in size and object count
- Access logs or CloudTrail data events for auditing
- Cross-region replication for critical buckets
- VPC Endpoint for traffic from EC2/EKS
- MFA Delete on backup buckets
- IAM with least privilege —
s3:*without restrictions is unacceptable
In short
- S3 is a durable place for database and file backups; Glacier Deep Archive lowers the cost of long-term storage to $1/TB per month.
- Versioning protects against accidental deletion and ransomware:
DELETEadds a marker, the versions stay and stay billed until lifecycle removes them. - MFA Delete and Object Lock block permanent deletion even when an account is compromised; replication to another account moves a copy out of its reach.
- Cross-Region Replication copies only new objects; for the ones already sitting there you need
aws s3 syncor S3 Batch Replication. - The bill is made up of storage, requests, and egress traffic ($0.09/GB to the internet) — and egress is what you cut first: VPC Endpoint zeroes out EC2 → S3 within a region, CloudFront zeroes out the S3 → cache leg, and Cloudflare R2 charges no egress at all.
- CloudWatch, access logs, and S3 Inventory cover the main monitoring needs.
What to read next
- Fundamentals — how S3 is built: bucket, object, key, storage classes.
- Spring + AWS SDK v2 for S3 — working with S3 from a Java application.
- PostgreSQL backups — what actually goes into such a bucket:
pg_dump,pg_basebackup, the WAL archive. - Operating Elasticsearch — S3 as a target for Elasticsearch snapshots.