As soon as your system gets users from Europe, the acronym GDPR shows up — usually paired with the words "fine" and "data deletion." Let's look at the regulation through a developer's eyes: what counts as personal data, when it applies to you, what it actually requires from your code, and why the main tactic is the same as with cards — store as little as possible.
The key fork shows up in the "right to be forgotten": how many places you have to visit on an erasure request is decided not by the request itself, but by what you collected in the first place and where you put it.
An erasure request has to visit every place the data settled in. Without minimization, a DELETE from users closes one place out of four; with anonymized orders, clean logs and user_hash in analytics the code touches two, and the other two are closed in advance.
What it is and when it applies to you
GDPR (General Data Protection Regulation) is a European Union regulation on the protection of personal data. It applies to you if you process the data of people located in the EU — regardless of where the company and servers are. It is not an industry standard like PCI DSS, but a law with tangible fines: up to 20 million euros or 4% of annual turnover.
Key concepts:
- Personal data — any information that can identify a person: name, email, phone number, IP address, device identifier, geolocation. Special categories (health, biometrics, opinions) are singled out separately — the requirements for them are stricter.
- Data subject — the person whose data is being processed.
- Controller and processor — who decides why data is processed (the controller) and who processes it on their behalf (the processor, for example a cloud provider).
- Lawful basis — processing must have a reason from a fixed list: consent, performance of a contract, legal obligation, and others. No basis — no processing allowed.
The main strategy: store as little as possible
GDPR is built on minimization: collect only the data that is genuinely needed for the task, and keep it only as long as needed. This also removes most of the risk — data you don't have can't be stolen and doesn't need to be deleted.
- Don't collect just in case. Every form field should answer the question "why does this feature need it." A date of birth is needed to verify age — store an "18+" flag, not the date itself.
- Pseudonymization. Replace direct identifiers with a surrogate: analytics works with a
user_hash, while theuser_hash → emailmapping sits separately under strict access. A leak of the events table then doesn't reveal identities — this is the same move as card tokenization in PCI DSS. - Retention period. Every data set has its own lifespan, after which it is deleted or anonymized automatically, rather than "kept forever just in case."
Data subject rights and what they mean for code
GDPR grants a person a set of rights, and almost every one turns into a concrete task in code. Build them into your model up front — bolting a "right to erasure" onto a system where data is smeared across ten tables and logs is expensive.
| Right | What it requires from the system |
|---|---|
| Access | export all data about a person in a readable form |
| Rectification | let incorrect data be edited |
| Erasure ("right to be forgotten") | delete or anonymize data on request |
| Portability | hand over data in a machine-readable format (JSON/CSV) |
| Restriction and objection | pause processing, unsubscribe from mailings |
The "right to be forgotten" is the trickiest for architecture. Deleting a row in the users table isn't enough: a person's data has settled into orders, logs, analytics, backups, and processors. A realistic approach is anonymization instead of deletion where the record can't be removed: the order stays for accounting, but the name and contacts in it are replaced with "deleted user."
// GDPR erasure: delete where we can, anonymize where we can't (needed for accounting).
void forget(long userId) {
profileRepo.deleteByUserId(userId); // profile and contacts — subject to deletion
orderRepo.anonymizeCustomer(userId); // orders stay, but without personal data
auditLog.record("gdpr_erasure", userId); // we record the fact of erasure itself
}
Exporting data (the right to access and portability) is a walk through the same places, but read-only: gather everything tied to the user into a single export.
What this means for code
In practice GDPR boils down to a few habits familiar from secure development:
- Minimum data in the model. No field — no problem. Special categories (health, biometrics) aren't created without a clear need and a separate basis.
- Clean logs. Personal data in a request log or a stack trace is a classic violation. Same principle as with PII and secrets in logs: identifiers go into the log, not emails and names.
- Retention in the schema. A
created_atfield plus a background job that deletes or anonymizes records older than the term — not a manual cleanup "someday." - Encryption and access. Personal data is encrypted at rest and in transit, access follows the principle of least privilege, and access to it is under audit.
- Consent as data. If the basis for processing is consent, store what and when the person agreed to, so it can be shown and revoked.
Data residency
A separate requirement is where the data physically lives. For EU users, data often has to be stored and processed within the EU region. In practice this means choosing a region at your cloud provider (for example, a European AWS or GCP region) and preventing the "accidental" export of data to other regions via backups or analytics. This is an architectural decision — make it before launch, not after.
In short
- GDPR is an EU law on personal data; it applies to anyone who processes the data of EU people, with fines up to 4% of turnover.
- Personal data is anything that can identify a person (name, email, IP); processing must have a lawful basis.
- The main tactic is minimization: collect only what's needed and keep it for a limited time; data you don't have can't be stolen.
- Data subject rights (access, erasure, portability) are concrete tasks in code; build them into the model up front.
- The "right to be forgotten" is often implemented via anonymization where the record can't be deleted (orders, accounting).
- For code this means: minimum fields, clean logs, retention in the schema, encryption, access audit, and storing data in the required region.
What to read next
- PCI DSS for Developers — a related standard about card data; the same "store less" principle.
- PII and Secrets — how not to smear personal data across logs and configs.
- Auditing Administrator Actions — logging access to sensitive data.