As soon as your system gets users from Europe, the acronym GDPR shows up — usually paired with the words "fine" and "data deletion." Let's look at the regulation through a developer's eyes: what counts as personal data, when it applies to you, what it actually requires from your code, and why the main tactic is the same as with cards — store as little as possible.

The key fork shows up in the "right to be forgotten": how many places you have to visit on an erasure request is decided not by the request itself, but by what you collected in the first place and where you put it.

“delete my data” — where it already sits forget(userId = 42) profile orders logs analytics DELETE executedcontacts erasedname and phonestill sit inevery orderemail in logsand stack tracesevents keyedby emailprofile is empty — name and email remain in three places out of fourthe erasure request is not fulfilled deleteprofile, contactsanonymizename → “deleted”kept for accountinguser_id onlyno names, no emailnothing to cleankey is user_hashthe email linklived in profileand is gone with itthe code touches two places: delete the profile, anonymize the orderlogs and analytics need no fixing — names and email never got there how many places you visit is decided by the data model, not the requestwhat you never collected, you never have to find or delete

An erasure request has to visit every place the data settled in. Without minimization, a DELETE from users closes one place out of four; with anonymized orders, clean logs and user_hash in analytics the code touches two, and the other two are closed in advance.

What it is and when it applies to you

GDPR (General Data Protection Regulation) is a European Union regulation on the protection of personal data. It applies to you if you process the data of people located in the EU — regardless of where the company and servers are. It is not an industry standard like PCI DSS, but a law with tangible fines: up to 20 million euros or 4% of annual turnover.

Key concepts:

  • Personal data — any information that can identify a person: name, email, phone number, IP address, device identifier, geolocation. Special categories (health, biometrics, opinions) are singled out separately — the requirements for them are stricter.
  • Data subject — the person whose data is being processed.
  • Controller and processor — who decides why data is processed (the controller) and who processes it on their behalf (the processor, for example a cloud provider).
  • Lawful basis — processing must have a reason from a fixed list: consent, performance of a contract, legal obligation, and others. No basis — no processing allowed.

The main strategy: store as little as possible

GDPR is built on minimization: collect only the data that is genuinely needed for the task, and keep it only as long as needed. This also removes most of the risk — data you don't have can't be stolen and doesn't need to be deleted.

  • Don't collect just in case. Every form field should answer the question "why does this feature need it." A date of birth is needed to verify age — store an "18+" flag, not the date itself.
  • Pseudonymization. Replace direct identifiers with a surrogate: analytics works with a user_hash, while the user_hash → email mapping sits separately under strict access. A leak of the events table then doesn't reveal identities — this is the same move as card tokenization in PCI DSS.
  • Retention period. Every data set has its own lifespan, after which it is deleted or anonymized automatically, rather than "kept forever just in case."

Data subject rights and what they mean for code

GDPR grants a person a set of rights, and almost every one turns into a concrete task in code. Build them into your model up front — bolting a "right to erasure" onto a system where data is smeared across ten tables and logs is expensive.

RightWhat it requires from the system
Accessexport all data about a person in a readable form
Rectificationlet incorrect data be edited
Erasure ("right to be forgotten")delete or anonymize data on request
Portabilityhand over data in a machine-readable format (JSON/CSV)
Restriction and objectionpause processing, unsubscribe from mailings

The "right to be forgotten" is the trickiest for architecture. Deleting a row in the users table isn't enough: a person's data has settled into orders, logs, analytics, backups, and processors. A realistic approach is anonymization instead of deletion where the record can't be removed: the order stays for accounting, but the name and contacts in it are replaced with "deleted user."

// GDPR erasure: delete where we can, anonymize where we can't (needed for accounting).
void forget(long userId) {
    profileRepo.deleteByUserId(userId);          // profile and contacts — subject to deletion
    orderRepo.anonymizeCustomer(userId);         // orders stay, but without personal data
    auditLog.record("gdpr_erasure", userId);     // we record the fact of erasure itself
}

Exporting data (the right to access and portability) is a walk through the same places, but read-only: gather everything tied to the user into a single export.

What this means for code

In practice GDPR boils down to a few habits familiar from secure development:

  • Minimum data in the model. No field — no problem. Special categories (health, biometrics) aren't created without a clear need and a separate basis.
  • Clean logs. Personal data in a request log or a stack trace is a classic violation. Same principle as with PII and secrets in logs: identifiers go into the log, not emails and names.
  • Retention in the schema. A created_at field plus a background job that deletes or anonymizes records older than the term — not a manual cleanup "someday."
  • Encryption and access. Personal data is encrypted at rest and in transit, access follows the principle of least privilege, and access to it is under audit.
  • Consent as data. If the basis for processing is consent, store what and when the person agreed to, so it can be shown and revoked.

Data residency

A separate requirement is where the data physically lives. For EU users, data often has to be stored and processed within the EU region. In practice this means choosing a region at your cloud provider (for example, a European AWS or GCP region) and preventing the "accidental" export of data to other regions via backups or analytics. This is an architectural decision — make it before launch, not after.

In short

  • GDPR is an EU law on personal data; it applies to anyone who processes the data of EU people, with fines up to 4% of turnover.
  • Personal data is anything that can identify a person (name, email, IP); processing must have a lawful basis.
  • The main tactic is minimization: collect only what's needed and keep it for a limited time; data you don't have can't be stolen.
  • Data subject rights (access, erasure, portability) are concrete tasks in code; build them into the model up front.
  • The "right to be forgotten" is often implemented via anonymization where the record can't be deleted (orders, accounting).
  • For code this means: minimum fields, clean logs, retention in the schema, encryption, access audit, and storing data in the required region.