You paid for an order in a delivery app, and the money was charged twice. Almost always that could have been spotted earlier: someone had to walk the same scenario before you and say "stop, this is wrong".
Testing is checking that a program does what is expected of it, and doesn't do what it shouldn't.
Testing is one comparison: the expected result from the requirements against what actually happened. A match means the check passed; a mismatch is a bug.
What "checking" actually means
You press a button, something happens. Whether it is correct the screen won't tell you: an app shows a wrong total as calmly as a right one. The answer comes one way only: compare the actual result against a reference, against "how it should be". A match means the check passed; a mismatch is a finding for the team. By the reference, a taken email address gives the error "this address already exists"; a form that freezes or creates a second account is a bug.
Where does the reference come from when nobody wrote anything down? You look for it on a ladder, stepping down only when the current rung has no answer: the task description and acceptance criteria → the designer's mockup → how the feature works in production today → a similar screen next door → common sense (the app doesn't crash, money doesn't disappear) → a question to the analyst. No answer on the last rung is itself a finding: a hole in the requirements, caught before a developer guessed for everyone. The worst move is to invent the reference yourself and compare the product with your own imagination.
Checking doesn't stop at the screen: one action leaves traces in several places — a message, a database row, an email, a service response. The screen says "order created" while no row exists, and that bug is not caught by eye: hence developer tools, SQL and service requests later in the program. Checks are positive (everything entered correctly) or negative — an empty field, letters instead of digits; most bugs live in the second kind, and designing them is what test cases teach.
Why programs break
Most often it isn't the code that breaks but the understanding. The analyst wrote "a 10% discount is applied to the order", meaning the order including delivery: that is what the meeting decided, but it never made it into the sentence. The developer calculated from the price of the goods — delivery isn't goods. The code runs without a single error, a check from the same sentence passes, and a customer with an order of 3,000 plus 400 for delivery gets 300 off instead of 340.
The rest is simpler: people make mistakes (a flipped sign, a forgotten empty field); the product changes — you fix one thing and bump into its neighbour; the real world is unpredictable — slow internet, a name with an apostrophe, "Pay" pressed twice. Bugs are the norm, not an emergency: the question isn't "are there any" but "will we find them first". The chain "human error → defect → failure" is unpacked in what a bug is.
Exhaustive testing is impossible
An order form: three fields (name, phone, address) with five meaningful values each — normal, empty, too long, special characters, spaces at the edges — give 5 × 5 × 5 = 125 combinations. Add a delivery date thirty days ahead and four payment methods: 125 × 30 × 4 = 15,000. At a minute per check that is 250 hours, more than a month without a day off, for one form. And the "order comment" field accepts infinitely many values.
Hence the core of the job: testing is choosing what to check first, not an attempt to cover everything. The choice isn't random: test design techniques squeeze thousands of combinations into a dozen, and decision tables and pairwise handle options that affect one another.
Two effects help. Bugs cluster where the logic is complex, where the code changed recently, where a newcomer worked: three bugs found in checkout — look there for the fourth. A set of checks goes stale: an old checklist catches less and less, what it found was fixed long ago and new errors hide where it never looks. This is the pesticide paradox (insects get used to one spray); the cure is refreshing the set and exploratory testing.
The cost of a bug grows over time
The same discount costs the team wildly different sums depending on the stage it was noticed at.
- Requirements. The question "before or after delivery?" costs five minutes and one edited line: the bug never exists.
- Development. The test environment shows the wrong total, the developer fixes the formula in code written yesterday: an hour of work, nobody outside the team noticed.
- Before release. Besides the formula, everything nearby is re-checked: the cart, receipts, accounting reports, emails; the release slips by a day or two.
- After release. A nighttime patch, manual reconciliation of orders already created, refunds to everyone who overpaid, replies in support.
It grows because the later it is, the more has been built on top of the mistake: neighbouring features lean on the wrong formula, checks and documentation were written from it, support quotes it. Then come corrupted data: the database filled with wrongly discounted orders, and fixed code won't heal them — old ones are repaired in a separate pass. And some customers have already seen the error, a line of the bill no code can fix.
Testing is not debugging
A tester finds a problem and describes it so that it can be reproduced: with these steps you get B instead of A. Debugging is what the developer does next: digs into the code, finds the cause, fixes it. A tester doesn't need to know which line hides the error, only how to show it's there. And a tester doesn't break the product, they show it is already broken under certain conditions: the bug existed before them, nobody had met it.
What counts as quality
Zero bugs doesn't happen — even one form can't be checked exhaustively. Quality is the absence of serious problems on the scenarios that matter: a misaligned label in the settings can wait, an error on payment never can. How serious "serious" is depends on the product: in a bank transfer a wrong penny is a disaster, a game is forgiven a stuttering animation.
Verification asks "did we build what the requirements describe?", validation asks "did we build the right thing?". The formula matched the requirement exactly and every check passed — verification done; but the customer expected a discount on the whole order — validation failed.
The seven principles of testing
Everything above is scattered across sections, while the industry long ago boiled these observations down to seven points and refers to them by name — in arguments about how much more to check and what a green run really means.
- Testing shows the presence of defects, not their absence. A green run speaks about those checks and nothing else.
- Exhaustive testing is impossible. 15,000 combinations on a single form — hence the choice of what to check first.
- Early testing. A question about the requirements costs five minutes; the same mistake after release costs a nighttime patch and refunds.
- Defect clustering. Three bugs found in checkout — look there for the fourth.
- The pesticide paradox. The same set of checks stops finding anything new over time.
- Testing is context dependent. A bank transfer and a game are checked differently: the price of an error differs.
- The absence-of-errors fallacy. A product without a single bug that nobody needed has helped no one.
Next to them sits one more split — by whether the program is running while you check. Reading requirements, a mockup or code without running it is static testing; everything where the application is running and someone acts on it is dynamic testing. Static testing is possible before any code exists, which is why it yields the cheapest findings — exactly the ones the third principle is about.
In short
- Testing is one comparison: how it should be against what happened; a mismatch is a bug. And it goes past the screen: an email, a database row, a service response are the same check.
- The reference is found on a ladder: the task, the mockup, current behaviour, a similar screen, common sense, a question to the analyst. No answer anywhere is a finding, not a licence to invent one.
- You can't cover every combination even on one form: 15,000 variants at a minute each = 250 hours. The job is choosing what to check first.
- The later a bug is found, the more it costs: code and documentation are built on top of it, and the database holds corrupted data.
- Quality isn't zero bugs but the absence of serious problems on the scenarios that matter; a green run proves only that those checks passed — the first of the seven principles.
What to read next
- Who a tester is and their role on the team — the working day, and how QA differs from testing.
- Manual and automated testing — what a human checks and what code takes over.
- The software lifecycle and where testing fits — where the early, cheap findings appear.
- Test design techniques — thousands of combinations boiled down to a dozen.