A customer placed an order, the money was charged, and the confirmation email never arrived. Three messages in the chat: "the email function is covered by tests", "the mail service replies, the request is in its log", "I walked that scenario this morning and the email arrived". All three are telling the truth — they checked different things: a piece of code, the neighbouring service, the customer's whole path.
The scale of what you check in one go is a test level. These are not calendar stages but the size of the piece in your hands when you press "check": all four levels of one task can happen on the same day or stretch over months, depending on the development model. The questions differ, and an upper level never replaces a lower one:
- unit — does this function calculate correctly?
- integration — did two neighbours read the same agreement the same way?
- system — did the customer reach the end and get what they expected?
- acceptance — is this what was ordered, and can we release it?
We'll walk through each on one shop with a promo-code discount, a card payment and a confirmation email.
Going up, it isn't difficulty that grows but scale: one function, the joint between two pieces, the customer's whole path, the decision to ship. The higher the check, the closer it sits to the user — and the more every run of it costs.
Unit testing: one function under a magnifying glass
Inside, a program is assembled from small pieces, and the most common is a function: a named chunk of code that takes something in and gives a result back. "Give me the order total and the promo code — I'll give you the discount." Checking one chunk is unit testing: the function is called directly, with no browser, so the check runs in milliseconds — teams write thousands and run them automatically on every code change, long before it reaches a test environment. Feed it a negative total when the rule says "discount applies from 1000" and you expect a clear error, not a discount that turns into a payout. Developers write these, and one thing matters to you: a unit test checks the code as the developer understood it. If the requirement was misread, the test is green and the behaviour still wrong; "covered by tests" only means those lines were executed, not that anyone compared the result with the requirement.
Stubs and mocks: checking a piece away from its neighbours
The discount function goes to the database for the promo code, and a check cannot go there for real: the code expires tomorrow and the test turns red with no change to the code. So the neighbour gets replaced. The piece that always answers "what is promo code LETO25" with "25% off, valid until the end of the year" is a stub: it does one thing, return a pre-recorded answer. A mock is the same replacement with a memory, and why that matters shows in the requirement "a double click must not charge twice": the outcome is identical, the order is paid, and the only difference is how many times the payment service was called — and the real one cannot be called, that is real money. So the mock gets asked: how many times did anyone reach you? Once is fine, twice is a bug.
The same word covers replacements on a test environment: "payment is stubbed on staging" means that behind the "Pay" button sits not a bank but a piece that always answers "success" half a second later. A declined card, "the bank thinks for forty seconds", a double charge cannot be reproduced there: if your case "passed", you checked the stub, not the payment, and that is how it goes into the report — "payment checked against a stub; real card declines were not checked". A real neighbour answers slowly, sometimes with an error, sometimes stays silent until the wait runs out: a whole class of bugs lives on that border and is caught one level up.
Integration testing: checking the joint, not the piece
Two pieces were checked separately, both green; put them together and it doesn't work. The joint has content of its own: the agreement about what one piece hands the other and in what shape. The cart calculated 3400 roubles and sent the payment service the number 3400, while that service expects the smallest unit and reads it as 34 — both sides have green tests, and the customer pays 34 for a 3400 order. Checking two or more pieces wired together is integration testing, and joint bugs come down to four questions: what shape the data is in (roubles or cents, 05.09.2026 or 2026-09-05), who calculates what (the cart applied the discount, payment applies it a second time — "the discount doubled"), what to do when the neighbour refuses (an error or thirty seconds of silence: who retries, and will the retry charge twice) and order and repeats (the email went out before the order was stored, so it carries an empty number). Your half is where the joint shows from outside: a request to the service, the Network tab, a query to the database. The trap is always the same, checking only the happy reply — and a precise description of the joint makes a bug report land with the right person.
System testing: the buyer's whole path
Every joint has been checked pairwise, and the customer still fails to reach the end: a scenario is not the sum of its joints, because every step depends on what piled up before it. A system check takes the whole product and walks the path from finding an item to the confirmation email with no replacements: the real interface, the real database, real neighbours or their official sandboxes. What gets caught here is what separate pieces can't show: data dragged along the whole path (the address was chosen at step two and the email shows the profile one), state that changes on the way (the promo code expired between cart and payment), going back and breaking off (press "back" from payment, change the quantity, pay — which amount was charged). This is also where "how well" is measured: non-functional checks almost always live at the system level. It is the core work of a manual tester, and its enemy is how the test environment differs from production: stubbed payments, a hundred items instead of a million, one user instead of a thousand; that list of differences belongs in the report.
Two mirror-image cases explain why an upper level never replaces a lower one. The discount is correct for all forty values and the unit tests are green — yet the payment screen shows the amount without it, because the price was passed before the discount was applied: the function is innocent, it was asked in the wrong place. The reverse: the scenario passes end to end and is marked "passed" — while a rounding error hides in the formula and surfaces only on amounts ending in five cents, and the scenario honestly walked one amount out of a hundred. So when something is "checked", ask what was checked, a function or a path; the answer decides what goes into your test report.
Acceptance testing: the call to ship
A product can be assembled without a single bug and still not ship: the mistake was made earlier, in deciding what to build. So at the end comes a check of a different kind: not "are there defects" but "is this what was needed, and can we release it". It is acceptance testing, run against acceptance criteria — conditions written down in advance: the discount applies to the goods total without delivery, two codes cannot be combined. Without criteria agreed in advance, acceptance turns into an endless stream of taste-based edits. The customer or the future users do the checking, hence the name UAT (User Acceptance Testing). Beside it stands operational acceptance, run by the people who keep the thing alive: how to roll back a bad update, where the logs are at night, whether a backup exists and whether anyone has restored from it.
Your job at acceptance is to prepare the environment and the data and to sort the remarks into two piles: it behaves differently from what was agreed — a bug for a developer; it behaves exactly as agreed but the customer now wants something else — a requirement change with its own deadline. Leave the piles unsorted and the team gets two dozen "urgent bugs", half of them new wishes.
The pyramid, and why it flips upside down
The same error in the discount formula can be caught at any level — the difference is price. A thousand unit checks finish before you've opened the browser; the same error through the interface means minutes for the environment, the login, the cart and the promo code, and the check breaks when the button moves to another corner though nobody touched the calculation. Hence the test pyramid: many cheap unit checks at the bottom, fewer integration ones above, very few interface runs at the top. In real life it often flips into an ice-cream cone: unit checks are few — nobody has time to write them, an older product doesn't come apart — reliability is topped up by manual runs, and you feel it as hours-long runs before a release and failures caused by a moved button. From there it becomes a conversation about automation, starting with a number like "re-checking discounts costs two hours every release".
In short
- A level is the size of the piece you check in one go, not a calendar stage; the upper one never replaces the lower.
- A unit check calls one function and runs in milliseconds; a green run only means the code does what the developer intended.
- A stub returns a pre-recorded answer, a mock also remembers how many times it was called; "payment is stubbed" means declined cards and double charges cannot be reproduced there.
- An integration check is about the agreement at the joint: data shape, who calculates what, the neighbour's refusal, repeats. A system check is about the whole path, and its enemy is how the test environment differs from production.
- Acceptance answers "is this what we wanted and can we ship" against criteria written in advance; beside it sits operational acceptance — how to roll back, whether a backup restores.
- The pyramid is about the price of a check: with no cheap bottom it becomes an ice-cream cone with hours-long runs.
What to read next
- Software development models: waterfall, the V-model and iterations — where the "project stage ↔ its own level of checking" pairing comes from.
- Types of testing: functional and non-functional — what exactly is checked inside each level.
- Manual and automated testing — why the lower levels are handed to code, and what a green run means.
- API testing basics in Postman — how to look at a joint without opening the interface.