← back to the section

Monday morning, a message in the team chat: "New build is on the test server, please take a look." The build has one new feature — paying with loyalty points — and three fixed bugs. The release is two days away.

Where do you start? You could poke at the points feature, work through the two hundred checks that piled up over the year, or verify the three fixes and report back. Each answer is a different type of testing: the word describes not the screen you opened, but the question the run answers.

every build goes through the same three runs build 214: smoke failed — we stop right here ✗✗ catalogue won't load build 215: the new part works, a neighbour broke✓ ✓ ✗✗ search is broken build 216: all three green — ship it✓ ✓ ✓ ✓the build reaches real users smoke is it alive? new feature does it do it? regression old stuff intact? release back to the developers each run answers one question and stops the build when the answer is no

Three runs, three different questions: smoke asks whether the product is alive, the checks of the new feature ask whether it does what was promised, regression asks whether anything next to it broke. A cross at any of them stops the journey — the build goes back to the developers instead of moving on.

Functional testing: does it do what it should

The task says one line — "a customer can pay for up to 30 % of the order with loyalty points" — and it looks like a single check. But the line falls apart into a dozen questions: how many points will the product offer to spend on a 1000 order when the balance is 500, and when it's 200; what happens at exactly 30 % and one cent above it; are points taken from the original total or from the one a promo code already reduced.

That is functional testing — checking that the product does what is expected, one action at a time: "I did this, I got that", where "that" comes from the task, the mock-up or a neighbouring screen (where the reference comes from), not from your head. There are always more checks than fit the schedule, so values are picked with test design techniques: "every possible order total" becomes four or five meaningful ones. And a feature is not verified the moment the right number appears on screen: the result lives in several places at once, and points can look deducted in the cart without reaching the balance.

Positive and negative checks

A positive check is when everything is done right: enough points, a fitting amount, payment goes through. A negative check is when things are done wrong on purpose, to see how the product refuses: spend more points than you have; type a negative number; open the cart in two tabs and spend the same points twice. The point is not to break the product but to make it fail decently — a clear error, nothing typed lost, no money taken halfway.

The rule worth learning in your first week: most bugs live in negative scenarios. The happy path was walked by the developer before you, who tested the feature the way they designed it — the interesting part starts one step to the side.

Non-functional: it works, but it is unusable

The points are deducted correctly. It's just that the cart thinks for eight seconds after the click, the "Spend" button slides off the edge on a phone, and someone else's balance shows up if you change one digit in the address bar. Formally the feature works — nobody can use it.

Non-functional testing answers not "does it do it" but "how well": speed under a rush of shoppers, usability, security, work across browsers and phones, resilience — what happens if the connection drops between the points being deducted and the payment confirmed. A manual tester rarely runs this as a separate pass: they keep these angles in mind while walking through functional checks and file a finding when something stands out (non-functional types, cross-browser testing).

The smoke set: is it even worth starting

You spend half a day checking the new feature in detail, and by lunch it turns out that logging in as a second user doesn't work on that server at all: the build was assembled with a broken setting. So every new build starts with a short run through the main scenarios: does the product open, can a user log in, does search return anything, can an order be created, does payment go through.

That is the smoke set — named after plugging in an unfamiliar device: if smoke comes out, fine-tuning is not going to happen. It answers one question, is it worth starting the detailed checks, and is therefore wide and shallow: one shortest check per key feature, happy path only, 10 to 15 minutes for all of it. Builds arrive several times a day, and an hour-long run will stop being run.

Failed — the build goes back: while the catalogue won't open, the other results mean nothing. The main trap is growth: after every loud bug somebody adds a check to the set, a year later it takes an hour and people start skipping it. One rule cures it — one key feature, one check; the rest moves into regression.

Sanity check: a narrow pass after a small fix

The evening before release, a one-line change arrives: the wording of the error shown when there aren't enough points. Full regression means half a day you don't have; not checking is scary. Then comes a sanity check: not the whole product skimmed, as in the smoke set, but one area in full — spending points in every variation the change could have affected.

The difference runs along two axes: smoke is wide and shallow, across the whole build, after every deployment; sanity is narrow and deep, over one changed area. Teams use the words loosely, and "give it a sanity check" often means "have a quick look, see if anything fell apart" — arguing about terminology is pointless, asking what to run and in what time is not.

Re-testing a bug: fixed, or "fixed"

The bug moved to "fixed", which means only that the developer believes they fixed it. Someone else has to confirm it: that is re-testing, in three steps.

Make sure the fix is in this build. The bug is fixed in one build and you open a server still running yesterday's: "not fixed!" is the most common false alarm from a beginner. The build number is usually visible in the interface.

Follow exactly the steps written in the report — the same data, the same user, the same order, and on your own machine: someone else's screenshot only shows what happened on the developer's data. Hence writing a bug report in detail: in a month the person repeating those steps will be you.

Check around it — the same scenario on neighbouring data: a different amount, a different user, a cancellation instead of a confirmation. Fixes often address the described case rather than the cause, and the case next door stays broken.

If it no longer reproduces by the steps but the behaviour still looks wrong, don't close the bug as "cannot reproduce" — return it with a description of what you see now (the life of a bug).

Regression: the part nobody touched breaks

Loyalty points were added — and order confirmation emails stopped arriving: the order total calculation was touched for the points, and the receipt, the email and refunds all use it. The parts of a product are tied together more tightly than they look, which is why you ask the developer not "what did you do" but "what could this have touched".

Such a run is called regression: did our changes break something that used to work. It is confused with re-testing, though the questions differ — "was this bug fixed?" and "did the fix create new ones?"; both are done in one pass, so the second gets forgotten.

The regression suite builds itself: scenarios that have caught bugs before, and those whose failure costs the most — money, access to the account, data that must not be lost. It grows with every release until it no longer fits the schedule; then the picks go in order — what the changes touched, what is expensive to break, what has broken most often. It is also the first thing handed to automation.

And the most expensive trap: regression results live only until the next build. You ran it on Monday, five builds arrived during the week, you ship on Friday — and the ticks describe a product that no longer exists.

Release suites: what we run and when

By the fiftieth task there are hundreds of checks, and assembling them from memory means forgetting something every time. So checks are grouped into named suites: the smoke set runs after any build is deployed; the suite for a new feature lives while the task is in progress and then moves into the regression one; the release regression suite is assembled around what changed in that release. Separately, a short suite runs after the release on the live product: log in, pay, place one real order — needed because the test server and the live product differ in settings, data and external services.

Two things are agreed in advance: start after smoke has passed, or you are carefully checking a build already known to be broken, and finish not "when the cases run out" but "when blocking bugs are closed and there is a decision on the rest, whether we ship with them or not" (the test plan, test management tools).

All of it is the working language of a team: "smoke passed, I'm on the new feature", "re-check that bug on build 215", "we need regression on payments by Friday" — and you answer in the same words, not "I looked at everything".

In short

  • A type of testing is not a screen or a tool but the question a run answers: functional asks "does it do what was promised", non-functional "how well".
  • Most bugs live in negative checks: the happy path has already been walked by the developer.
  • The smoke set is wide and shallow, 10 to 15 minutes per new build; one that grew to an hour stops being run.
  • A sanity check is narrow and deep over one changed area; the word is used loosely, so agree on the scope out loud.
  • Re-testing means the same steps from the report on the build with the fix, plus the data next door; regression asks a different question — did anything next to it break.
  • Any run result belongs to a build number: once a new build arrives, part of the checks is out of date.