← back to the section

A customer put together an order for 3000 and spent 500 bonus points on it. The screen said "to pay: 2500", the card was charged 3000, and the points were gone anyway. You file a ticket, "we charge more than the screen shows", and the developer replies: "the total is calculated correctly on our side, show me where exactly it breaks".

What you can answer depends on what you can see: only the screen — you describe the symptom; the application log and the database — five minutes show which amount went to payment; the code — you find the exact line. These three modes are what people call boxes — black, white and grey: "box" is not about your job title or your level of skill, but about how many layers of the product are open to you while you hunt for a cause.

Smoke, regression and sanity often get mixed up with the boxes: they answer not "how much can you see" but "why are we running this set right now", and they are covered in the article on types of testing.

the more layers you see, the smaller the search area mode what you can see where the cause hides outsideblack boxscreenrequestlogdbcodethe whole order path insidegrey boxscreenrequestlogdbcodestep: applying the points throughwhite boxscreenrequestlogdbcodeline: the missing call grey box — where a manual tester spends most of the day

A mode is not a rank, it is how many layers are open to you. From the outside the cause can be anywhere along the order path; the log and the database narrow it down to one step; the code narrows it to a single line. That middle row is an ordinary working day for a manual tester.

Black box: working with what the user can see

Log access hasn't been granted yet, the database is closed, you don't read code — that is what the first weeks look like for almost everyone, and most findings come out of exactly this kind of work. Checking that relies only on the product's input and output is called black box testing: what you typed, what you clicked — and what you got back.

Ideas come from two sources. From the requirements: every promise turns into a check — "bonus points can cover no more than half the order" gives a check for exactly half, for half plus one, and for spending everything. From the behaviour: requirements are often missing, and then you start from what the product lets you do — the points field takes zero, a fraction, more than the account holds, letters, a minus sign; picking a dozen meaningful values out of that is what test design techniques are for.

The blind spot of the mode is an honest one: from outside you see that the result is wrong, but not where it went wrong. One screen hides three causes — not calculated; calculated but not displayed; displayed but not saved. Your ticket will carry the same sentence, "the total is wrong", while the fix lives in three different places.

White box: looking at the code itself

Branches are invisible from the outside. The function that calculates a discount may hold five "if — then" conditions, two of them nearly unreachable from outside: they fire when the bonus service returns an error, or when points are granted and spent within the same second. Checking built from the text of the program rather than the screen is called white box testing; mostly developers do it — unit tests are white box testing.

Every branch is visible, and so is which lines were executed during a run. The latter is called coverage, and it is the most dangerous word in the subject: a line was executed does not mean anyone compared the result with the requirement.

The main trap is reading the code and then checking only what is written there. The code stops being the thing under test and becomes the reference — and the reference must be the requirement: if the developer forgot the case "the account holds more points than the order costs", it is not in the code, not in your checks either, and all of them are green. The cure is the order of work: first checks from the requirements and from the behaviour, then the code — and only to add. Crossing items off because "the code has no such case" is not allowed: that is the finding itself.

Grey box: check from outside, peek inside

Between the symptom from outside and reading code lies the mode where a manual tester spends nearly all of their working time: you check the product from the outside, but you also see part of the internals. This is grey box testing, and there are usually four places to look.

The Network tab in the browser shows what was sent to the server and what came back — the cheapest way to check the "the screen shows something different from what it sends" theory. The application log on the test environment is what the service recorded about itself: took the request, called a neighbour, waited thirty seconds for an answer. The database is what was finally saved: an order with the status "draft", points spent while the total is unchanged; a couple of queries tell "not saved" from "saved and not displayed". A direct request bypassing the screen through Postman: if the total is correct when sent directly, the calculation is not the problem — what the screen sends is.

You stop describing the symptom and start naming the place: "the screen says 2500, the payment request carries 3000, and the points are already spent in the database" — a bug report like that reaches the right person immediately. The mode also decides what you can promise in a report: "checked from the outside, didn't look in the database" and "the data in the database matches" are different promises.

There is one trap here, and it is expensive: not looking at the log and guessing for an hour. The screen spins after "Pay", the tester swaps cards, signs in again, clears the cache, calls a colleague over — an hour of work, while the first line in the log says "bonus service did not respond within 30 seconds". The rule: five minutes from the outside, and if the cause isn't obvious, go inside. Not the other way round.

Looking inside is free, changing things inside is not: editing data straight in the database gives you a state the product itself cannot produce, and half a day goes into a bug that does not exist. A line from the log is an observation, not a diagnosis: "the log shows an error from the bonus service" is a fact, "the cause is in the bonus service" is already a guess.

One bug, three modes: the order with bonus points

Back to the order: the screen says 2500, the card is charged 3000, 500 points are gone.

From the outside. You repeat the path, changing one condition at a time. A different total — the same result. Points typed by hand — no bug. Points spent with the "use all" button — it reproduces every time. You never saw the code and never looked in the database, yet you found the key fact: different paths give different results, and the ticket will not bounce back with "cannot reproduce".

From the inside. The Network tab shows 3000 in the payment request while the screen says 2500. The log has a successful line about spending 500 points. The database holds the order with a total of 3000 and a mark that points were applied. The conclusion: the points were spent and the amount due was not recalculated — the search area shrank from "the whole order path" to a single step.

All the way through. The developer opens the code: the recalculation is triggered by the input field's handler, while the "use all" button sets the value past that handler. One branch that neither the unit tests nor the requirement-based checks ever reached.

A useful finding already existed in black box mode: grey box is not mandatory, it saves the team a day and spares the developer the digging.

In short

  • A box is not a job title or a level of skill; it is how many layers of the product you can see while hunting for a cause. The more you see, the smaller the area to search.
  • Black box works from the requirements and from the behaviour. It gives you the symptom but not the place: "not calculated", "not displayed" and "not saved" look identical from outside.
  • White box is checking from the text of the program; coverage means only that a line was executed. The danger is that the code becomes the reference instead of the requirement: a forgotten case will never be found.
  • Grey box is the ordinary working mode: the screen plus the Network tab, the log, the database and a direct request to the service. The five-minute rule: if it isn't obvious from outside, go inside.
  • Looking inside is free, changing things inside is not: editing data in the database breaks the conditions of reproduction, and a line from the log is an observation, not a diagnosis.