On Friday you spent an hour on the new "pay with points" feature and wrote "checked, works" in the ticket. Three weeks later a customer complains: the points were taken but the amount due never changed. What was in the order, how many points were on the account — nobody has an answer, and the case still has to be reproduced somehow.
A check written down is called a test case: what is ready before you start, which actions with which data, which result counts as correct. Memory lasts a week, and the author is not the only one who runs the check. Half the questions appear while you write it: the requirement never says what happens when there are more points than the order costs — and a change in a requirement costs a conversation, in a finished payment flow a release.
At the top — a step only its author understands: it carries neither data nor anything to compare the result against. It always unfolds the same way: what must be ready, which actions with which numbers, and one verifiable outcome instead of "it works".
What a test case is made of, and what breaks without each part
The fields look like paperwork until one of them is empty.
Identifier BONUS-14 — otherwise the case cannot be referenced in a bug or a report: "third from the top" breaks the first time the list is reordered. Title is read in a list of two hundred rows: "Order payment: 500 points cover part of the amount" is findable, "Payment check 2" is not. Preconditions — the state of the world before the first step: without them the case drowns in preparation or cannot be run at all, because the points come from nowhere. Expected result — the thing you compare against: without it "passed" is awarded for nothing having crashed. Test data: one runner takes an order of 3400, another 600, where the "no more than half" limit kicks in. Priority matters when four hours are left before a release and there are two days' worth of cases: without it the list is cut from the top and the money check goes first. Link to the requirement answers "that's by design" with a clause, not a memory of a chat.
The core is mandatory always — title, preconditions, steps, expected result.
A step someone else can repeat
A step is written for a person who is not in the room: they cannot ask what you meant, and they see the feature for the first time. Four rules remove the room for guessing. One action per step: with "open the cart, apply points and pay" on one line, the mark "step 2 failed" says nothing about what broke. Elements are named by their on-screen label — not "enable spending" but "turn on the toggle Use points": the names in the code, in the requirement and on the button rarely match. Data goes inside the step: "a valid number of points" is guesswork, "leave 500 in the How many to spend field" is data. And a step is an action, not a conclusion: "check that everything works" says neither what to click nor what to compare against, and the runner will sooner mark it passed than invent a check. Next to it sits the trap of invisible steps: "wait until the points are credited" is rewritten into a visible sign — refresh "My points" and look at the balance. You also discover the wait is sometimes a full minute, which is a question for the requirement.
Where the expected result comes from
Writing down what the product did is the easiest way and the most common mistake: such a case describes the product's behaviour in terms of that same behaviour, makes a calculation error official, and turns red exactly when the error is fixed. The result comes from the requirement — the ticket, acceptance criteria, the mock-up, the analyst's answer. It rarely answers every question a case raises: the case may be undescribed, described twice and differently, or described in unverifiable words. Inventing an answer is no better than copying one off the screen: the question goes to the analyst, and the shapes those gaps take are covered in testing against requirements and properties of good requirements.
The main property of a result is verifiability: "total 2900, status Paid, balance 0" is compared with reality without discussion, while "the discount was applied" fits any outcome. The screen is not the only place to look: that the points were deducted once and not twice is visible in a database query, the amount that actually went out in the Network tab — if the result is visible only there, the case says so.
Paying with points: a bad case and the same case rewritten
Formally there is nothing wrong with it: title "Payment check 2", preconditions empty, steps "go to the shop; place an order; pay with points; check that everything works", result "payment goes through correctly". The runner sources the points and picks the order themselves — that is, decides which case they are checking. And the case cannot fail: "payment goes through correctly" fits both spending 500 points and spending none with the amount untouched. It breaks on specifics, not on formatting. The same case with specifics:
BONUS-14Order payment: 500 points cover part of the amount. Priority high — the customer's money; requirement BON-7 "points cover no more than half the order amount, 1 point = 1 unit of currency"- Preconditions: customer
anna@example.comis signed in; the points balance is 500; the cart holds one item priced 3400; delivery is selected and free - Steps: in the cart click "Proceed to payment"; turn on "Use points"; leave 500 in the "How many to spend" field; click "Pay by card"; enter card
4111 1111 1111 1111, expiry12/30, code123and confirm - Expected result: after step 3 the "Total" line reads 2900 with the note "500 points applied"; after step 5 the "Order paid" page opens with a number; in "My orders" the status is "Paid" and the amount 2900; on "My points" the balance is 0
The preconditions pinned down the case where the limit does not fire: 500 is below 1700, and the limit is the neighbouring case's job. Data inside the steps makes two runs into runs of one and the same case. And the result became a number — while the emptied balance catches the case where the amount was recalculated on screen but the points were never deducted.
One check — one case, and why you still need several
Adding the email, the refund and the half-the-amount limit "while we're here" is an expensive temptation: "case failed" stops telling you what broke, and the case halts at the first divergence. Next to the rule one check — one case lives independence: a case must not lean on its neighbour, or the run order becomes part of the check (how that is cured). Hence several cases per feature: one positive — how things go right — and negative ones for what goes wrong; they find more bugs, since the happy path is what the developer walks themselves. Around BONUS-14 there are at least five: 500 points on an order of 600 — no more than 300 applied; an empty balance — the toggle disabled and labelled with the reason; 501 against a balance of 500 — a message, not a payment going negative; a cancellation after payment — the same 500 come back; two clicks on "Confirm" — one deduction. You do not invent those at random: the values come from test design techniques, a live buyer's chain from a scenario.
All of this is expensive — fifteen minutes of writing for a two-minute run — so on a familiar feature a checklist takes the case's place: what to check, with no steps and no detailed expectations (where the line runs). The detailed case stays where a mistake is expensive or exact repetition matters; it lives in a test management tool, and once it fails it is an almost finished bug report.
In short
- The core of a case is title, preconditions, steps, expected result; identifier, data, priority and the requirement come once there are many cases.
- A step is written for someone who is not in the room: one action, elements named as on screen, data inside the step; "check that everything works" is not a step.
- The expected result comes from the requirement, not the product's behaviour, and has to be verifiable: "total 2900" can be compared, "payment went through correctly" cannot.
- One check — one case, and a case does not lean on its neighbours; a feature needs one positive case and several negative ones: the limit, an empty balance, a cancellation, a second click.
- A checklist is cheaper and fits a familiar feature; a detailed case fits an expensive mistake or a check that will be automated.
What to read next
- Test case quality: balance, power to find bugs and typical mistakes — why a well-formatted case may still find nothing.
- Checklists and test data — when a detailed case is unnecessary and where the values for checks come from.
- Testing against requirements — the source of the expected result: acceptance criteria and questions for the analyst.
- Test design techniques: equivalence classes and boundary values — how to tell which cases are worth writing and which are pointless.