A change to the checkout form landed on your desk on Thursday evening. Writing full test cases for it makes no sense: the checking takes half an hour, describing it takes an hour and a half, and in two weeks the form will change again. So you write down what to go through: "cart", "promo code", "payment".
On Monday a colleague picks up the same list and gets stuck on the very first line. "Cart" — does that mean open it? add an item? remove the last one? He invents his own version, runs the list in twenty minutes and writes "all good". Something did get checked — just not what you meant.
A list like that is not a checklist, it is a table of contents. And that is half the trouble: without customers, orders and codes prepared in advance, even a good line stays an intention.
A requirement unfolds into lines, and every line stands on data of its own: without it the line is a reminder, not a check. That data is single-use — a promo code spent in the first run turns the same line red the next day, with nobody having touched the product.
A checklist line is a check, not a topic
The main way checklists break: the line names an area instead of a check. Everyone fills "cart" with their own meaning, and two runs of one list check different things. A working line is built from three parts: state — action — what must be visible.
| Don't | Do |
|---|---|
| cart | empty cart: the "Pay" button is disabled |
| promo code | expired code: the total is unchanged, the reason is shown below the field |
| sign-up | email already taken: "This account exists", the form keeps the data |
A good line closes with "yes" or "no"; the only answer to "cart" is "well, I looked at it".
- One line, one check. "Promo codes and points" leaves a mark nobody can interpret afterwards.
- A rule, not the route to a button. A rule survives a redesign; written-out steps do not.
- Data goes in where it changes the outcome. "$6 and 500 points" against "$34 and 500 points": the first line hits the half-of-the-total cap.
- Negative lines are mandatory. Empty, someone else's, twice, expired, too long: the developer walks the happy path himself.
How a checklist grows out of a requirement
A list written from memory repeats what you already remember. A case nobody ever discussed will never make it in — so the list is grown from the requirement: the ticket, the acceptance criteria, the mockup, the analyst's answer. Take a short one: 10% off orders of $10 and up, one code per customer, codes valid up to and including 31 August, cannot be combined with points.
It takes four questions. What must work? A $20 order with the code costs $18. Where are the edges? Every number and date has one, and that is where the mistakes live: $10 and $9.99, 31 August and 1 September (test design techniques). What is forbidden? Every constraint is at least one line. What happens around it? The order is cancelled — does the code come back or is it gone?
- [ ] $20 order with a code → $18 to pay, the discount is shown in the order
- [ ] order of exactly $10 → the discount applies; $9.99 → no discount and the reason is stated
- [ ] code on 31 August → works; 1 September → refused with a clear message
- [ ] the same code on a second order by the same buyer → refused
- [ ] code together with bonuses → either bonuses are unavailable or the code does not apply
- [ ] cancelling a paid order → the code goes back to the buyer (check with the product owner)
That line with the question mark is the most valuable one: the requirement gives no answer, and inventing one is not allowed. It stays as a question for the analyst, best asked before the code is written (reviewing requirements).
The length is measured in time, not lines: a run has to fit into one sitting, thirty to sixty minutes — five to fifteen lines for a small change, twenty to forty for a feature, sixty to eighty for a section regression. A list that cannot be finished in one go will never be finished at all. So out go the lines green for two years (candidates for automation), the duplicates ("the discount is applied" and "the total drops by 10%") and the incantations like "check the overall quality of the interface".
Data: a cast of characters, not one "user"
The line "a blocked customer sees a clear message instead of being let in" takes a minute to check — if such a customer exists. Usually he does not, and the check turns into half an hour of looking for someone you can afford to block, or is simply skipped. So the cast is created in advance:
- no orders at all — empty screens are the most productive place there is: "You have no orders yet" instead of an empty table, statistics that do not divide by zero;
- orders in every status — paid, awaiting payment, in delivery, cancelled, refunded;
- a zero points balance — the spending toggle must be disabled and labelled with the reason;
- awkward data — a long surname, an apostrophe (
O'Brien), non-Latin script, a two-hundred-character address.
Plus a blocked and a deleted customer (login closed, old links do not open someone else's data), one account per role and awkward objects: a product with one left in stock, an empty file, a file of a hundred megabytes. Keep a one-line note on each character: a month later nobody remembers why customer test-07 must not have his address changed.
Creating them by hand through the interface is fine while there are five of them. A query produces a row instantly but goes around the product's own rules: an INSERT creates an order that cannot exist — paid, with no payment behind it — and half a day goes into a "bug" that is not there. The rule is to create data the same way the product does, through its own API, and use queries only to adjust what exists: shift a date, zero a balance, clear a flag (SQL for testers, API testing).
A production dump is never taken as it is: the personal data of real people on an environment the whole team and its contractors can reach is unsafe and in most cases against the law. Anonymising means replacing the values in the data itself, not putting asterisks on the screen — the interface may hide a phone number, but the mailer will read it from the database and send a message to a real person. Emails are replaced with the example.com domain, phone numbers with numbers that cannot exist, and payment details are never exported at all.
A red line without a bug: stale data or the wrong environment
Yesterday the checklist passed end to end; today three lines are red and nobody has touched the product. The most interesting data is single-use and the very first run spends it: the "once per customer" code is gone, the order that was awaiting payment is paid, the points are spent, the sign-up email is taken. Add time (a card "valid until 2026") and a deployment with a fresh copy of the database, after which the characters are gone. Hence the rule: a check must not rest on one specific record.
- Create the state at the start of the check: need an order awaiting payment — place one now instead of hunting for yesterday's.
- Generate uniqueness:
qa+2026-09-13-01@example.comlands in the same inbox (everything after the+is dropped on delivery) and will never be taken. - Keep a batch, not a single item: not one promo code but twenty, not one order per status but five.
- Write the trait, not the number: "an order in the 'Awaiting payment' status can be cancelled" survives a rebuild, "order #1024" dies with it.
The second reason is the environment. It is not just a build but five layers: the code, the data, the settings and flags, the stubs for external systems, and the surroundings (time zone, date, browser, cache). Usually only the first one matches — hence "it works for me": his customer has three orders, yours has three hundred; his payment is a stub that always answers "paid", yours is the real sandbox that sometimes refuses.
So a red line is first checked against the data and the environment and repeated from clean — a new character, a new order, a browser window with nothing left over — and only then filed as a bug. The bug report carries the environment, the build version, the account, the order identifier and the time to the minute.
Where this applies
In an ordinary week a checklist shows up three times: after a deployment — a quick walk through the main places; before a release — a section regression with marks that feeds the report; in an exploratory session — as a frame of what must be touched. Data is prepared a week before it is needed: on deployment day nobody will quickly make twenty promo codes. The first line of a list is usefully a check of the environment itself: is the right build deployed, are the test accounts in place, is the payment stub answering.
Where beginners stumble:
- Writing topics instead of checks. Everyone runs their own list and no two runs can be compared.
- Living off a single character. After the first check he is blocked, his code is spent and his points are gone.
- Tying lines to specific numbers of orders and customers — the list dies within a month along with them.
- Filing a bug without checking data and environment — and spending someone else's time on "cannot reproduce".
In short
- A checklist line is a state, an action and what must be visible, and it closes with "yes" or "no". "Cart" is a topic; "empty cart: the Pay button is disabled" is a check.
- A line describes a rule, not the route to a button: that is why a checklist survives a redesign and a case with steps does not.
- The list grows out of a requirement through four questions: what must work, where the edges of numbers and dates are, what is forbidden, what happens around it. What the requirement does not answer becomes a line with a question for the analyst.
- A run must fit one sitting, thirty to sixty minutes; lines green for two years, duplicates and "check the overall quality" are removed.
- Data is prepared in advance and as a set: no orders, orders in every status, a zero balance, an awkward name, plus a blocked customer and an account per role.
- Create data the way the product does and use queries only to adjust it. A production dump does not go onto an environment without anonymising: the interface hides the phone number, the mailer reads it from the database.
- Data is single-use: the code is spent, the order is paid, the email is taken. A red line is checked against data and environment first and only then becomes a bug.
What to read next
- How to write a test case: steps and expected result — what to do with a line you need to hand over or automate.
- Test design techniques: equivalence classes and boundary values — which values are worth putting into your data and which are pointless to check.
- SQL for Testers — how to find and prepare data with a query without creating orders that cannot exist.
- Exploratory Testing and Degrees of Formalization — where a checklist works as a frame for free exploration, and where it gets in the way.