← back to the section

Two days before the release. You have spent three days on points payment, and on Thursday a change to delivery turns up in the same release with nobody checking it: you assumed it was someone else's task, the developer assumed you would. The manager asks whether the release can go out, and you have nothing to answer with — nobody agreed on the signal that says the work is finished. That written-down agreement about boundaries is a test plan.

a plan: the scope boundary and two visible signals entry criterion build is deployed smoke is green exit criterion 18 of 18 passed no blockers left what could be checked points payment partial redemption order cancellation card payment admin reports points payment partial redemption order cancellation out of scope in scope test run 18 checks we ship not the calendar risk: no cancellation data on stagingmore likely than not — we move itthe boundary moved:cancellation — next release order cancellation

On the left the plan decides what is taken into testing at all: three items pass through the scope funnel, two honestly stay outside. Then two pairs of gates — the signals that show the run may start and that it is finished. At the bottom, a risk that fired before the run even began: the scope boundary moved, and that is a decision the plan made, not a surprise on the last day.

What we check — and what we honestly don't

"We check points payment" sounds like an answer, but it settles no argument: points payment is redemption, and returning points on cancellation, and a zero balance. So scope is written as two lists, and the second matters more: what we take on — and what we deliberately leave out this time, with one line saying why. The line "we are not touching card payment, the code has not changed" turns a future omission into an agreed decision and catches planning mistakes: the developer answers "actually, I fixed the rounding there" — and the boundary moves on Monday instead of Friday. "We test everything" is not a boundary but the absence of one.

Where we check: environment, build, data, stubs

Two people check the same feature, it works for one and fails for the other, and half a day goes into discovering they were on different builds. So the plan names things exactly: the environment address and who else uses it; the build number — without it "18 passed, 2 failed" has nothing to attach to, and a bug report will not be reproduced; the data the checks need and where it comes from. A separate line says what is a stub: the payment gateway in a test environment is usually not real, and emails go to a service mailbox — without that line someone waits twenty minutes for an email that never arrives.

Entry and exit criteria

The build arrived, you rushed to check, half of it fails — and an hour later the points service turns out never to have come up in the environment. To avoid spending days on something that was never ready, the plan holds an entry criterion: the build is deployed, smoke is green, the data is there. If login is broken the run is suspended, otherwise you get twenty red marks for one reason; the resumption signal goes in the same place: "login is fixed, build 4.19, smoke is green."

The exit criterion answers when we are done, and it is the item most often spoiled: "test everything" cannot be answered yes or no, so the work is stopped not by the criterion but by Friday. A checkable one is built from three parts: a specific list is finished ("all 18 checks from the 'Points payment' suite"), there are no open blocking defects, and the team has a decision about the rest. The word "blocking" is the key one: "no open defects at all" is unreachable on a live product, and people abandon such a criterion entirely. "90% of cases passed" does not qualify: it says nothing about which ten percent are left, and if payment was in there, it means nothing.

Risk: what has not happened yet but will break the plan

The "risks" section is usually filled in as a formality: "tight deadlines" — and on we go. But a risk is a future event, while a current problem is already a fact. You write it as a triple — what may happen, what it costs, what we do in advance: "there is no cancellation data; we lose half a run; we request it on Monday." Without the last part it stays a complaint. A risk that fires moves the scope boundary: no data means cancellation goes to the next release, and the exit criterion goes with it — which is what the animation shows.

A whole plan for one task

On an ordinary task the plan is a few lines in its description; a separate document is needed when there are several participants or when tasks from different teams went into one release.

  • Task: BON-7 "points cover no more than half the order total, 1 point = 1 unit of currency".
  • We check: redemption, partial redemption, the "no more than half" limit, returning points on cancellation.
  • We don't check: card payment — the code has not changed; admin reports — the second team takes them.
  • Where: environment staging-2, build 4.18, a customer with a balance of 500 points, the gateway is a stub.
  • Who: the checks are mine, from Tuesday; the data team prepares cancellation data by Monday; the release manager approves the release.
  • We start when: the build is on the environment and order placement works.
  • We finish when: 18 checks have been run, no blocking defects are open, the rest have a decision.
  • Risks: the cancellation data may not arrive — if it is missing by Wednesday, cancellation moves to the next release.

It is normal for a plan to change; changing it silently is not — an edit travels together with what changed and who agreed to it.

Test suite and test run

The checking is done not by the plan but by specific test cases. Once there are more than a dozen, cases are grouped in advance into test suites — sets built for one run goal: by feature, by goal (smoke, regression), by priority. One case may sit in several suites. A test run is a suite plus a build plus the marks: passed, failed, blocked. "Blocked" is not "failed": there was nothing to run the case with, so it is unfinished work, not a defect — mix the two and the report lies in both directions. The main trap is assembling a suite and never touching it: a year later regression runs forty cases about the old cart and none about points, so a suite is revisited after every large task.

In short

  • A test plan is an agreement about boundaries, not paperwork for an archive; half its value comes from the conversation that happens while it is written.
  • Scope is written as two lists, and the second matters more: "what we are not checking" turns an omission into an agreed decision.
  • An exit criterion is verifiable in a minute: the list is finished, no blocking defects are open, the rest have a decision. "Test everything" and "90%" do not qualify.
  • A risk is written as a triple: what may happen, what it costs, what we do in advance; a risk that fires moves the scope boundary.
  • A run is a suite plus a build plus the marks; "blocked" and "failed" are different lines in the report.