There are four hundred and twenty cases in the suite, a full run takes two days, and the report shows thirty-one failures. Exactly one of them is a real bug: in the rest a button moved, the test data ran out, and eight cases check a section removed back in the winter. The anatomy of a case is flawless here. The problem is not the formatting but the properties the cases do not have.
A suite ages on its own even if nobody touches it: a button label changed, the test data was deleted, a feature was removed — and three cases out of seven fail for reasons that have nothing to do with the product. After the clean-up there are fewer cases, but every one of them actually checks something.
A detailed case is not the same as a good one
A twelve-step case looks more solid than a four-step one, but detail answers "how do I run this," while quality answers "what can this case catch." That property is fault-finding power. "Start the application; it is expected to start" has none: startup happens hundreds of times a day, and a break would be noticed without any case. "Start the application with the working folder path pointing to a nonexistent drive" has plenty: nobody lands here by accident, and the developer wrote that branch once.
That answers the question about size too: twelve steps, nine of them preparing data, is not detail but missing preconditions. The question to ask any case is the same: what has to break in the product for this case to go red? Writing a case costs half an hour, running it three minutes every release, fixing it after a redesign another half hour — for ten cases at once. So spend detail where an error is expensive: calculations, money, access rights. Elsewhere a checklist is cheaper.
A case that cannot run on its own
Monday's run was green; on Tuesday five cases failed, and the product had not changed — the order had: two people split the suite and went in parallel. A hidden dependency is worse than an explicit one: the precondition "a paid order exists" looks honest, except that the order appears only because case TC-12 sits higher and somebody runs it. There are three kinds: shared state, a shared account test17 (the "delete profile" case breaks the next twenty) and a single-use resource — a promo code, a balance of 500 points, three seats on a flight. The fix: the precondition describes the state of the world and somebody actually prepares it, the case owns its data, and a chain that is genuinely the point becomes an explicit sequential suite.
Repeatability is the same trouble from the other side: a case must run many times in a row with the same outcome. It is destroyed by unique data (cured by generation: anna+2026-03-14-01@example.com), by state that was changed and never restored (cured by a postcondition — every test management system has a field for it, and it is the emptiest one) and by a consumed resource. Checking both properties takes a minute: run the case first on a fresh environment, then immediately a second time.
A fragile case and a durable one: the same check written twice
A case fails either because the product has an error or because the case is tied to something allowed to change. The second kind is pure loss: time spent, nothing to show. Take one check — a customer requests a return for a delivered order — and write it twice.
Fragile. "Return #4". Preconditions: sign in as test17@example.com, password qwerty123; order 10482 dated 14 March exists in the database. Steps: open https://stage-7.shop.internal/orders/10482; click the second grey button from the top on the right; pick the third item in the drop-down; click "Submit." Expected: a green banner "Request #8 accepted."
Durable. RET-3 "Return of a delivered order: the request is created and visible to the customer", requirement RTN-2 "a return can be requested within 14 days of delivery". Preconditions: a customer from the returns-ok data set with one delivered order, delivered yesterday. Steps: open that order's page under "My orders"; click "Request a return"; select the reason "Wrong size"; submit the request. Expected: the request page opens with its number, under "My returns" it is listed as "Under review" and points to the same order, and the order's status becomes "Return requested."
The second case is no longer than the first — it carries six edits, and each removes a way to fail without a bug:
- The environment address left the steps: rename the environment and the case quietly checks the wrong build.
- The element is named by its label, not its position: "the second grey button from the top" vanishes on a reshuffle, and "the third item in the drop-down" starts checking a different return reason.
- A specific record was replaced by data properties: order
10482lives until the next rebuild, andtest17is shared by the whole team. - The password left the text: credentials change more often than anything else.
- The date became relative: "dated 14 March" drops out of the 14-day window in a month, and the suite goes red by the calendar.
- The expected result does not demand a literal match: the colour and wording of the banner change without a bug, and the number
8depends on how many requests exist on the environment.
The case did not become vague. The rule is this: pin down the data that decides which branch of the requirement you are in — 500 points against an order of 3,400 checks partial redemption, not the "no more than half" cap; describe everything else as a property ("a file name with spaces and non-Latin characters"); and the environment, the build and the account belong to the run, not to the case.
A suite ages by itself: duplicates and clean-up
The product changes every week while the cases stay six months old. First come the cases that "always fail for a known reason" and get marked "passed" from memory, then the share of blocked ones grows — and a green run stops meaning anything, even though "380 of 420 passed" still feeds the release decision. An important detail: if the preconditions could not be prepared (no user, the environment is down), the case did not fail — it is blocked, the product has not been checked at all. Marked as "failed", such runs paint an epidemic of bugs in the report out of thin air.
Three habits keep a suite alive: cases are fixed in the same task as the behaviour; a failing case is investigated down to its cause rather than having its expectation adjusted to the product (after that adjustment the case is green forever and the bug is the norm — expectations follow the requirement); and the suite has an owner who works through repeated failures once per release.
Duplicates come from different authors and titles, from copying a case out of smoke into regression verbatim, and from mechanical use of test design techniques: −1, −5 and −100 against a "not less than zero" boundary sit in one equivalence class, so one is enough. Group cases under requirements and duplicates end up side by side — along with the clauses that have no cases at all. Keep the stronger one and merge the expectations. Delete a case when what it checks is gone, when the check moved to a lower level, and when it checks somebody else's code — the browser, the operating system. Deletion goes to the archive, with a reason and a date, and never alone.
Reviewing cases in pairs
Authors cannot see their own cases: they remember the state of the environment and fill the gaps in their head. Fifteen minutes of reading together catches what would otherwise surface a month later mid-run. The questions come in this order, wording last:
- Where the expected result came from — the requirement or the screen? A case copied off the screen stays green right up to the day the error is fixed; invented behaviour is cured by a question to the analyst.
- Can the case be executed without asking the author anything. The reader says the steps out loud while the author stays silent; every "and how does this work?" is a finding.
- What happens if the case runs at a different time and place — first in an empty run, twice in a row, in a month and on another environment.
- Will the case be found by search, and is it a duplicate.
And only then the wording: impersonal verbs, present tense in the expected result. One more rule belongs here, the kind you confirm at a glance: the expected result always describes the product working correctly. Even in the harshest negative case the expectation is not "the application crashes and loses data" but "the message 'Cannot save the file: not enough free space' appears."
In short
- Case quality is not detail but fault-finding power: what has to break in the product for this case to go red?
- A case must not lean on its neighbours and must run twice in a row. Check it: run it first on a fresh environment, then immediately again.
- Fragility is a tie to whatever is allowed to change: a button's position, a database record, a calendar date, the number of a new request.
- Pin down the data that decides the branch of the requirement; describe the rest as a property. The environment, the build and passwords belong to the run.
- "Blocked" and "failed" are different: the second means a defect was found, the first means the product has not been checked yet.
- Delete a case when what it checks is gone, when the check moved to a lower level, or when it checks somebody else's code: archive it, with a reason, never alone.
What to read next
- How to write a test case — the anatomy of a case and where the expected result comes from.
- Checklists and test data — when a detailed case is not needed, and where to get data instead of one record.
- Test management: TestRail, Qase, Zephyr — where the suite lives and how archiving differs from deleting.
- User scenarios and case suites — when dependence between cases is legitimate.