← back to the section

Thursday, six in the evening. The manager asks in the chat: "So, are we shipping?" — and you answer the way it feels: "Yeah, looks fine, a couple of small bugs left." On Monday it turns out that some customers cannot pay with points, and that nobody tested card payments at all: the gateway on the test environment had been down since Tuesday.

The manager heard one person's feeling, and the decision is his: he needs a picture — what was checked, what was not and why, what we risk by shipping today. That picture, written down, is a test result report.

one and the same run — twenty checks, two of them failed 90 % passedhow it sounds in chat — ship it the same run, broken down by importance paying with pointsmoney goes missing2 of 4 failed — the money path cart and catalogueannoying but survivable6 of 6 — clean copy and iconsnot everyone notices10 of 10 same run: nothing to ship yet90 % green is the small stuff — the red sits where people pay

At the top — a run folded into a single number: twenty checks, two red, "90 % passed". Releases get approved on that line. Below is the same run broken down by importance: both failed checks landed in payments, which is half of everything about money. The percentage did not lie; it simply counted a payment check and an icon check with the same weight.

Who reads the report and what it must contain

Four readers, each with their own question. The manager decides "do we ship or not": what breaks for customers and how bad is it. The developer needs what failed, on which build, and what to fix first. Another tester, or you a month later, needs what has already been checked. Support needs what to say when the first call comes in on Monday. The first of them reads one paragraph and stops, so the conclusion goes first and the details below.

A report can be a page in the ticket or ten lines in chat, but the questions in it are always five: what it is about (build, environment, period, who ran the checks), what was checked, what was not checked and why, which defects are open and how heavy, risks and the conclusion. The first two are spoiled most often. Without a build, "18 passed, 2 failed" are numbers about nothing, and that same field kills the "well, it works on mine" argument. And "what was checked" is not "testing was performed" but which suites were run, and which marks are by hand and which from automated checks: the trust differs.

The paragraph about untested areas is the most valuable

The reader learns about tested areas from the text and about untested ones from nowhere, so silence reads as "checked". That is exactly how card payments get lost on a Thursday evening.

Three reasons, and the decisions differ. No time left — give it another day, or ship and check afterwards. Nothing to test with — no data, the gateway is down; an extra day will not help, someone else's help is needed, and this line should reach the manager as early as possible. Deliberately out of scope — the area sat outside the boundary back in the test plan; a reference is enough, but without it, in a month it looks like an omission.

The wording is a triple — what, why, what it could cost: "Card payments were not checked: the gateway on the test environment has been down since Tuesday. If something is broken there, every customer hits it at the payment step." An untested area left unwritten becomes your fault a week later; written down, it is a team decision.

Defects: not a count but a weighted list

"12 defects open" means nothing: twelve misaligned icons and twelve broken payment paths are different releases behind the same number. The weight comes from severity and priority, so the report carries a breakdown by severity, and each group unfolds into three states: what blocks the release, what is in progress and already estimated, and what we ship knowingly and until when. Without that date, in six months nobody can tell "tolerated on purpose" from "forgotten". And every defect is a record number in the tracker, not a retelling in your own words: a retelling cannot be found.

A conclusion you won't be ashamed of

Everyone reads the conclusion, and it is written last and tired. A spoiled one is a feeling ("looks fine"), over-insurance with no reason, a decision made for the manager ("we are postponing the release") or an unverifiable "quality is acceptable". A useful one is built from state, condition and a named risk.

The "Paying with points" suite passed on build 4.19: 18 checks out of 20, no blockers. Ready to ship with two caveats: card payments were not checked (the gateway on the test environment has been down since Tuesday) and with a zero points balance the redeem button stays active (BR-118, minor, in progress). Risk: if card payments are broken, every customer hits it at the last step, and rolling the build back takes about an hour.

The state rests on a specific list and build, so it can be verified in a minute, and the risk is described through whoever will hit it. A "too early to ship" conclusion is built the same way: what fails, how often, what the developer estimated the fix at and how long a re-run takes.

Metrics: what kinds exist and how they get broken

"Quality" is not measured by one number: each metric answers one narrow question and is useful as long as that question is remembered.

  • Pass rate — is this build ready. Broken by mixing statuses: "blocked" (could not be run) is not "failed".
  • Execution rate — how far along we are. Without it the previous one lies: 100 % pass rate at 40 % execution means we are at the beginning.
  • Requirements coverage — are there areas nobody looked at. A shallow check per requirement yields the same hundred percent as ten demanding ones.
  • Defect density — where trouble concentrates: defects per size of the area (a screen, a module, a thousand lines of code), or the big part always looks worse. A part rewritten last week shows more, and where nobody searched there is always little.
  • Defects found after release — also called defect leakage: how much slipped past us to customers; honest because the team is not the one counting. The "before and after" pair folds into defect removal efficiency (DRE) — the share caught before release out of everything known: nine out of ten is 90 %.
  • Defect age — does anyone fix what we find. Median rather than average, and read by severity: critical ones in three days and ordinary ones in half a year means "ordinary ones are never fixed".

What breaks metrics is a number turning into a judgement about a person — this is Goodhart's law: once a measure becomes a target it stops being a measure: improving the number is faster than improving what it measured. Count found defects as a tester's work and one defect splits into five and cosmetics go into the tracker; add a developer measured by defects let through and the two go to war. A round number as a goal breaks things the same way: "100 % coverage" yields a case for every line of the requirement, "zero open defects" yields small stuff nobody logs.

Why percentages lie without context

A percentage is short, and shortness throws away what a decision needs. The unknown denominator: "90 % passed" — out of twenty checks or two hundred, and where did the list come from; about what is missing from the list a percentage stays silent. No weight — the case from the picture; cured by one line next to it saying what landed in the remaining ten percent. No threshold and no trend: "coverage 71 %" — good or bad, there is no answer until you say "up from 63 %, against an agreed minimum of 60 %".

A short daily update and a release report

The environment died on Tuesday, the data never arrived on Wednesday, and it surfaces on Thursday evening when nothing can be done. So alongside the release report people write a daily update — the same stand-up, only in lines and with numbers.

Build 4.19: ran 12 checks from the payments suite — 10 passed, 2 failed (BR-118, BR-121). Blocked by: the gateway on the test environment has been down since Tuesday, card payments are on hold.

The value is in the third line, the obstacle: named on Tuesday it gets solved on Wednesday instead of surfacing in the release. The release report answers not "how is it going" but "do we ship", and it lives in the release ticket or in a test management tool rather than in chat, where it sinks in two days; if the daily updates were written, it takes twenty minutes to assemble. A half-hour review after the release produces the numbers for leakage and DRE.

In short

  • Four readers — manager, developer, another tester, support; the first one reads a single paragraph, hence the conclusion goes first.
  • Five parts are mandatory: what the report is about, what was checked, what was not checked and why, defects with weight, risks and the conclusion.
  • Silence reads as "checked": untested areas are written with a reason — no time, nothing to test with, deliberately out of scope.
  • Defects are shown as a breakdown by severity plus three states: what blocks, what is in progress, what we ship knowingly and until when.
  • A good conclusion is state, condition and a named risk, not a feeling that it "looks fine".
  • A metric stays honest while nobody is judged by it, and means something only with a threshold and a trend; the industry names of three of them are defect density, defect leakage and DRE.