"Instalment payments are on staging, please take a look." There are no requirements — only a thread with the analyst and two mockups on the ticket. There are no test cases either, and nothing to write them from. The demo to the customer is a day and a half away.
One person clicks everything for an hour and writes "seems to work." Another spends forty minutes and finds three places where the amount due disagrees with the price on the product page. It is not about seniority: the second one had a frame — a goal, a time limit, and an agreement written down in advance about what counts as a finding. That is exploratory testing: you invent the checks as you go, but not at random.
A test case leads you in a straight line: what is written is what gets checked, and the report adds up. A session follows the same path right up to the first oddity — then turns off to where no list exists. The price of that turn: part of the route stays unvisited, and that goes into the report too.
Three degrees of formalization: scripted, exploratory, at random
"Go check the new feature" is done in three different ways, and one thing separates them: where the list of checks came from. That is the degree of formalization.
Scripted. The steps of the test cases are written in advance: nothing is forgotten, anyone on the team can repeat the run. The price is high: a case verifies exactly what is written in it, and a bug one step off the path stays unseen.
Exploratory. The list is born as you go, each step prompted by the previous one: a long name is truncated on the order screen → is it truncated in the customer email too? → does search by that name find the order at all?
Ad hoc. No goal, no list, no notes — ten minutes of fresh eyes on a build that has just landed. Reconnaissance, not a method.
People confuse the second with the third, and the difference shows before the first click: a session has a goal, a time limit and a trail.
The session charter: goal, time limit, what counts as a finding
The boundaries are written down beforehand, in three or four lines; that note is called a charter.
Exploring: applying a promo code in the cart — stacked discounts,
expired and other people's codes, repeated entry.
How long: 40 minutes, staging, build 217, Chrome.
A finding: total disagrees with the product page; total changes with
no action by the customer; an error message that says nothing.
Not today: delivery and card payment — that is a separate session.
A goal that is too wide ("payments") turns the session into wandering; one that is too narrow ("the Apply button") leaves nothing to branch into. The limit is 40 to 90 minutes: less and you never get going, more and attention fades. Without the "a finding" line you spend half an hour arguing with yourself whether something is a bug or intended. "Not today" keeps the scope from creeping.
Notes as you go: so the bug reproduces later
The most painful loss is a finding you cannot repeat. It is not about memory: in a test case the state at each step is known in advance, while in a session you spent twenty minutes accumulating it. The cure is one habit — write the line before you click.
Notes taken as you go: the time, what you did, what you saw. Someone else can reproduce a bug from a record like this — from "the discount applies twice" only the author can.
A bug report assembles itself from those four lines in a minute. Look at 14:07: repeating from a clean state has to be done right away — it is what turns "I saw something" into a finding.
Two things cannot be recovered afterwards. The data — not "a product" but "Coffee grinder, 3400, code LETO25": half of all irreproducible bugs hang on one specific value. The time, to the minute — a developer uses it to find your request in the server log.
Where the ideas come from, and in what order to use them
"Explore the product" sounds to a beginner like "make it up yourself." Stock moves help — the ones that apply to almost any screen.
- Boundaries — zero, one, the maximum, the maximum plus one: an empty cart, 99 items against a limit of 99, a total one cent below the free-delivery threshold (test design techniques).
- Repeat — press the button twice, enter the code again: developers plan for one call, and the second one produces a duplicate — two orders, two discounts, two emails.
- Abandon halfway — walk away from a half-filled form, close the tab on the payment step: what is left behind is a stuck reservation or points deducted but never credited.
- Other people's data — someone else's order number in the page address: not a broken screen, but somebody's data shown to a stranger.
Next come the back button and two tabs (a cached page still showing the old total), throttling and cutting the network in DevTools, and time around deadlines — a reservation held for 15 minutes, an operation at 23:59. All of that is error guessing: the collection fills up fast if after every finding you ask where else the same mistake could live.
The order is kept plain: simple positive checks first, then simple negative ones (an empty field, obviously invalid data), then complex positive (two promo codes, a partial refund) and only at the end complex negative. Start with the exotica and you get a product that survives clever attacks and fails in the elementary case. The main path also doubles as your map: an oddity is only visible against the ordinary.
What is left after a session
Going through the notes is the last four minutes of the same forty. In our example the question "where else could the same mistake live?" led to bonus points: the code plus the points brought the amount due to "0", and the order went through. Four things come out of a session:
- findings — bugs filed in the tracker, the money one first;
- coverage — "promo codes done, delivery and card payment untouched"; the second half matters more than the first;
- ideas for cases — everything that reproduces reliably;
- questions for the requirements — "a lowercase code is not accepted, is that intended?".
How much went into checking and how much into filing is added if the team tracks metrics.
Where this applies
In a sprint the sequence is: the ticket lands on staging → a session of 40 to 60 minutes → findings into the tracker → the stable checks written up as cases → the cases move into regression, while sessions go on looking for what regression does not have.
A session gives the most where cases do not exist yet or no longer help: a new feature without requirements, a large rework (cases confirm what was described, a session catches what was broken sideways), green runs alongside live complaints. And it does not fit where repeatability and provability are required: a regression run has to be identical from version to version (cases and regression), twenty near-identical tariff combinations are a job for a decision table, and a duty check repeated by someone else needs a checklist.
Where beginners stumble:
- They fall into the first rabbit hole. An oddity found in minute five gets investigated for forty, and the goal of the session is never touched.
- They do not repeat the finding immediately. Twenty minutes later the state cannot be restored, and the finding degrades into "I must have imagined it."
- They start with the exotica. Half an hour on emoji in the name field, while the ordinary form save has never been walked end to end.
- They go into unfamiliar territory unprepared. There is nothing to branch from: requirements first, or half an hour with the analyst.
In short
- The degree of formalization is about where the list of checks came from: written in advance (cases), born as you go (exploration), or absent entirely (ad hoc).
- The charter is written before the start: the goal, a limit of 40 to 90 minutes, what counts as a finding, and what you are not touching.
- The note is written before the click: time, data, action, result — four such lines assemble into a bug report in a minute.
- Notice an oddity and walk to it again from a clean state right away: the repeat separates a finding from an impression.
- Ideas come from stock moves: boundaries, repeat, abandoning halfway, other people's data, the back button and two tabs, a dropped network, time. The order runs from simple positive to complex negative.
- A session does not replace cases where repeatability and provability are required: regression, acceptance, duty checks, near-identical combinations.
- A session ends by going through the notes: findings, coverage (including where you did not get to), ideas for cases, questions for the requirements.
What to read next
- Test design techniques: equivalence classes and boundary values — where check ideas come from, the ones a session uses without any tables.
- Checklists and test data — the lighter frame that sits between cases and free exploration.
- How to write a bug report — how to turn a line of notes into a finding a developer can reproduce.
- Browser DevTools for a tester — the Network tab and throttling, what a session uses to look beneath the screen.