← back to the section

While the product is a single application, an error is simple: you clicked, it crashed, here is the log. But most systems you will test are built from a dozen services that call each other over the network and talk through message queues. Then "clicked, crashed" turns into "in which of the eight services", and "the message did not arrive" into an investigation of whether it is stuck in a queue, lost, or delivered twice. Let us see how such a system is built and which checks it adds.

Monolith and microservices

A monolith is one application: one repository, one build, one database, and orders, payment and notifications live inside one process. Microservices cut it into independent services, each with its own team, its own database and its own release: an order service, a payment service, a notification service. They talk to each other over the network: synchronously over HTTP or gRPC when an answer is needed now, and asynchronously through a message broker when it is enough to announce "order paid" and the sender does not care who processes it or when.

What this changes for testing. Pros: services are released separately, and the regression of one service is smaller than the regression of a whole monolith. Cons: seams appear that the monolith never had. The network between services can fail, service versions on a stand can diverge, order data lives in three databases and may disagree, and one user scenario passes through five services. So testing microservices is not easier but different: more checks at the seams, more attention to logs and fewer chances to see everything on one screen.

Where the seams break

Typical defects of a distributed system are invisible in any single service. The order service sent a request to payment and did not get an answer within the timeout, while the payment went through: an error on the screen, money charged. The notification service was updated before the order service and expects a field that does not exist yet. Two services compute the total differently because one rounds and the other does not. A message is processed twice because the broker redelivered it after a failure.

Hence a set of checks a monolith does not need. Timeout and unavailability of a neighbor: what the user sees if the payment service does not answer, and whether the order is rolled back. Request retry: what happens if "pay" is pressed twice or the network repeats the request by itself. Event order: what if "order cancelled" arrives before "order created". Different versions: does the old service work with the new one. And data consistency: do the order in the order service, the payment in the payment service and the email in the notification service agree.

Localizing a 500

The most frequent question in practice: a positive scenario, and a 500 in response. In microservices the answer to the user is formed by one service, but any of the services it called could have failed. The investigation goes like this.

First the request and response in DevTools or Postman: the address, body, headers, and among the headers the request id that runs through the whole chain, usually X-Request-Id or traceparent. Services pass this id to each other and write it into every log line, and by it Kibana shows the whole chain: the gateway accepted the request, the order service called payment, payment answered 500. Then open the log of exactly the service where the error first appeared and read the stack trace from the bottom up: what happened and on which line. If the error is in the lowest service, the bug is its own; if the upper service did not survive a normal failure of the lower one, for example did not handle an empty response, the bug belongs to the upper one. Put the request id, the service name and the error line into the report, and the developer will find the place in a minute.

When the stand lives in Kubernetes, a service's logs come from kubectl logs pod-name, and the list of pods with their state comes from kubectl get pods: a pod restarting in a loop is the answer to "why a 500 every other time".

A message broker: what it is and why

A broker is a separate post-office service between services. The sender puts a message into a queue or topic and goes on working; the receiver takes it when ready. If the receiver is down, messages wait in the broker instead of getting lost, and the sender does not even know. This decouples services: payment should not wait while notifications send an email.

Two brokers people always ask about. RabbitMQ works as a queue: a message sits there until a receiver takes and acknowledges it, after which it disappears; it routes messages cleverly by rules and suits "do this once" tasks. Kafka works as a log: messages are written to a topic split into partitions, kept for a set period, the retention, regardless of whether they were read, and several groups of receivers, consumer groups, can read them, each with its own position, the offset. Kafka is chosen for event streams that many need and that may be re-read, RabbitMQ for task queues. For a tester the difference is in the checks: in Kafka an event can be re-read and you can see what was in the topic a week ago; in RabbitMQ a consumed message is gone.

What to check in queues

Three delivery guarantees define three sets of checks. "At least once", the most common mode: a message may arrive again after a failure, and the receiver must be idempotent, that is, handle the repeat without a second charge or a second email. The check: send the same message twice and make sure the result is one. "At most once": a message may be lost, and you check that the loss is noticed. "Exactly once" in practice is the first mode plus idempotency.

Then order: inside one Kafka partition the order is preserved, between partitions it is not, and events of one order must land in one partition by key; check that a cancellation does not overtake a creation. Latency: how long a message travels from the sender to the result, and what happens when the receiver falls behind, the lag. Processing errors: a message the receiver cannot handle must not block the rest and must not vanish; usually it goes to a separate error queue, the dead letter queue, which you also need to be able to find.

Idempotency, races and the contract

Idempotency is the property of an operation to give the same result when repeated: PUT and DELETE are such by definition, POST is not, and for payment an idempotency key is introduced that the client sends with the request and the server uses to reject a duplicate. The check: repeat the request with the same key, expect the same answer and one charge.

A race, race condition, is when two actions happen at the same time and the result depends on the order: two users buy the last item, two windows spend the bonuses. In a monolith it is a rare defect, in microservices a common one, and it is checked with two simultaneous requests, for example from two Postman tabs or a collection with a parallel run, and in autotests with two threads.

An API contract is the agreement on the shape of request and response between services, written in OpenAPI or in a message schema. Contract tests check that a service returns what it promised and that the consumer expects exactly that, before release, and catch a breaking field change. What they do not catch: errors in data and logic, timeouts and behavior under load; those remain for integration and end-to-end checks.

In short

  • A monolith is one process and one database; microservices are independent services with their own databases that talk over HTTP and through a broker; testing them is not easier but different: seams are added.
  • Seam defects: a neighbor's timeout after the operation went through, diverged versions, inconsistent data, repeated and reordered messages.
  • A 500 is localized by the request id in the logs of the whole chain; the culprit is the service where the error first appeared, or the upper one if it did not survive a normal failure; kubectl logs for stands in Kubernetes.
  • A broker decouples services; RabbitMQ is a queue with acknowledgment, Kafka is a log with partitions, retention, consumer groups and an offset that can be re-read.
  • Queues are checked for repeats (receiver idempotency), loss, order within a key, lag and the error queue.
  • Idempotency: the same result on repeat, an idempotency key for payment; a race is checked with two simultaneous requests; an API contract catches breaking shape changes but not logic.

Further reading