All posts

4 min readĐọc bằng tiếng Việt

Why we run a fix three times before we call it fixed

One green run can be luck. VibeQA runs the repro test three times to verify bug fixes, then turns the outcomes into one verdict. Here is what each one means.

A developer moves a Jira ticket to Ready for QA. Somebody now has to open the app, follow the steps in the ticket and decide whether the bug is gone. If the bug only showed up some of the time, that somebody runs it once, sees it work, and closes the ticket. A week later it comes back.

VibeQA does that check for you. To verify bug fixes, it runs the same test three times.

What happens when a ticket moves

When a defect is tracked in VibeQA, it has a repro test: an approved test that fails while the bug is there. When its Jira ticket moves to the status you picked, VibeQA queues that test three times as one re-verify group. Each run gets a fresh sandbox: a new browser with no cookies and no signed-in state left over from the run before.

When the last run finishes, the three outcomes become one verdict, and VibeQA posts it as a comment on the ticket.

Re-verify, 3 runs
  1. #48211m 38s
  2. #48221m 41s
  3. #48231m 36s
Fixed
Three passing runs: the bug no longer reproduces.

Why one run is not enough to verify a bug fix

Many bugs that reach QA are intermittent: a race between two requests, a cache that hides the problem on the second load, test data that only breaks on some days. Say a bug still shows up in half of all runs. One run misses it half the time. Three runs all miss it one time in eight.

Three is not magic. A bug that shows up in only 30% of runs still slips past three runs about a third of the time. But it moves "we got lucky" from the most likely outcome to the unlikely one, for the price of two extra runs of a test you already have.

How the verdict is decided

The verdict is a small, deterministic function of the outcomes. No model reads the logs and gives an opinion; the same three outcomes always give the same verdict.

Outcomes of the runsVerdict
Every run that finished passedFixed
Every run that finished failedStill reproducing
Some passed and some failedFlaky
The test itself broke, or a run was cancelledBlocked
Fewer than two runs reached a pass or a failBlocked
Earlier runs: Release 2.14 check failed, 2 of 4 passed; re-verify groups for SHOP-452, SHOP-447 and SHOP-418; nightly runs
On the Runs page each re-verify group is one row with its verdict: flaky at 2 of 3, still reproducing at 0 of 3, fixed at 3 of 3.

Two details matter here.

Our failures are not your failures. If a run dies because a sandbox ran out of memory or the model provider timed out, that is an error on our side, not a failing test. The queue retries it, and it never counts as a pass or a fail. If two runs still reach a real result, the verdict stands on those two.

Flaky goes to a person. When the runs disagree, VibeQA does not pick a side. It says the result is flaky and leaves the call to you. A flaky verdict is useful on its own: it tells you the fix changed something, but not everything.

Re-verify, 3 runs
  1. #51071m 52s
  2. #51080m 47s
  3. #51091m 49s
Flaky
Runs that disagree: VibeQA reports flaky and leaves the call to a person.

One group at a time

A defect has at most one re-verify group running. If Jira sends the same update twice, or someone drags the ticket back and forth while the runs are going, the extra requests are skipped instead of piling up more runs. And once you close a defect in VibeQA, it stops re-verifying it.

Defect SHOP-418: three re-verify runs passed, verdict Fixed with 3 of 3 runs passed, Tokay saying 3 for 3. Fixed.
The defect page keeps the latest group and its verdict, next to Re-verify now and Close defect.

Running three times costs more than running once. We think it is the cheaper option once you count the reopened tickets.