All posts

5 min readĐọc bằng tiếng Việt

What is different about VibeQA

What sets VibeQA apart from other AI testing tools: you review a test design, a person approves every version, and a red run means your app broke.

There are already several ways to get end-to-end tests for a web app. You can write Playwright by hand, record your clicks with a record-and-replay tool, or ask an AI testing tool to write the test code for you. Each of them ends with code that runs.

VibeQA ends there too. What it changes is what your team reviews, who says a suite is ready, and what a red result means. Here are those differences, each with the mechanism behind it.

You review a design, not a file

When Tokay, our agent, writes a suite, it starts with a test design, not code. The design goes from requirements to a classification tree (every input split into valid and invalid classes, with the boundaries marked) to Gherkin scenarios whose Examples rows are the actual test cases. For the phone-number field on a Storefront checkout, that is a short table: 10-digit mobile, +84 format, 9 digits, 11 digits, contains letters. A QA lead reads it in a minute and sees what is missing.

Coverage is computed from the design by a fixed rule: each class at least once, every boundary, at most one invalid class per row, and pairs of valid classes for high-risk requirements. Tokay does not declare coverage, and neither does the reviewer. The same design always gets the same answer. The details are in Design the tests first, write the code second.

Checkout coverage view: Every case is covered, 19 of 19 checks met; each phone-number class lists the cases that test it
Coverage is computed, not claimed: every requirement and every input class links to the cases that test it.

A person approves every version

Tokay drafts the design, answers comments with changes you accept or reject, and writes the code: one Playwright test per automated row, named after it, like TC-2.3. It never approves its own work.

Approving a version takes a person with the owner or admin role, and the platform checks that the version is ready: the code was generated from that exact design, coverage is complete, no comment thread is open, every case was reviewed by a person in its current form, and a dry run finished and reported every automated row. A failing row is allowed, because a repro test is supposed to fail.

Pass and fail come from the runner

A run passes or fails on two things the test runner leaves behind: its exit code and its report.json. Not on the log stream, and not on a model reading the output and giving an opinion. Tests can include AI-driven steps, but the verdict is still the runner's.

Six outcomes, not two

A red run should mean your app broke. So a finished run has one of six outcomes:

PassedFailedSuite errorInfra errorModel errorCancelled
Infra and model errors are ours. The queue retries them, and they never count as a test failure.

A sandbox that runs out of memory or a model provider that times out is an error on our side. Only those two are retried, and only by the queue. A real failure is never re-run until it turns green. More in A red run should mean your app is broken.

Suites table, last results: Sign in and Search Passed, Cart Our side (violet), Checkout Failed, Order tracking Not run yet
In the Suites table, Cart's last run is marked Our side, not Failed; only Checkout's red is about the app.

A fresh sandbox for every run

Every run starts in a new sandbox with its own browser and its own network, and it is thrown away afterwards. The sandbox holds no platform secrets; the only credential inside is a token that expires with the run. Your project's test accounts reach only that project's sandboxes, and Tokay sees their names and usernames, never the passwords.

Jira fixes, checked three times

When a defect is tracked in VibeQA and its Jira ticket, say SHOP-388, moves to Ready for QA, VibeQA runs its repro test three times and posts one verdict back to the ticket: fixed, still reproducing, flaky or blocked. Runs that disagree are reported as flaky and left to a person. Why three, and what three cannot catch, is in Why we run a fix three times.

Where your team already works

Tokay reads specs from Jira, Confluence and Notion, and results go out to Slack and Telegram. A connection belongs to the whole organization or to one project, which overrides the organization's. Everything Tokay reads from a ticket, a page, an uploaded document or the app itself reaches the model labelled as untrusted data, not as instructions.

Your own server, if you need it

VibeQA can run on your own server with Docker Compose and your own Anthropic key, so the apps you test never leave your network.

Compared with hand-written tests and other AI testing tools

None of the usual options is wrong. They make different trade-offs.

Record and replay is the quickest way to capture a flow you already know. You get the path you clicked. The negative and boundary cases around it, and knowing which ones are missing, are still work you do by hand.

Hand-written Playwright gives you full control, and it lives in your repository. It is also all manual: someone decides the cases, writes them, keeps track of coverage and wires up where they run. VibeQA produces Playwright too; the difference is the reviewed design above it and the runs around it.

An AI that writes test code directly is fast. But to know what it tested, you read the code, and a few hundred green lines tell you little about the cases nobody thought of. If the same agent also reports whether the run passed, nothing independent stands behind the result.

What it does not do yet

VibeQA tests web apps today; Android and iOS come later. If your app signs in with a magic link or a social login, you export a signed-in session for tests to use, and replace it by hand when it expires. The Jira and Confluence connections use Atlassian API tokens that act as one person. And writing a design first is slower than asking for code. We think what you get to review is worth the wait.