3 min readĐọc bằng tiếng Việt
A red run should mean your app is broken
Test failure vs infrastructure failure: why VibeQA gives a run six outcomes instead of pass and fail, shows its own errors in violet, and retries only those.
Most test tools have two colours. A test passed, or something went wrong and the run is red. The trouble is that "something went wrong" covers very different things: your checkout button stopped working, your test clicked a button that was renamed, or the machine running the test ran out of memory.
Only the first one is news about your app; the last is an infrastructure failure, not a test failure. If all three show up as the same red, people learn that red does not mean much, and start re-running until it goes green. That is how real bugs get waved through.
Six outcomes
Every finished run in VibeQA has exactly one of these outcomes:
- Passed: every test passed.
- Failed: a test ran and an assertion did not hold. This is the one that is about your app.
- Suite error: the tests could not run as written, for example a broken config or a suite that filled its disk. Your test needs fixing, not your app.
- Infra error and Model error: the sandbox crashed, ran out of memory, timed out, or the model provider failed. That is on us.
- Cancelled: someone stopped the run.

Telling a test failure from an infrastructure failure
The outcome is decided from two things the test runner leaves behind: its exit code and its
report.json. Not from the log stream, and not from an AI reading the output and guessing.
A run with no report at all is an infrastructure error, because the runner never got far enough to write one. A run that was killed for using too much memory, or ran past its time limit, is an infrastructure error too. A run that exits with failing assertions in the report is a failure.
Because it is a plain function of those inputs, the same run always gets the same outcome, and we can test that function the way we test everything else.
What gets retried
Only the two outcomes on our side are retried, and only by the queue: an infrastructure or model error gets one more attempt in a fresh sandbox. If the retry passes or fails for real, that is the result you see.
A failed run is never retried. Re-running a real failure until it passes does not fix anything; it hides a flaky test or an intermittent bug. If a test really is unreliable, you can quarantine it for up to 30 days: it still runs and its history still shows when it recovers, but its failures alone no longer fail the run.
Why our errors are violet
In the dashboard, infrastructure and model errors are violet, never red. Red is reserved for a test that found something. Every status also has an icon and a word, so the meaning never depends on telling colours apart.

When you see red in VibeQA, go look at your app. When you see violet, the problem is ours, and the queue has usually tried again already.