5 min readĐọc bằng tiếng Việt
What automation testing is, and why end-to-end suites decay
Automation testing explained: what unit, API and end-to-end tests each check, what they buy you, and why browser suites so often stop being trusted.
Every release asks the same question: did this change break something that used to work? On a small app, a person can answer it by clicking through the main flows. On a real product, with a checkout, a dozen forms and three kinds of user, nobody clicks through all of it before every deploy. Something gets skipped, and it is usually the thing that broke.
Automation testing is the practice of writing that check down as code, so a machine can run it the same way every time.
Three kinds of automated test
Most teams end up with a mix of three kinds. They differ in how much of the system they touch.
| Kind | What it checks | Typical speed | What it misses |
|---|---|---|---|
| Unit | One function or class, in isolation | Milliseconds | How the pieces fit together |
| API | One service through its HTTP or RPC interface | Tens of milliseconds to seconds | What the user actually sees |
| End-to-end | A real flow in a real browser, against a running app | Seconds to minutes | Very little, which is why it is slow and fragile |
A unit test might check that a phone-number validator rejects 09123 as too short. An API test might
send that number to POST /orders and expect a 422. An end-to-end test opens the Storefront in a
browser, adds two products to the cart, types the short number into the delivery form and checks that
the page shows an error and does not let the customer pay.
The end-to-end test is the only one of the three that would notice if the error message rendered behind the Pay button on a small screen. That is its value, and also its cost.

What automation testing buys you
Repeatable regression checks. A test runs the same steps with the same data every time. It does not get tired on the fortieth form field or forget the step that only matters for cash on delivery. Once a bug is fixed, a test that reproduces it keeps it fixed, or at least tells you when it is back.
Fast feedback before release. A suite that runs on every pull request finds a broken checkout while the change is still fresh in the author's head, not two days later in a QA queue, and not from a customer.
A shared definition of done. A test is a precise statement of what the app should do. When it is written well, a new team member can read it and learn how the feature behaves.
Unit and API tests deliver these reliably. End-to-end suites are where things tend to go wrong.
Why end-to-end suites decay
Most teams that have run a browser suite for a year know the pattern. It starts with twenty tests and a lot of confidence. A year later it has three hundred tests, a nightly run that is red most mornings, and a habit of clicking "re-run" until it goes green. A few specific problems get it there.
Flaky tests
A flaky test passes and fails on the same code. The cause is usually timing: the test clicks before the button is enabled, reads a total before the discount request returns, or depends on data another test changed. Each flaky test is a small tax. A suite with enough of them teaches the team to ignore red, which defeats the point of having it.
Brittle selectors
End-to-end tests find elements by something: a CSS class, a position in the DOM, a piece of text. When a designer renames a class or moves the promo-code field into a collapsible panel, tests fail even though the feature still works. The fix is mechanical, but somebody has to do it, often for dozens of tests at once.
Maintenance cost
Every feature change touches the tests that cover it. For unit tests that cost is small and local. For end-to-end tests it is large, because one flow touches many pages. Teams under deadline pressure start skipping or deleting the tests that are hardest to update, and those are usually the ones covering the most complex flows.
Tests nobody can read
A 400-line test file with helper functions three levels deep tells you what the code does, not what the test is for. When it fails, the person on call has to reverse-engineer the intent before deciding whether the app is broken or the test is. When nobody can read a test, nobody can tell whether it is still worth keeping.
Green means nothing, red means less
An end-to-end run depends on a lot of machinery: a browser, a network, a test environment, seed data, sometimes a third-party sandbox. When any of it hiccups, the run fails, and the failure looks the same as a real bug. After enough false alarms, a red run stops meaning "the app is broken" and starts meaning "check whether it is the environment again". A green run is not much better if nobody knows what the suite covers.
Coverage nobody can state
Ask a team which cases their checkout suite covers. Does it try a phone number with nine digits? With eleven? An address at the length limit? Usually the honest answer is "let me look at the code". Code coverage tools measure which lines ran, not which inputs and boundaries were tried, so they do not answer the question either. If nobody can say what a suite covers, nobody can say what it missed.
None of this is a reason to stop
These problems are not arguments against end-to-end testing. They are the reason good suites take deliberate work: test design that someone can review, stable ways to find elements, isolated runs, and a clear line between "the app broke" and "the test environment broke".
In the next post we introduce VibeQA, which is our attempt to build those habits into the tool rather than leave them to discipline.