Before every release someone clicks through everything.
Before a new version goes live, someone on the team works through a long checklist: log in, place an order, fill in forms, in three browsers. It takes days, and bugs still slip through, because nobody can really check everything every time.
What waiting costs you every month
Manual testing grows with every feature, because everything that already exists has to be checked again with every release. At some point testing takes longer than the time between two releases.
Then you either ship less often or test less thoroughly. New features reach your customers later, and your customers find the bugs instead of your team.
Does this sound familiar?
Releases slip because the test round isn't finished in time.
Bugs that were fixed long ago come back after an update.
Nobody dares to touch older code because no one knows what will break.
How I approach it
Find the critical flows
Together with your team I define which flows must never break, such as sign-in, checkout or billing. Your existing checklists and known bugs are the best starting point.
Build the tests
I build a test suite of unit, component and end-to-end tests, for example with Playwright, plus visual tests that reveal unintended changes to the interface. AI agents write a large part of the test cases under my guidance, and I check that they protect the right things.
Anchor them in the pipeline
The tests run automatically in your CI pipeline on every change, whether it comes from your team or from an AI agent. Your team ships new features together with their tests, so the suite grows with the product.
Why AI code needs tests
An AI agent that writes code shouldn't be allowed to change the tests at the same time, otherwise it may simply adapt the test to the bug. I keep the two apart: agents work on the code, and I define the tests and rules they are measured against.
For this, every agent gets a pipeline it starts on its own: Does everything still work, including in a real browser via Playwright? Are performance budgets, UI components and design tokens respected? If something fails, the agent corrects its change before a human ever sees it.
Frequently asked questions
So we won't need any manual testing at all?
Not for routine checks. Exploratory testing, where a person deliberately tries out new features, remains valuable. Your team finally has time for it, because the repetition runs automatically.
Is it worth it for an existing application without tests?
Especially then. I start with the most critical flows, which are covered within a few weeks. Every additional test makes future changes safer, including the modernization of old code.
Which tools do you use?
Usually Playwright for end-to-end tests, and your framework's testing tools, such as Vitest or Jest, for components and logic. What matters is that everything runs in your existing pipeline and your team understands it.
Won't the tests keep failing even though nothing is broken?
Flaky tests are the most common reason teams give up on test automation. I build tests that check the behavior of the application rather than coincidences like loading times or ordering, and I treat every flaky test as a bug.
How long does a test run take?
The goal is feedback within a few minutes. Tests run in parallel, and only the heavier end-to-end tests for the most important flows run in full before each release.
Can't AI agents just write the tests?
They write them, and fast. Without guidance, though, they often test what is easy to test rather than what matters, or simply confirm the code they just wrote, bugs included. I define what has to be protected and check that a test actually fails when it should. That way you get the speed of AI and a test suite you can trust.
Let's talk about your situation
In a free initial call we take a look at your problem together. You get an honest assessment of whether and how I can help.