I Taught Claude to Test Our App by Letting It Watch Me Use It
How a business analyst with no QA background built a post-deployment test suite from a requirements doc and a browser window.
Every deployment used to end the same way for me. The developers would push the release, someone would post in the channel that it was live, and then I’d open the application and start clicking. Log in, create a record, edit it, move it through the workflow, check the reports, log out, and repeat for the next module. I’m a business analyst on a legal operations platform, and somewhere along the way, “making sure nothing broke” had quietly become part of my job. It was never formally assigned to me. I just knew the product best, so I was usually the first one to notice when something felt off.
The trouble is that manual checking doesn’t scale, and it doesn’t stay honest for long. By the fourth deployment in a month, you start skipping the screens you think are safe. You check the feature that changed and trust that everything around it still works. That’s exactly where regressions like to hide.
I’d always assumed the answer was test automation, and I’d always assumed test automation was someone else’s job. It needed engineers, frameworks, selectors, and a QA team. However, they go deep in requirements and I just wanted a quick check. What I learned over the last month is that most of what a good test suite needs, I already had. I just needed a way to hand it over.

Starting with the requirements doc
The first thing I gave Claude was the requirements document, the same one I’d written and kept updating throughout the build. It describes what each module is supposed to do, which fields exist, which validations apply, the states a record moves through, and who is allowed to do what.
That gave Claude the “what should happen.” It could read that a record needs certain fields filled before it can be submitted, or that a particular status unlocks a particular action. But a requirements doc has a blind spot I know very well, because I’m the one who writes them. It describes the product as it was designed, not as it’s actually used. It doesn’t tell you which screen people land on first, which filter everyone reaches for, or the shortcut that quietly replaced the official flow two sprints ago.
Still, starting with the requirements was the right call. It meant every test could trace back to something the product was supposed to do, rather than to whatever happened to be on screen that day. If a test failed, I could point to the requirement it was protecting. That’s the difference between a test suite that checks the app and one that checks the app against what we agreed to build.
Letting it watch me
The second step was the one I was most curious about. I opened the application in the browser and asked Claude to observe how I interacted with it. Then I simply used the app the way I normally would. I logged in, went through the modules, created and edited records, and walked through the same workflows I check after every release.
This turned out to be the most important part of the whole process. Watching me filled in exactly what the document couldn’t. It picked up the real order of steps, the screens I go to by default, the places where I wait for something to load before clicking, and the paths through the app that actually matter day to day. In a way, it was doing what a new QA hire does in their first week: sitting beside the person who knows the product and learning by watching.
There’s a name for this kind of knowledge: tacit knowledge. It’s the stuff you know how to do but would struggle to write down. Every product team runs on it, and it usually lives in one or two people who have been around long enough. When those people are busy or on leave, the checks get thinner. Letting the AI learn from watching was the first time I had a way to capture that knowledge without turning it into another long document nobody reads.
It also changed how I thought about my own knowledge. A lot of what makes me useful on this project isn’t written down anywhere. It lives in how I move through the product. Being able to show that knowledge, instead of having to document every last bit of it, made it transferable for the first time.
Reading the code behind the screens
Watching me showed Claude how the app is used, but not how it’s built. So I added a third source of context. I pointed Claude Code at the web repo and asked it to read through the frontend to understand the UI better. It didn’t change anything there. It just studied how the pages were put together.
This filled a different gap. From the code, it could see how screens and components were structured, how elements were named, which fields only appear under certain conditions, and how the pages connect to each other. Those are things you can’t fully pick up from clicking around, and they matter a lot when a test has to find the right button on a page every single time.
By this point, Claude had three views of the same product. The requirements doc said what the app should do, my walkthrough showed how it’s actually used, and the codebase explained how it’s built. Each one covered something the other two missed, and together they gave it a far more complete picture than any single source could.
From observation to documentation and test cases
Once it had the requirements, the workflows it had observed and an understanding of the code, I asked Claude to write the documentation and the test cases. The documentation described each flow in plain language, and the test cases turned those flows into Playwright tests across the modules.
The first draft was closer than I expected. The structure made sense, the flows matched what I’d shown it, and the coverage followed the modules laid out in the requirements doc. It wasn’t perfect, and I didn’t expect it to be. But it was a real starting point rather than a blank page, and for someone who had never written a test framework from scratch, that difference is everything.
My role at this stage looked a lot like my usual job. I read the test cases the way I’d read any spec, checking whether each one tested something meaningful, whether the expected results matched the requirements, and whether any module had been left thin. It turns out that reviewing a test case isn’t very different from reviewing a user story. You’re asking the same questions: what is this supposed to prove, and how will we know it worked?
Handing off to Claude Code for the fixes
When I ran the tests, a handful of small things broke. Some selectors didn’t match the actual elements on the page, a few steps needed to wait for content to load, and login credentials had to be handled properly across modules. None of these were problems with the logic of the tests. They were the fiddly details of making code run reliably against a real application.
For that part, I went back to Claude Code, which already knew the codebase, and the split felt natural. Claude in the browser was good at understanding the product: reading the requirements, watching me use the app, and describing what needed to be tested. Claude Code was good at the engineering detail: going through the test files, fixing selectors and timing, and getting everything to pass consistently. Using the right tool for each part made the whole thing faster than trying to force one tool to do everything.
This is something I’d tell anyone trying a similar setup. Don’t expect one tool or one prompt to take you from a requirements doc to a working test suite in a single step. Break the work into the parts that need product understanding and the parts that need engineering precision, and treat them as separate jobs. The first draft gets you most of the way there, and a focused round of fixes gets you the rest.
Two reports for two kinds of readers
The last piece was reporting. I created a markdown report of the test results, and a Playwright report was generated alongside it.
The two serve different people. The Playwright report is detailed, with step-by-step results and the information a developer needs to dig into exactly why something failed. The markdown report is the one anyone on the team can open and understand in a minute: what was tested, what passed, and what failed. As a BA, I spend a lot of my time translating between technical detail and people who just need to know whether things are okay. Having both reports meant I no longer had to do that translation by hand after every run.
It also changed what a deployment update looks like. Instead of a vague “looks fine from my side” message after a release, there’s now a report that shows what was checked and what the result was. That makes the conversation after a deployment more factual and a lot less dependent on whoever happened to be clicking around that day.
What I’d tell someone trying this
Keep your requirements doc current before you start. The AI treats it as the source of truth, so anything outdated in the doc will quietly turn into an outdated test. A few minutes spent cleaning it up saves a lot of confusion later.
When you walk through the app for the AI, cover the flows you’d be most worried about after a release, not just the happy path you use every day. Whatever you show it is what it learns, and whatever you skip is what it won’t test. It’s worth thinking about this walkthrough the same way you’d think about a demo for a new team member.
Finally, review the output like a BA, not like a spectator. The AI is fast, but it doesn’t know which flows matter most to your users or which edge cases have caused trouble before. That judgement is still yours, and it’s exactly where your product knowledge earns its keep.
What it looks like now
Now, after every deployment, I run the suite instead of clicking through the app myself.
It isn’t magic. The tests are only as good as the requirements and the workflows I showed it, so if I skip a flow while demonstrating, that flow won’t be covered. Selectors still break when the UI changes, and someone has to maintain the suite. But that maintenance is far lighter than doing the whole check by hand, and it’s honest in a way manual checking never was. It doesn’t skip the screens it thinks are safe.
The biggest shift for me wasn’t the time saved. It was realising I’d been thinking about automation the wrong way. I assumed it needed engineering skill first. In practice, it needed product understanding first: knowing what the app should do and how people actually use it. That was the part I already had, and the AI handled the rest.
I think this says something bigger about where roles like mine are heading. The line between the people who define a product and the people who verify it is getting thinner. A BA who understands the requirements and the users can now turn that understanding into something that runs, without waiting for a separate team to translate it. That doesn’t replace QA engineers, but it does mean the person closest to the requirements can take much more ownership of whether they’re actually being met.
If you’re a BA or PM who has quietly become the person who “just checks” every release, give this a try. Start with your requirements doc, then show the AI how you actually use the product. You might find you already know most of what a test suite needs.
I Taught Claude to Test Our App by Letting It Watch Me Use It was originally published in Bootcamp on Medium, where people are continuing the conversation by highlighting and responding to this story.