I Used Claude to Critique My Own Design, Then Built a Way to Pitch New Clients With It

I ran the same evaluation on my own design and a real product, then built an outreach method around what Claude found.

Quick note before you dive in: every tool or feature I mention here, I’ll tell you straight up whether it’s free or paid as of today. I’ve lost count of how many times I’ve watched a video, gotten excited about some AI result, tried it myself, and only then found out the person was on a paid plan the whole time. Not doing that to you.

I wanted to test something specific: can Claude actually run a Nielsen heuristic evaluation on a real screen? If you haven’t heard the term, Nielsen heuristics are ten general usability principles (things like how visible system status is, how consistent the design stays, whether the interface helps you recognize options instead of forcing you to remember them) that designers have used for decades to catch usability problems in a structured way, not just a gut-feeling review.

I ran it twice. Once on a medical guidance webpage I designed myself in Figma, and once on PostHog, a well-known product analytics tool. PostHog recently leaned hard into a retro, desktop-inspired redesign, so this was a chance to run an evaluation on something genuinely new, not an old, settled design.

Medical Guidance Webpage

Test 1: A Medical Guidance Webpage I Designed

I uploaded a screenshot and asked for a full Nielsen evaluation, no other context given, deliberately. Here are the 5 main findings:

  1. Visibility of system status: Weak. No indication of how many steps the symptom-to-solution flow takes.
  2. Match between system and the real world: Mixed. The hero copy is plain and human, but further down the page shifts into clinical shorthand like “APAP + Caffeine” and “bio equivalent active formulas,” terms the target user likely doesn’t know.
  3. Consistency and standards: Strong. The card-based layout repeats consistently across every section, features, symptom categories, product comparisons, article previews.
  4. Recognition rather than recall: Strong. Quick-tag chips under the search bar (Muscle Sprains, Tension Headache, Dry Cough) let someone recognize their issue instead of typing it out.
  5. Aesthetic and minimalist design: Weak. The homepage stacks a hero, a feature grid, a symptom grid, a full molecule comparison matrix, article previews, a feedback widget, and a newsletter banner, all before a first-time visitor fully understands the product.

It also flagged one thing outside the ten heuristics entirely: a stat in the hero reading “800+ Cured Patients.” This site gives health information, it doesn’t treat anyone. “Cured” is a clinical overclaim a compliance reviewer would likely catch too.

What it got right:

  • Both the consistency and recognition-pattern calls match what I’d actually catch myself as the designer
  • The “800+ Cured Patients” catch was sharp, that’s the kind of finding worth acting on

Where it honestly couldn’t say much:

  • Heuristics needing live interaction, error prevention, error recovery, control inside a flow, all came back “can’t assess,” because a static screenshot doesn’t show what happens when something goes wrong
  • It correctly avoided guessing at things it had no way to know, like actual client goals or user research, instead of inventing an answer to sound complete
PostHog Homepage

Test 2: PostHog

Different domain, different result shape. PostHog’s whole homepage leans into a playful, irreverent voice, a “Trash” link in the nav (a nod to the old desktop Recycle Bin), a fake urgency banner mocking marketing tactics, a joke “Not endorsed by Kim K” line, all part of the recent redesign. Here are the 5 main findings again:

  1. Visibility of system status: Strong. Work items in the product’s “Inbox” carry priority labels (P1, P2, P3) and timestamps, a specific, clear answer, not a fuzzy one.
  2. Match between system and the real world: Strong. The core metaphor, comparing the product to a self-driving car, “you’re still the driver,” stays consistent everywhere it’s used, plain language throughout, no unexplained jargon.
  3. User control and freedom: Strong. Built directly into the product’s core pitch: work happens in a safe testing space, nothing merges without a human clicking approve, stated plainly instead of left implied.
  4. Consistency and standards: Good. The desktop metaphor (Trash, folder-style language) and the hedgehog mascot repeat consistently across sections.
  5. Aesthetic and minimalist design: Weak. The page is packed, feature lists, pricing tables, social proof, docs links, joke sections, all stacked on one homepage.

It also flagged a content duplication: one full list of product features appeared twice, word for word, in the same section, the underlying code repeated it exactly.

What it got right:

  • The user-control finding is a strong, specific catch, most tools bury that kind of reassurance in fine print, this one built it into the pitch itself
  • The status-system finding (P1/P2/P3, timestamps) is concrete and checkable, not a generic “good use of status indicators” line

Where it was more uncertain:

  • The duplicated feature list could be a genuine repeated-content problem or just how the page is built, it said so instead of stating it as a confirmed bug
  • Where exactly personality (the jokes, the mascot) crosses into clutter on an already-dense page is subjective, and it flagged that instead of pretending there’s one right answer

What This Actually Tells You

The useful pattern here isn’t “Claude finds problems.” It’s that the gaps were consistent and predictable: anything needing live interaction, user research, or business context got flagged as uncertain instead of invented. That’s the part worth trusting. The parts confidently stated, consistency patterns, recognizable UI conventions, actual copy problems sitting right there in the content, held up well in both tests.

This is a fast first pass on any design, your own work, a competitor’s site, something you’re pitching to redesign. Upload a screen, ask for the heuristics by name, and you get a solid starting point in minutes instead of an hour of manual audit.

How to Turn This Into Client Outreach

This is a strategy, not something I’ve landed a client with yet myself, but it’s a real tactic freelancers use: find a problem on someone’s website, then pitch the fix. Here’s how the process works, step by step:

Pick a niche you already have real work in:
Healthcare, e-commerce, SaaS dashboards, whatever you already have 2 to 3 actual designs for in your portfolio. That way, if a prospect asks “can you show me something similar,” you already have an answer ready

Run Nielsen evaluation on Claude:
Give Claude the site’s link (or a screenshot) and ask it to run a Nielsen heuristic evaluation,the same way I did in this post.

Verify every finding yourself:
Drop anything unclear or anything Claude flagged as uncertain, only keep what you’d stand behind if someone pushed back on it

Reach out with something short, warm, and specific:
Don’t send a formal report. Busy founders skim. Something like:

Hi [name], I checked out [company]’s site and ran a quick usability review on it. A few things might be costing you sign-ups:
[issue 1]
[issue 2]
[issue 3]
I’ve designed a few [niche] products recently, happy to show you a quick redesign idea for one of these, no charge. Fixing even one of these could mean more conversions and fewer confused visitors.

Follow with an actual direction, not just problems:
A rough fix idea, a redesign concept, even a single reworked screen. Just telling someone what’s wrong with their site feels like criticism, easy to ignore. Telling them what’s wrong and showing them a rough idea of the fix feels like an offer, much more likely to get a reply

The part that actually gets you noticed isn’t Claude’s raw evaluation, it’s what you personally do with it afterward: writing a real message instead of a generic one, and showing up with an actual idea instead of just a list of problems.

Manual heuristic evaluation normally takes a while, going screen by screen, checking each principle by hand. This speeds that part up considerably. Worth running on a few sites you already know well first, so you get a feel for where it’s sharp and where it needs a second look, before you ever put it in front of a stranger.

One thing worth being direct about: verify every point yourself before you send it anywhere. Some of what Claude finds will be solid, checkable problems, worth sending as-is. Some of it will just be Claude honestly saying it isn’t sure. Only send the confirmed kind. Passing off an unsure guess as a confirmed problem makes you look sloppy, not helpful, use it as a first pass, not a final audit.

I’m Usama, a UI/UX designer based in Islamabad. I write about what design tools actually do, not what they claim to do. If you’ve tried something like this yourself, I’d like to hear how it went, drop it in the comments. More of these coming.


I Used Claude to Critique My Own Design, Then Built a Way to Pitch New Clients With It was originally published in Bootcamp on Medium, where people are continuing the conversation by highlighting and responding to this story.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论