Failure modes of Claude Code in a supervised but code-blind 60h project
100% human-written. Copyedited by Claude.
*Epistemic status: experience report of one ~60h project with Claude Code, plus longer for the write-up.
- Scope: I supervised Claude's thinking but didn't look at the code. Spec was an exploratory, high-level design refined iteratively.
- Method: I introduced runtime and design self-checks as Claude failure modes appeared, and ran manual regression tests against a known result. I noted the behaviors and quotes myself; I linked them to existing research and Anthropic's docs when I found correspondence; the rest I listed as potential research questions. The project arc is consistent with YC alum, AI tooling founder Dex Horthy's public account.
- Background/bias: 18+ years as a software engineer in critical infrastructure, lately moving towards R&D in formal methods. I am skeptical of human software engineering practices, so my bar for LLM code is just "about as good as human code". I tried taking the LLM-coding claims at face value. I discussed parts of the post with two AI-adjacent researchers and two senior software engineers, all broadly bullish on LLM usage, yet they recognized several of these failure modes from their own use.
Below is a summary focused on what I think is most relevant to LW. Full post at the link.
I guided Claude Code (Pro plan; Opus 4.8, Sonnet 5, Opus 5, Opus 5.5; at or above Anthropic's effort recommendations, whose inconsistencies I document) to build a semantic fuzzer to find bugs in Obsidian Sync. I supervised its thinking but never read the code. Handwaving a lot, we got a functional prototype in ~30h, which I estimate would take ~40h by hand.
The next ~30h I tried to finish off this prototype for publication, which turned into a bug treadmill: at every step, Claude kept breaking as much as it fixed. I gave up and declared the project finished at the ~60h mark.
The fuzzer works: It does find sequences of operations on synced notes that you can repeat manually to cause data loss on your own system. (I reported an example to the Obsidian devs but got no response.)
I rarely reached the Pro plan token limits, probably helped by the project's flexibility.
During the whole development process I took notes of how Claude Code failed: how its various kinds of confidently wrong explanations misled me; how it kept me busy enough to lower my standards; how Claude failed to provide (sane) design ideas and instead required mine, while leaving me no time to learn new ones; how all of this affected the overall productivity.
Additionally, at the end of the process I looked for correspondence with existing research. Turns out that even Anthropic's docs already cover some of the problems, which is hard to reconcile with their advertised use cases of "research assistant" and "thinking partner".
Some observations I couldn't match to existing research, so I collected them as potential research ideas. Some examples:
- There are empirical studies that hired teams to independently implement the same system to study variations in various factors affecting human maintenance effort. Would be interesting to compare them to LLM code, hopefully as a repeatable eval. What exactly makes LLM code difficult for humans? What makes LLMs get lost in their own code? If it boils down to context rot, would using a dense language like APL or even LISP help?
- Relatedly, looks like many people are looking at formal verification to improve LLM code, no matter if they look at the code or not. Obviously this makes more sense for those not looking; but how does it intersect with the previous point?
- I saw a case where the summarized CoT said "I won't do XYZ because " only to do XYZ seconds afterwards. Is it a more extreme case of CoT unfaithfulness? Or maybe relatable to other cases of the Claude CoT summarizer misinterpreting who is talking to whom and security infrastructure leaking into the conversation?
- Claude has a pattern of forgetting instructions (even repeated) while recalling useless trivia (even if told to purge it). Looks to me relatable to a recent paper on endogenous authorization laundering, but wider. Is there a way to measure Claude's "attachment" to each piece of information and how that correlates to survival in the needle-in-the-haystack experiment and persistent memory mechanisms?
Finally I carve out use cases that seem salvageable with enough discipline, and techniques to do so: basically short, well defined tasks that avoid context growth and rot (unlike long conversations), while explicitly keeping the human as the memory and control thread of the task.
The linkpost is the 2nd part of the experience report, which focuses on Claude Code's failure modes.
For more context on the project itself, please see the 1st part.
The code, with a human-written README explaining how to use it, is published on GitHub.
I'm looking to move into technical AI safety; I'll be grateful for feedback on the research ideas at the end of the post.