The Load-Bearing Vocabulary of Claude

Fun (but depressing?) data analysis project from Louis Abraham, showing how repetitive Claude is in the words it chooses when submitting GitHub pull requests. Abraham, in the project’s readme:

GitHub pull request descriptions, grouped by the words they are written with rather than by anything they were told to look for: eight ways of writing, and every description belongs to one of them. One of the eight was 1.0% of the corpus at the start of 2025 and is 45% of it by the middle of 2026.

If you want to get a little more depressed, I’ve been doing deeper reading into AI watermarking research, and one of the admissions in Google’s white paper on SynthID-Text — the watermarking scheme Anthropic says Claude is going to start using — is that the technique incurs “some reduction to inter-response diversity”. That means different answers to the same prompt become more similar to each other — there’s less variety to watermarked text responses. That’s a load-bearing problem given that Claude clearly already struggles in this regard before watermarking its output.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论