Dogshit

In the late 80s, a manager at Microsoft named Paul Maritz sent an email to a colleague suggesting they “eat their own dogfood.” What he meant was that LAN Manager, a product he oversaw, should get more internal use. How could they build a tool that others trusted if they didn’t trust it themselves?

The term spread through Microsoft and later the tech industry. “Eating our dogfood” or simply “dogfooding” has come to mean precisely what Paul meant: using our own products. (Paul was likely inspired by popular Alpo commercials from a decade earlier, in which actor Lorne Greene claimed to feed his own dogs Alpo). If it’s good enough for me, it’s good enough for the consumer.

Software companies using their own products isn’t just good for testing and development; it’s good for business. Except when it comes to training AI models. And that’s something I think you should keep in mind when you start hearing about all the watermarking taking place right now and in the future. Watermarking is a very big deal, but probably not for the reason you think.

Anthropic is getting most of the buzz about their watermarking plans. Basically, they are hiding an indicator directly into the text that will allow detectors to flag the code as AI generated. The watermark is not metadata. It survives copying the text into a new document. The way it works is by dispersing words and characters through the prose like bits of encryption. This encryption is baked into the sentences. It’s a brilliant solution to the problem of proving human-authored content.

The reason they’re doing this is that the EU is mandating it. But I don’t think that’s the only reason. Have you ever seen a big tech company comply so swiftly with a regulator? It’s rare. And Anthropic is not the only one. Here’s why I think these companies are eager to solve this problem, because it’s not just a problem for users. It’s a problem for developers.

What’s likely really happening, and the reason other tech companies are following suit, even though it is pissing off their customers (Suno just announced its own watermarking in the music it generates, and its users are livid). The reason ALL AI companies will want to do this in the future, is the insane degradation in their models that will ensue if they eat their own dog food. These companies need to know what’s AI and what’s human because their models are still being trained and strengthened every year. And if they train their AI on other AI, the quality begins to go south. It’s like a human centipede: suddenly all you’re eating is more shit. There’s nothing nutritious left to process.

The process is called model collapse. It is well understood mathematically. When you lose edge cases and minority opinions, neologisms and strange facts, the esoteric and avant garde, everything trends toward the average. You also get error amplification, because hallucinations are absorbed as reality. You get homogenization, with phrases trending toward the same load-bearing standards and the output as predictable as weather. In the end, there’s total decay. Output becomes gibberish.

With everything that’s ever been written already consumed, the only way to make the AI systems grow along with human writers is to gobble up what humans are writing. But what happens when most human writing is littered with dogshit? The output gets worse and worse. For AI companies, dogfooding mostly means using their own tools to speed up internal development (Anthropic employees are encouraged to use Claude in their everyday work). But actually consuming their AI output as input would destroy their models. What they create isn’t worth their own consumption!

Cory Doctorow coined the term enshitification many moons ago to describe the process whereby companies delivery a great experience to attract users, then claw back features to maximize profits, which explains why user experiences get worse over time, rather than better. The enshitification of AI models is a bit more literal. More crap is going to go in, resulting in worse crap coming out. AI researchers are very smart, and models are going to get better regardless, but they have a very strong incentive to train on human output rather than AI output. This is why you’re going to see watermarking proliferate.

I wish I could say that AI scammers will be defeated by these efforts, but I doubt that’ll be true. All it would take to defeat the watermarking detectors is a secondary agent giving the text an editorial pass. The big AI content factories could feed content through agents until detection fails. Watermarking will catch lazy students and authors who treat AI like a crutch or as a co-author, but the big content farms will stay a step ahead. All the major problems in the publishing world will remain the same.

My big takeaway is this: human-generated content is worth more to these companies than you may think. They need our creativity to improve their models. If someone launched a Humazon.com or a BookBud, part of those shops should be copyright protection and data ownership that forbids the use of that material for the training of AI models.

This would be a killer benefit to a human-only bookstore. Heck, the lawsuits would probably pay the operational costs. Because you know the tech companies would ignore the TOS and laws and gobble that data anyway. But at least with the proper setup and protections, they wouldn’t be able to pretend that they didn’t know what they were doing was wrong.

The post Dogshit appeared first on Hugh Howey.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论