Think Your Anonymous Account Is Safe? AI Might Figure Out Who You Are

Illustration of a face obscured by speech bubbles, with the right side of the face dissolving into digital pixels.

According to the famous joke, “On the internet, nobody knows you are a dog.” Now that may no longer be true—thanks to large language models.

In a recent study, researchers found that AI can identify social-media users who don’t reveal their real names simply by piecing together clues from their posts and matching them with details found in other accounts online.

One test involved users of tech-industry forum Hacker News who had posted their LinkedIn profile in their bio. The researchers stripped out any identifying details, and asked the AI to figure out whom each account belonged to using only public information. In the cases where AI took a guess—252 cases out of 338 total—it was correct 90% of the time. And the matching was done more quickly and cheaply, and on a larger scale, than a human investigator could.

The Wall Street Journal spoke with Florian Tramèr, an assistant professor of computer science at ETH Zurich and one of the co-authors of the paper, about how AI can now automate this kind of digital detective work, who should be worried and whether anonymous posting is dead.

Below are edited excerpts from the conversation.

WSJ: Who should be worried about this?

FLORIAN TRAMÈR: Anyone who writes under a pseudonym online, myself included. Someone might say, in a pseudonymous account, they studied at MIT in 1997 in one post, and mention having a dog named Bluey in another, and on their own those details feel harmless. But put them together and it’s clearly the same person.

WSJ: How soon before an average person, not a government or researcher, can realistically do this to someone else?

TRAMÈR: Governments have a clear advantage in data access that an individual using AI can’t match. But for what’s publicly available, the technology is essentially already there. Anyone can now use AI to do an inexpensive, thorough profiling of someone by gathering up posts they have made under their real name and other publicly available data and try to match it against an anonymous account.

WSJ: Was there a piece of information that came up again and again as the thing that gave people away?

TRAMÈR: It really depends on the platform. On Hacker News, it was often job history, people saying how long they worked somewhere or where they studied, which lines up easily with a LinkedIn profile.

In a different test we conducted on Reddit, we asked AI to analyze posts by anonymous or pseudonymous posters, and based on that information we asked the AI to match up those users with a different set of posts. We looked at people discussing movies, for instance, and sometimes just a shared taste in obscure films was enough to link two pseudonymous accounts.

WSJ: What if you keep your social-media accounts private? Since these systems can’t access private accounts or sites that block scraping, does that make private accounts meaningfully safer?

TRAMÈR: Getting around scraping limits is pretty easy, since most of those limits are designed to block AI bots and not actual people. So you can just scrape the data yourself and feed it to an AI agent yourself. Private accounts are different. If something isn’t accessible to the person doing the deanonymizing, there’s not much to fear. If my Facebook is private, the real risk is a friend of mine sharing their access with an agent, not a stranger’s bot getting in directly. We’re also seeing in follow-up work that people tend to underestimate how much of their information is actually public.

WSJ: Is the AI mostly matching facts, or is writing style doing real work too, the way researchers once analyzed writing style to identify anonymous authors?

TRAMÈR: That’s a common misconception. Writing style comes in mostly as a secondary check. Our main approach has a model summarize someone’s pseudonymous posts into a profile built from facts, then compares that profile against a pool of other accounts to find the closest match. At the final verification step, the model might notice stylistic patterns too, but it’s a supporting signal, not the main one.

WSJ: Is there anything people should change about how they post, in anticipation of this becoming more common?

TRAMÈR: Just being aware that pseudonymity has limits is a good start. If you have an account you don’t want linked to your real identity, share less, especially unusual or specific details. Some people have suggested defeating this by posting false information about yourself. That’s plausible, but the challenge is doing it convincingly.

If your profile says you studied computer science in one place and art history in another, that inconsistency stands out, and these models are good at catching contradictions. And you have to keep the story straight over time.

WSJ: Do you think pseudonymity and anonymity are dead on the internet, or dying?

TRAMÈR: I don’t think it’s dead. In most cases, it was always technically possible to tell two accounts belonged to the same person, and often nobody cares. But there are real risks, blackmail at scale, or exposing people who work somewhere sensitive. What’s changed is that this capability is now cheap and available to almost anyone, so stalking and similar misuse gets easier. But they’re not going to figure out an account is yours if there’s genuinely nothing connecting it to you.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论