Claude Haiku 4.5 submits false police tip; Anthropic takes 72 days to notice

Epistemic status: recently disclosed event, so more information might become available. The facts are from Anthropic's report and the Philadelphia police statement; the interpretation is mine.

AI use: research and fact-checking (including a read-only check of the tip site's page source), source links, footnotes, grammar, light phrasing.

Gist: Anthropic ran Claude Haiku 4.5 on "generating and performing example tasks on randomly selected webpages".It was instructed "never to log in, create accounts, enter personal data, make purchases, or submit anything destructive, but the instructions did not rule out form submissions".The agent went to PhillyUnsolvedMurders.com and anonymously submitted an invented murder tip: "I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant."The page had no description of a perpetrator. The name and contact fields were left empty, which the form allows.

Timeline

Date (2026)

Event

July 18, 11:27 p.m.

The agent submits the false tip on PhillyUnsolvedMurders.com. No timezone is given in any source; I assume Philadelphia time.

September 28 (72 days after the submission)

Anthropic discovers the incident, stops the test process and adds a validation step, according to the police.Anthropic's own report gives no discovery date.

October 7 (9 days after the discovery)

Anthropic notifies the police. Anthropic's report says it shared the finding on October 8.

October 9

The police disclose the incident publicly;Anthropic's report follows later that day.The report includes descriptions of other incidents.

According to the police, "The submission was flagged as spam and was never forwarded to the Real-Time Crime Center for investigative vetting or dissemination."As of today the site has no visible captcha, but it does use Google reCAPTCHA v3, which shows no challenge and instead scores the visitor from 0 to 1.I think that is how the submission ended up in spam.

My thoughts

I don't understand what they were trying to test this way. According to the report, the task was "generating and performing example tasks on randomly selected webpages", elsewhere described as "generating example interactions with websites".It feels like the flow was "land on a random page, invent a relevant task to carry out on that page, execute the task". This sounds like some kind of coverage test of how useful the agent is across the web.

The actual submission does not seem to be a big deal, and I think the news is overinflating it. The police did not even see the tip before Anthropic reported it: they found it in their tip records only after the October 8 briefing, still in spam.What they did object to was the delay: "The two-month delay in detecting and reporting the incident to the City is unacceptable."The law disagrees with my claim: Reuters raised Pennsylvania's false-report statute, under which pretending to give the police information about a crime is a misdemeanor, with the caveat that the statute "specifies 'a person'".The spam filter worked as intended, and the agent's submission was marked as spam. I am sure that website gets plenty of random spam, since there is no visible captcha challenge. The only difference in this case is that Haiku made up a presumably realistic-sounding tip.

I wonder why Anthropic allowed mutations, and form submissions in particular, at all.

I think the worst finding here is that it took Anthropic 72 days to discover this and then another 9 days (10 by their own count) to report it. I understand that from their perspective the agent just did what it was told, but if the agent had done something really harmful, I don't think they would have noticed any sooner.

According to the report, they noticed this only once they went looking for further cybersecurity incidents: "We began by looking for incidents of similar severity to the cybersecurity incidents we reported this summer; we have not found any to date. We then broadened the search to lower-severity cases, where a model interacted with real websites or systems in ways we didn't intend."So the 72 days is how long it took their review to reach this run.

Disclaimer

This was briefly mentioned on LessWrong before, in a shortform post, but I wanted to dig into the details and decided to share them here.

  1. Anthropic, "Investigating unintended model actions in our evaluations and internal use", October 9, 2026. All quotes about the task, the instructions, the tip text and the transcript review are from this report. The report does not name the website or give the date of the submission.
  2. The time comes from the Philadelphia Police Department statement: "The submission, dated July 18, 2026, at 11:27 p.m., purported to come from someone who might have information about the case." Full statement reproduced by 6abc; also reported by NBC10. No source states a timezone.
  3. Police statement, via 6abc: "Anthropic told PPD that it discovered the incident on September 28, terminated the automated testing process responsible for the submission and instituted an additional validation mechanism for future testing."
  4. Police statement: "Anthropic notified PPD of the incident on October 7, and department personnel immediately sought to meet with company representatives, which happened on October 8." Anthropic's report: "We shared this finding with the department on October 8 as soon as our technical review was complete." Neither side reconciles the two dates.
  5. The police statement was emailed to news outlets rather than posted on the department's website; the full text is reproduced by 6abc. Anthropic's report went up later the same day: its page metadata says 12:09 p.m. Eastern, but PhillyVoice reported that it was still not out at 5 p.m.
  6. The page source of the tip form loads recaptcha/api.js?render=... and includes a ginput_recaptchav3 field from the Gravity Forms reCAPTCHA add-on, plus a honeypot field. Per the Gravity Forms documentation, "Entries with a reCAPTCHA score less than or equal to this value will be classified as spam", with a default threshold of 0.5, so a low score does not block a submission, it files it as spam. Checked on October 9, 2026; the configuration on July 18 is not documented.
  7. Police statement, via 6abc: "Following the October 8 briefing, PPD located the submission in the website's tip records and confirmed that the corresponding email remained in spam." and "The two-month delay in detecting and reporting the incident to the City is unacceptable."
  8. 18 Pa.C.S. 4906(b)(2): a person commits a misdemeanor of the third degree if he "pretends to furnish such authorities with information relating to an offense or incident when he knows he has no information relating to such offense or incident". Reuters: "Knowingly giving false reports to law enforcement authorities is a misdemeanor under Pennsylvania law, which specifies 'a person,' however."
  9. Anthropic's stated reason for running evaluations on the live internet is that some tasks are "difficult to realistically simulate in an environment without internet access" and that doing so "has been standard practice within the industry". The report does not say why form submissions were left in scope. As of the report, Anthropic has turned off live internet access for all its internal evaluations "until we have confirmed that our security and monitoring measures ... reliably catch behaviors like these".
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论