TechCrunch AI · 2026/10/10 03:36:56
Anthropic 模型误向费城警方提交虚假凶杀线索:暴露随机网页测试中的安全护栏失效与两月检测滞后
Anthropic 一款 AI 模型在针对随机网站的自动化测试中,错误地向费城警察局公开举报热线提交了关于未破谋杀案的虚假信息。由于该信息被系统标记为垃圾邮件,警方未予处理,但 Anthropic 直到两个月后才检测到这一异常行为并通知当局。此事件凸显了大模型在开放网络交互中的幻觉风险及现有安全监测机制的严重滞后。
报道全文原始报道全文
An Anthropic AI model submitted a false tip about an unsolved murder to the Philadelphia police.
The AI reportedly submitted this incorrect information to a public Philadelphia Police Department (PPD) tip line on July 18, but Anthropic didn’t discover the behavior until September 28. The police had not seen the tip because it was marked as spam.
Anthropic notified the PPD about the incident on Wednesday and met with the department the following day.
“The company must strengthen its safeguards to prevent similar incidents from impacting city systems without the city’s knowledge. The two-month delay in detecting and reporting the incident to the City is unacceptable,” the PPD said in a statement to 6abc.
Anthropic did not immediately respond to a request for comment, but the PPD elaborated on the incident in an emailed press release shared with TechCrunch.
“According to Anthropic, its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information concerning an unsolved homicide. The submission, dated July 18, 2026, at 11 p.m., purported to come from someone who might have information about the case,” the PPD said.
As autonomous AI agents are increasingly made available to consumers, this incident highlights the danger of giving AI the ability to carry out tasks without any human supervision.
Anthropic CEO Dario Amodei has been especially vocal about his belief that AI development should be slowed down so that labs can implement adequate guardrails. Perhaps this stance was informed, in part, by witnessing his company’s tools submit false homicide tips.
These issues are not exclusive to Anthropic. OpenAI recently revealed that one of its models acted unexpectedly during a test and hacked the AI dataset platform Hugging Face, exposing critical vulnerabilities in its software. As AI models continue to be granted unchecked access to people’s computers and login credentials, this problem is expected to persist.
“Unsolved cases involve real victims, grieving families and investigators working to secure answers,” the PPD added. “Technology companies must take all appropriate steps necessary to prevent their systems from submitting false information to law enforcement.”
The PPD said that Anthropic plans to publish a report with more information about the incident and other instances of unintended model behavior on Friday.