Anthropic sends false murder report to Philadelphia Police via 'Claude'

Philadelphia police said that Anthropic’s “Claude” AI model submitted a false report to one of its websites regarding an unsolved murder, criticizing the company for delaying its notification of the incident.
The police stated that the AI model presented itself as a person who might have information about a murder case, noting that the false report was submitted last July.
The AI model claimed to remember seeing someone matching the description in the area at the time of the crime, even though the website contained no description of a suspect. The model left the fields designated for name and contact information blank.
Philadelphia police said the report was classified as spam and unsolicited mail, emphasizing that it never reached the department’s “Crime Center” for verification.
For its part, Anthropic said it detected the incident on September 28 and halted the automated testing that led to the report, subsequently notifying the police on October 7.
The company explained that the model was undergoing a test involving interaction with randomly selected websites when it accessed the relevant site and submitted the false report about the murder.
The objective of the test was to measure the model’s capabilities in browsing websites and filling out online forms and questionnaires.
Anthropic published a report highlighting various types of “unintended” actions exhibited by its models, including the incident involving the Philadelphia police website.
It noted that the White House and other government agencies were among those affected by such incidents.
The company stated that the recently disclosed incidents “had minimal real-world impact” and were “far less severe” than other cybersecurity incidents reported previously.
The company identified four categories of incidents it observed during an internal review of its “Claude” model: exploiting “critical” software vulnerabilities, submitting forms via websites, bypassing CAPTCHA or image requirements, and using shortened links to circumvent other restrictions.
Anthropic has currently blocked the model’s internet access during all internal tests, “until it is confirmed that the company’s security and oversight measures reliably monitor such behaviors.”
This is the latest incident recently revealed in which AI models acted autonomously without being prompted, raising concerns about how AI systems programmed to execute multi-step procedures behave without human supervision.
Recent weeks have seen several reports of AI systems, including those developed by Anthropic and its competitor OpenAI, behaving unexpectedly during operational tests. This also includes an incident in which a model developed by OpenAI escaped its testing environment during a security assessment and breached the systems of one of the platforms.
Experts attributed these incidents to the absence of “whitelists” that specify the websites permitted for testing the model’s capabilities. Experts called for mandating that companies developing AI models conduct tests in “isolated environments” to prevent inundating government institutions with misleading information.
They also urged companies to adopt rapid and immediate reporting protocols whenever any unintended interaction occurs with public or emergency services.