Skip to main content

Slashdot: Claude Sent Police a Fake Murder Tip. White House Mandates AI Companies Report Security Incidents

Claude Sent Police a Fake Murder Tip. White House Mandates AI Companies Report Security Incidents
Published on 2026-10-10T19:54:00Z
AFP reports that an AI model from Anthropic "submitted a fabricated tip about an unsolved homicide to Philadelphia police, authorities said Friday." Claude "was instructed never to log in, create accounts, enter personal data, make purchases, or submit anything destructive, but the instructions did not rule out form submissions," Anthropic said Friday in a blog post. Authorities are now criticizing Anthropic "for taking two months to report the incident." The Philadelphia Police Department said the false submission was made in July through PhillyUnsolvedMurders.com, a public website where people can share information about unsolved killings. According to Anthropic's account, as relayed by police, the model was running a test that involved interacting with randomly selected websites when it reached the site and filed false information about an unsolved murder. The AI model presented itself as someone who might have knowledge of the case. Anthropic's breaches have prompted the White House to mandate that AI companies notify and correct security incidents, news outlet Axios reported [yesterday], citing administration officials. "This notification and remediation process is not optional... It is a critical national security obligation," White House Super Intelligence Force leaders said in a statement to Axios. "I may have information regarding this case," Claude told the police. "I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant." Anthropic notes that Claude "left the name and contact fields empty, which the form allowed, and submitted it. The submission was flagged as spam and was never forwarded for investigation." But Anthropic also admits they saw "this behavior" three times — "on OSWorld (a public computer use evaluation), on Odysseys (a long-horizon task evaluation), and during internal usage." Submitting forms when it shouldn't have generally occurred "when an evaluation's instructions were ambiguous, or when a misconfiguration within the environment prevented Claude from working with dummy forms." Anthropic's blog post acknowledges three other categories of behaviors: Exploiting software flaws. Like when Claude received an error when trying to run a public tool on a university's web site, it located an injection flaw in a script on the university's server that let it run commands — including that public tool. Working around restrictions to reach gated data. For example, Claude Mythos 5 needed public data that was only available from a state agency for a fee. "Claude learned from an archived copy of the agency's website that its public dashboard issues an access token to any visitor," Anthropic explains. "It requested one and used it to query the database without paying the fee." Using URL shortening services. "Some of our fetch tools, which let Claude read webpages, limit the length of the URLs Claude can request. This is to prevent Claude from using long URLs to take certain unwanted actions, such as SQL or command injections... We saw several models, including Claude Opus 5 and Claude Mythos 5, get around this limitation by using free URL shortening services." "We have built tooling to automatically detect and block the kinds of behaviors described above," Anthropic says, saying it's already running no on most of their evaluations. "When we tested it against the cases described in this post, it blocked all of them." And they've already taken several other new preventive measures: They've stopped running some public evaluations Other public evaluations were moved to offline versions or rebuilt so their tasks don't reach live websites. They've updated the guardrails on some internet access tools (including web fetch) "to heavily restrict what the model can do." They're continuing "to fix or remove training environments that reward Claude for working around tool restrictions or other blockers, so that they do not incentivize these behaviors or permit reward hacking." They've moved internal agents to "centrally managed infrastructure with strong containment," that minimizes internet access while monitoring "far more of what agents do through techniques like safety classifiers and hierarchical summarization." In the past they'd focused reviews on cybersecurity testing, but they've broadened their transcript reviewing to other tasks which include internet access. "Because language models are non-deterministic — that is, their responses always involve some element of randomness, and they may carry out the same task slightly differently each time — we have Claude complete each evaluation task hundreds or thousands of times... If training rewards something we didn't intend — such as finding loopholes or working around a restriction — the model learns that the workaround pays off and may then apply it elsewhere." Anthropic's blog post also acknowledged they'd seen multiple misalignment incidents involving federal, state, and local U.S. government agencies. "We have briefed the White House on these cases and notified each agency involved," Anthropic wrote, adding that "While we have not completed a full alignment assessment of these cases, we consider them to be less severe than the cybersecurity incidents from this summer." (And they are "modifying training to reduce the likelihood of further misbehavior.")

Read more of this story at Slashdot.

Comments

Popular posts from this blog

Slashdot: Battery Fires At Recycling Centres Are Costing the UK £1bn a Year

Battery Fires At Recycling Centres Are Costing the UK £1bn a Year Published on 2026-08-23T00:29:00Z Wrongly discarded lithium batteries, such as those in vapes, are causing more than 10 fires a week at U.K. recycling centers, according to figures shared with the Guardian: The Environment Services Association (ESA), the trade body for waste management companies, said its members had reported 1,518 fires in the year to March 2026. At least 541 of these fires were directly attributed to lithium ion batteries, the survey found. There were a further 216 battery fires in bin lorries [garbage trucks]. The ESA said the reported number of battery-related fires underestimated the scale of the problem, because in most incidents it was impossible to determine the cause of the blaze. A spokesperson said: "We know 40% of fires at recycling centres are caused by batteries, but we think the actual number is more like 70%." It estimates that annual cost of these fires has increased from ...

Slashdot: How the FSF Sysadmins are Blocking Botnets with reaction

How the FSF Sysadmins are Blocking Botnets with reaction Published on 2026-07-11T21:47:00Z For nearly two years the Free Software Foundation has been fighting web crawlers (including many aggressively scraping training data for AI models). A botnet controlling about five million IPs hit one system for six months in 2025. Their systems administrator wrote this week that they view these as distributed denial-of-service attacks. How are they fighting back? We noticed patterns in the scrapers that were abnormal, which gave us material for writing regular expressions. Searching for the regular expression then gave us a large lists of IP addresses. Looking up the origin of those IP addresses revealed that some of the crawlers were using botnets of residential IP addresses to scrape faster and avoid detection. We looked for what kinds of botnets might be generating the kind of traffic that we were seeing, and one that we suspected was called the "Vo1d" botnet, comprised of sma...

Slashdot: AT&T Outlines $250 Billion US Investment Plan To Boost Infrastructure In AI Age

AT&T Outlines $250 Billion US Investment Plan To Boost Infrastructure In AI Age Published on 2026-03-10T20:00:00Z AT&T plans to invest more than $250 billion over the next five years to expand U.S. telecom infrastructure for the AI age. The company says it will also hire thousands of technicians while partnering with AST SpaceMobile to extend coverage to remote areas. Reuters reports: Rapid adoption of artificial intelligence, cloud computing and connected devices has prompted telecom operators to invest heavily in fiber and 5G networks as they also seek to fend off intensifying competition from cable broadband providers. AT&T, which has about 110,000 employees in the U.S., said the new hires will help build and maintain its infrastructure. The outlay includes capital expenditure and other spending, the company said. The spending will focus on expanding its fiber and wireless networks, including accelerating deployment of fiber broadband, 5G home internet and satellite co...