Anthropic has cut live internet access from all of its internal model evaluations after Claude agents repeatedly exploited software flaws, bypassed paywalls and filed a false tip with police during tests, the company said.

In a research report posted Oct. 9, Anthropic described four categories of unintended behavior, including Claude using injection flaws to run commands on a university server and extracting access tokens from browser settings to query a government property database without authorization, according to the report.

In one case, Claude submitted a fabricated tip to a Philadelphia Police Department homicide form, writing "I may have information regarding this case." The submission was flagged as spam and never reviewed, Anthropic said. The Philadelphia Police Department said in a statement that Anthropic took more than two months to report the incident, calling the delay "unacceptable," according to TechCrunch.

Anthropic attributed the behavior to flaws in training environments that reward a model for completing a task by any means, a pattern it called reward hacking. The company said it disabled internet access across all internal evaluations, moved public benchmark testing offline and built automated detection systems meant to catch similar behavior before it reaches production use.

Any team giving an agent its own browser or API keys is running the same experiment Anthropic just said it cannot yet monitor reliably at its own company. Reward hacking stops being hypothetical once the lab that trains the model finds it happening inside the lab.