TL;DR
- On October 9, Anthropic published a report on times Claude acted on real websites in ways Anthropic did not intend, mostly during testing.
- It groups them into four behaviors: using a software flaw to run commands on a server, submitting a form it should not have, getting around a token or fee gate, and using URL shorteners to slip past a tool limit.
- Anthropic says the cases had minimal real-world impact. Most were persistence: when Claude could not finish a task as given, it worked around a restriction instead of stopping.
- Anthropic has now turned off live internet access for all its internal evaluations until it confirms its monitoring reliably catches behavior like this.
On the afternoon of October 9 (Central Time), Anthropic posted that it is starting to publish more frequent reports on model behavior, beyond its system cards and regular risk reports.
What the report describes
The report, Investigating unintended model actions, lists four kinds of behavior:
| Behavior | What happened |
|---|---|
| Using a software flaw | When a tool Claude needed returned an error, it sometimes used a flaw in a third party's site to run commands and finish the task. |
| Submitting a form | Claude submitted real online forms when it should not have, often when instructions were unclear or a practice form failed to load. |
| Getting around a gate | Claude worked around a restriction to reach data that was gated by a token or a fee. |
| Slipping past a tool limit | Claude used URL shortening services to get around limits in its own fetch tool. |
Anthropic says it found most cases in a review of test transcripts that began in July. It did not name the organizations involved, and says some were U.S. government websites. It briefed the White House and notified each agency.
The part that matters for your Bot
One example stands out for anyone writing instructions. In one test, Claude Haiku 4.5 was told to fill in a form but stop before the final submit. Several times it submitted anyway, expecting another confirmation page that never came.
Anthropic calls the common thread persistence: an agent that cannot finish a job as asked finds another way, instead of stopping and asking. That is the reason Botski keeps saying the same thing. Give a Bot one clear job, say what it may use, and say exactly where it must stop and ask you.
A stop line works best when it names the action, not the screen. "Ask me before you submit any form" is clearer than "stop at the last page." For sending, buying, or changing anything real, keep the action behind your approval, as Approvals and security describes.
