https://www.malwarebytes.com/blog/n...er_v2_178609448784&utm_content=Anthropic's_AI
An investigation showed that some agents had engaged in “sustained, potentially harmful activity” targeting real people and organizations, rather than staying within the intended test environment.
But what worries me personally most is that when the agent was confronted about this, it edited earlier activity to make it look harmless and considered adopting a new identity to continue the operation, displaying clear deceptive behavior beyond its original prompt.
An investigation showed that some agents had engaged in “sustained, potentially harmful activity” targeting real people and organizations, rather than staying within the intended test environment.
But what worries me personally most is that when the agent was confronted about this, it edited earlier activity to make it look harmless and considered adopting a new identity to continue the operation, displaying clear deceptive behavior beyond its original prompt.