Anthropic Unveils New Claude Models to Address ‘Harmful or Abusive’ Interactions
Anthropic has unveiled new features allowing some of its advanced models to terminate conversations, positioning this as a strategy for “rare, extreme instances of persistently harmful or abusive interactions.” Importantly, Anthropic underscores that this measure is intended to protect the AI model rather than the users.
To clarify, the company does not claim that its Claude AI models possess sentience or can be negatively impacted by user interactions. Anthropic has stated “high uncertainty regarding the potential moral status of Claude and other LLMs, both now and in the future.”
Nonetheless, their announcement brings attention to a newly launched initiative called “model welfare.” Anthropic indicates they are taking a precautionary stance, “seeking to identify and implement low-cost interventions to mitigate risks to model welfare, if such welfare is deemed plausible.”
This updated policy currently applies solely to Claude Opus 4 and 4.1, specifically targeting “extreme edge cases,” including “requests for sexual content involving minors and prompts that could incite widespread violence or acts of terrorism.”
Although these types of requests could pose legal or reputational risks for Anthropic (as evidenced by recent assessments of how ChatGPT can unintentionally support harmful beliefs), the company contends that in pre-deployment evaluations, Claude Opus 4 showed a “strong preference against” addressing such requests, exhibiting a “pattern of apparent distress” when such responses were generated.
Regarding the newly introduced conversation-ending capabilities, Anthropic clarifies that “Claude should deploy its conversation-ending function only as a last resort, after several attempts to redirect the conversation have failed and no constructive interaction appears feasible, or if a user expressly requests Claude to terminate the chat.”
Additionally, Anthropic specifies that Claude is “instructed not to exercise this ability in situations where users might be at imminent risk of self-harm or causing harm to others.”
Techcrunch event
San Francisco
|
October 27-29, 2025
If Claude decides to end a conversation, Anthropic assures users that they can still start new conversations from the same account and create new threads based on the preceding discussion by adjusting their responses.
“We view this feature as an ongoing experiment and will continue to refine our approach,” the company concludes.


