OTHER

Anthropic Unveils New Claude Models to Address ‘Harmful or Abusive’ Interactions

Anthropic has launched new features that allow some of its advanced models to conclude conversations, describing this as a measure for “rare, extreme instances of persistently harmful or abusive interactions.” Importantly, Anthropic states that this approach aims to protect the AI model rather than the users.

To clarify, the company does not claim that its Claude AI models exhibit sentience or can be negatively impacted by user interactions. Anthropic has expressed “high uncertainty regarding the potential moral status of Claude and other LLMs, both now and in the future.”

Nonetheless, their announcement spotlights a newly implemented initiative they refer to as “model welfare.” Anthropic indicates they are taking a precautionary stance, “working to identify and implement low-cost interventions to mitigate risks to model welfare, should such welfare be deemed plausible.”

This updated policy currently applies solely to Claude Opus 4 and 4.1, specifically addressing “extreme edge cases,” including “requests for sexual content involving minors and inquiries that may incite widespread violence or acts of terrorism.”

While these types of requests could pose legal or reputational threats for Anthropic (as demonstrated by recent analyses of how ChatGPT can inadvertently promote harmful beliefs), the company asserts that during pre-deployment testing, Claude Opus 4 showed a “strong preference against” responding to such requests, indicating a “pattern of apparent distress” when such responses were generated.

Regarding the newly introduced conversation-ending capabilities, Anthropic explains that “Claude should utilize its conversation-ending ability only as a last resort, after multiple attempts to redirect the conversation have failed and no constructive interaction seems possible, or if a user explicitly requests Claude to end the chat.”

Furthermore, Anthropic specifies that Claude is “instructed not to use this ability in situations where users might be at imminent risk of self-harm or harming others.”

Techcrunch event

San Francisco
|
October 27-29, 2025

If Claude opts to end a conversation, Anthropic assures users they can still start new conversations from the same account and create new threads based on the prior conversation by altering their responses.

“We consider this feature an ongoing experiment and will persist in refining our approach,” the company concludes.