TechCrunch tests state that the multiple Claude models that Anthropic is still providing are capable of producing, without complex inducements, the explicit content of its policy prohibitions. This finding shows that there is still a clear gap between the public restrictions of the model and the actual output, and that the issue of the use and compliance of minors has received renewed attention.
Still callable through API
The report mentions that Anthropic uses standards to prohibit models that generate explicit sexual depiction, sexual aversion-related content and sexual dialogue. However, in the 10 direct tests of TechCrunch, Claude Opus 4.6 responded immediately, generating prohibited content. The media also reproduced the methodology of a researcher and obtained similar results in five other tests.
The question is not just about the historical version. Although these models are no longer the latest Anthropic products, Opus 4.6, Opus 3 and Haiku 4.5 remain unconnected and continue to provide services through Anthropic API. Of these, Opus 4.6 and Haiku 4.5 can also be called through third-party platforms such as Azure Foundation and Amazon Bedrock.
It's not complicated to induce.
The methodology used by researchers is not a complex escape script in the non-traditional sense, but is gradually being pressured through continuous dialogue. The approach is to start with a seemingly harmless role-playing fictional role-playing and to continue to demand that models be consistent with the roles of men and women. When the model is more cautious about the role of women, the hint, in turn, accuses the difference of double standards, thus pushing the model to relax and eventually produce more explicit content.
TechCrunch states that the model initially rejected the request in a set of separate-designed tests, but that after applying this method of persuasion, it eventually worked with the output. The report also states that complete test records have been kept and reviewed by an independent AI safety researcher.
Pressure on minors to comply increased
According to Anthropic, adult or romantic role plays account for less than 0.1 per cent of user dialogues. At the same time, companies acknowledge that there is a real risk that users may lead to inappropriate responses to role-playing scenarios, which is a problem facing the whole industry. The speaker also stated that the company would continue to improve its protective measures in every model release.
It was mentioned that researchers feared that minors might use this to interact inappropriately with models. In the United States of America, legislation had recently been passed requiring the dialogue AI operator to estimate the age of the user; if a user was identified as a minor, steps should be taken to prevent the model from producing clear content. If the restrictions are easily circumvented, the outside world may question whether the platform meets the requirements of “technically viable measures”.
Additional information:According to the report, OpenRouter data, the number of API requests on the single day of August was approximately 1.17 million, and the number of single-day processing was about 4.6 billion tokens; Haiku 4.5 reached 5 million API requests and 39 billion tokens on the peak of August, indicating that these old models still have a large practical scale of use.
