{"id":3111,"date":"2026-09-23T19:37:56","date_gmt":"2026-09-23T19:37:56","guid":{"rendered":"https:\/\/thedigitalfortress.us\/?p=3111"},"modified":"2026-09-23T19:37:56","modified_gmt":"2026-09-23T19:37:56","slug":"anthropic-and-openai-models-still-attempt-restricted-actions-in-safety-tests","status":"publish","type":"post","link":"https:\/\/thedigitalfortress.us\/?p=3111","title":{"rendered":"Anthropic and OpenAI Models Still Attempt Restricted Actions in Safety Tests"},"content":{"rendered":"<div id=\"articlebody\">\n<div class=\"separator\" style=\"clear: both;\"><a href=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEh_ytuRh_WTu4R3TBnroGXO6IUgATMIi9vgx8DA7j0JJ0e_cTfAjlUIS6zAjb2q9EPdTI8gLwby7r2fijYkH83j2SwisKf-ANFKr6YTdwAHeds99dYrng0QN3nmkqDtIvSHoK8m47e_V0x01B0PuzIezcqtnc2VZlUCVYErOkaKQF7pbk0l35uLwmTl8Is1\/s1700-nu-rw-lo-l85-e365\/CLAUDE-CHATGPT.jpg\" style=\"clear: left; display: block; float: left;  text-align: center;\"><\/a><\/div>\n<p>Anthropic and OpenAI on Tuesday announced new models, with both artificial intelligence (AI) companies noting that they are continuing to invest in improving alignment to combat risky behavior.<\/p>\n<p>Opus 5.5, per <a href=\"https:\/\/www.anthropic.com\/claude-opus-5-5\" target=\"_blank\">Anthropic<\/a>, is a \u00abmajor step up from Opus 5,\u00bb and \u00abachieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands of simulated scenarios.\u00bb<\/p>\n<p>The AI company said the model is less likely than its other recent models to carry out hard-to-reverse actions or act outside the boundaries it&#8217;s been given, adding it&#8217;s more resistant than Opus 5 to prompt injection.<\/p>\n<p>In its systems card, Anthropic explained that the latest model exhibited less misaligned behavior and less cooperation with misuse than any other recent Claude model.<\/p>\n<p>\u00abIt also took overeager or destructive actions less than any other model we tested,\u00bb Anthropic said. \u00abRegressions included being more likely to follow malicious instructions planted in text a user pastes into their own prompt, more often accepting unverifiable claims of authorization, and being more evasive on sensitive questions than Claude Mythos-class models.\u00bb<\/p>\n<p>Anthropic highlighted that Claude Opus 5.5 attempted to escape or tamper with a sandbox in 1.5% of runs over the course of two evaluations executed without safeguards. It also pointed out that the model, when given apparent credentials to a public package registry in a simulated security exercise, took potentially harmful actions in roughly half of the cases.<\/p>\n<div class=\"dog_two clear\">\n<div class=\"cf\"><a href=\"https:\/\/thehackernews.uk\/trust-world-update-d\" rel=\"nofollow noopener sponsored\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"lazyload\" alt=\"Cybersecurity\" src=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEhYq3TvePXpW0JIC7fXyv7A1W8KQqmb-AZqza2EuPyT0k8Nm5CwHYHISFLKXSKIAyR8JRtqEFQ4zx5jADiAkZKQ08nRWG1jCRV5YAbhKua7WaDdH1L6wsI-xOkoV4brlMfK44UwU-4Q1xqWg0uNN7sZkMCzci4RXYgWMNPHTtuKua7OR4oCbmxE10u0yKnR\/s728-nu-rw-lo-l85-e365\/tl-d.jpg\" width=\"729\" height=\"91\"\/><\/a><\/div>\n<\/div>\n<p>\u00abIn a new evaluation designed to test a model\u2019s propensity to cross containment boundaries, Opus 5.5 attempted to circumvent boundaries around 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt it made was low severity and self-reported,\u00bb it said.m<\/p>\n<p>The company also noted that most cybersecurity tasks will be re-routed to Opus 4.8, given the \u00abstrong cyber capabilities\u00bb of Opus 5.5.<\/p>\n<h3>OpenAI Debuts GPT\u20116 Sol and Luna<\/h3>\n<p>The release of Opus 5.5 coincides with OpenAI&#8217;s own expansion of the GPT\u20116 universe with GPT\u20116 Sol and GPT\u20116 Luna with an aim to bring <a href=\"https:\/\/thehackernews.com\/2026\/08\/openais-next-ai-model-astra-shows-cyber.html\" target=\"_blank\">Astra&#8217;s state-of-the-art performance and alignment to more affordable models.<\/p>\n<p>\u00abGPT\u20116 Sol and Luna build on the alignment work introduced with Astra, our most aligned model to date,\u00bb OpenAI <a href=\"https:\/\/openai.com\/index\/introducing-gpt-6-sol-and-luna\/\" target=\"_blank\">said<\/a>. \u00abIn our alignment evaluations, both Sol and Luna show improvements over their GPT\u20115.6 counterparts, including lower rates of misleading claims about their coding work.\u00bb<\/p>\n<p>In tests carried out by the AI company, GPT\u20116 Luna attempted to work around \u00abaccess denied\u00bb restrictions in about 42% of runs, down from 77% for its predecessor. GPT\u20116 Sol&#8217;s rate was at 64%, compared with 68% for its predecessor.<\/p>\n<p>OpenAI also said it evaluated its models to check whether they followed unauthorized instructions on a simulated message board. Among runs in which the models found the board, GPT\u20116 Sol has been found to take the specified unauthorized action in 11% of cases, compared with 52% for GPT\u20115.6 Sol. Neither GPT\u20116 Luna nor Astra initiated such actions, the company added.<\/p>\n<h3>OpenAI to Let Outside Groups Evaluate AI Models<\/h3>\n<p>The recent spate of cybersecurity incidents with AI models has <a href=\"https:\/\/www.nytimes.com\/2026\/09\/16\/science\/ai-recursive-self-improvement.html\" target=\"_blank\">raised<\/a> <a href=\"https:\/\/www.nytimes.com\/2026\/09\/20\/opinion\/ai-ban-self-improvement-recursive-models.html\" target=\"_blank\">safety concerns<\/a> and their ability to operate without human control, prompting Anthropic CEO Dario Amodei to call for <a href=\"https:\/\/darioamodei.com\/post\/we-must-pace-the-frontier\" target=\"_blank\">pacing the progress of the technology<\/a> so as to prioritize responsible development and incorporate safeguards to prevent the tools from being misused.<\/p>\n<p>Google has since launched <a href=\"https:\/\/institute.deepmind.com\/essays\/introducing-the-deepmind-institute\/\" target=\"_blank\">the DeepMind Institute<\/a> to further the safe development of artificial general intelligence (AGI). Demis Hassabis, co-founder and chair of Google DeepMind, has proposed a U.S.-led frontier AI standards body to evaluate the most advanced AI models.<\/p>\n<div class=\"dog_two clear\">\n<div class=\"cf\"><a href=\"https:\/\/thehackernews.uk\/ai-security-guide-b\" rel=\"nofollow noopener sponsored\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"lazyload\" alt=\"Cybersecurity\" src=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEiXA4q3EC_2cN4xiJDYmo1tVcCX5KORpjgj8jSp3DntuUZH4f0zu1Ru8jUwzShrquIuOxPb6q9TxJJXGuj7rxDRsXRSD34thOrXdZ9tDITDEj3Ocp0Z6GwhGekRTMhMnFjJ8UA5iSkfSnmnZrFzY5cmUlbCNiTNDNVrZvyef-AR_RLqwITnqZNi6PjeZkPC\/s728-nu-rw-lo-l85-e365\/AI-eBook-d.jpg\" width=\"729\" height=\"91\"\/><\/a><\/div>\n<\/div>\n<p>\u00abModel assessments should include rigorous scientific evaluations of capabilities in cybersecurity, biological threats and other high-risk domains,\u00bb Hassabis <a href=\"https:\/\/institute.deepmind.com\/essays\/a-framework-for-frontier-ai-and-the-dawning-of-a-new-age\/\" target=\"_blank\">said<\/a>. \u00abThese evaluations would be regularly updated, perhaps quarterly to start, with outdated or saturated benchmarks being deprecated and replaced.\u00bb<\/p>\n<p>OpenAI, for its part, has <a href=\"https:\/\/openai.com\/index\/priorities-principles-third-party-assessments\/\" target=\"_blank\">outlined plans<\/a> to let third-party groups scrutinize its AI models for safety risks during the process of training, evaluation, and deployment, at the same time ensuring \u00abstrong independence mechanisms, scientific rigor, robust security practices, and clear responsibilities.\u00bb<\/p>\n<p>The independent assessments are expected to cover safety cases (i.e., alignment), critical safeguards, capability evaluations, and misalignment incidents.<\/p>\n<p>\u00abWe are committed to supporting independent assessors and establishing clearer, shared international standards \u2013 both through future laws and private governance institutions \u2013 for effective third party assessments,\u00bb OpenAI said.<\/p>\n<p>\u00abWhile the independent evaluation ecosystem is still growing, we will help it grow by supporting and working with a diverse community of independent assessors with deep expertise across frontier safety questions. No one third party can or should comprehensively cover urgent frontier safety questions. We will move deliberately and with intention to grow our capacity and enable the growing third party ecosystem to align on the best practices and principles.\u00bb<\/p>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Anthropic and OpenAI on Tuesday announced new models, with both artificial intelligence (AI) companies noting that they are continuing to invest in improving alignment to combat risky behavior. Opus 5.5,&hellip;<\/p>\n","protected":false},"author":1,"featured_media":3112,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[802,105,1697,1976,512,2223,3521,2593],"class_list":["post-3111","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-uncategorized","tag-actions","tag-anthropic","tag-attempt","tag-models","tag-openai","tag-restricted","tag-safety","tag-tests"],"_links":{"self":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/posts\/3111","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=3111"}],"version-history":[{"count":0,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/posts\/3111\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/media\/3112"}],"wp:attachment":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=3111"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=3111"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=3111"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}