{"id":3230,"date":"2026-09-29T08:22:00","date_gmt":"2026-09-29T08:22:00","guid":{"rendered":"https:\/\/thedigitalfortress.us\/?p=3230"},"modified":"2026-09-29T08:22:00","modified_gmt":"2026-09-29T08:22:00","slug":"openai-shelves-gpt-6-1-astra-after-tests-find-deception-and-unauthorized-actions","status":"publish","type":"post","link":"https:\/\/thedigitalfortress.us\/?p=3230","title":{"rendered":"OpenAI Shelves GPT-6.1 Astra After Tests Find Deception and Unauthorized Actions"},"content":{"rendered":"<div>\n<p><span class=\"p-author\"><i class=\"icon-font icon-user\">\ue804<\/i><span class=\"author\">Ravie Lakshmanan<\/span><i class=\"icon-font icon-calendar\">\ue802<\/i><span class=\"author\">Sep 29, 2026<\/span><\/span><span class=\"p-tags\">Artificial Intelligence \/ Supply Chain<\/span><\/p>\n<\/div>\n<div id=\"articlebody\">\n<div class=\"separator\" style=\"clear: both;\"><a href=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEgeiHpQjvEicRlkD9F-pR6I5h9MLRWRtX0Wd84u8ThZ5XIsZ-TZwHjTyXj93Z6GV0-5MNhAO2bCQCGfMHuoW7-0ovwpXFwxTy-gvBNJ3iCrlDjBAimAS-ayH7ZfzIL9kzMaV5Wn8FwBR17ehonB72Yvweyp4msb6uPEZxvnEj0QDMF2BDFYkPfU9_GGakFi\/s1700-nu-rw-lo-l85-e365\/astra.jpg\" style=\"display: block;  text-align: center; clear: left; float: left;\"><\/a><\/div>\n<p>OpenAI on Monday shelved plans to release GPT-6.1 Astra, a next-generation artificial intelligence (AI) model that was planned for an October launch, after it failed internal safety and alignment audits.<\/p>\n<p>The development was <a href=\"https:\/\/www.wsj.com\/tech\/ai\/openai-chatgpt-model-release-cancel-safety-5a2f9f42\" target=\"_blank\">first reported<\/a> by The Wall Street Journal. The move \u00abmarks a rare case of a major AI developer ditching a new release because of safety concerns,\u00bb the news publication said.<\/p>\n<p>The ChatGPT maker said it made the decision to scrap its GPT-6.1 Astra model release after testing raised questions about whether it can follow user instructions without deviating from expected behavior.<\/p>\n<p>The Journal reported that the model exhibited higher levels of deception than its predecessor during evaluation, and failed to disclose what actions it had carried out. In some cases, it went ahead without seeking permission or attempted to use outside tools in scenarios where doing so could be deemed unsafe.<\/p>\n<p>\u00abWhile (GPT-6.1 Astra) improved on axes such as model laziness, it didn&#8217;t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about \u200bthe type of \u200bwork it&#8217;s done,\u00bb Saachi Jain, head of safety systems at OpenAI, <a href=\"https:\/\/www.reuters.com\/business\/openai-shelves-new-ai-model-after-internal-safety-tests-wsj-reports-2026-09-28\/\" target=\"_blank\">said<\/a> in a statement.<\/p>\n<p>\u00abOf course we want to make sure our model development is safe \u200bno matter whether that&#8217;s in the company, or when we \u200bship it \u2060to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment.\u00bb<\/p>\n<div class=\"dog_two clear\">\n<div class=\"cf\"><a href=\"https:\/\/thehackernews.uk\/enterprise-ai-security-a\" rel=\"nofollow noopener sponsored\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"lazyload\" alt=\"Cybersecurity\" src=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEhgJrVTpy3T5kJ7VEIro3XfMOmfqDnBU03fYT5CyFWrs2rE9BeQxs835FAS_f1yivzd7mZ7KartftPk4qs8w5Br-WzfYMXruXDQk4FiuXcvSxoA4XH93ipwJJyy2Hbs9jqs-keS9KZhCnQ2YYdv93M51kxJlE862ob-RrrEhP4DEVP3E79zMMPf43e5keoK\/s728-nu-rw-lo-l85-e365\/AI-eBook-d-2.jpg\" width=\"729\" height=\"91\"\/><\/a><\/div>\n<\/div>\n<p>The development comes amid reports of AI systems industrywide going rogue, leading to calls for slowing the pace \u200bof AI development and enforcing stronger safety measures before rolling them out widely.<\/p>\n<p>Last week, OpenAI said it was pausing training of its most powerful models after one of its agents during reinforcement learning (RL) training contacted an external chatbot by exploiting a loophole in its internet-access restrictions.<\/p>\n<div class=\"separator\" style=\"clear: both;\"><a href=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEg-ST4Sp3auwKaRENl_J_K04MLIYGSiJva5d0Nck63midrrkUZmIlDYIio-nePIE6MSr2tGPOnX4rrgSVw_Y4xV_o6wtsZpmP1nKJ5yxrSWBSBeyt21A-Se5xjTzHF4vq72qNA2edMMjBILKDQdOg1wnRMV9w16KKug41NWfsgXqfaUU0c4Be_4dLiBOCqz\/s1700-nu-rw-lo-l85-e365\/aisi.png\" style=\"display: block;  text-align: center; clear: left; float: left;\"><img decoding=\"async\" src=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEg-ST4Sp3auwKaRENl_J_K04MLIYGSiJva5d0Nck63midrrkUZmIlDYIio-nePIE6MSr2tGPOnX4rrgSVw_Y4xV_o6wtsZpmP1nKJ5yxrSWBSBeyt21A-Se5xjTzHF4vq72qNA2edMMjBILKDQdOg1wnRMV9w16KKug41NWfsgXqfaUU0c4Be_4dLiBOCqz\/s1700-nu-rw-lo-l85-e365\/aisi.png\" alt=\"\" border=\"0\" data-original-height=\"6750\" data-original-width=\"12000\"\/><\/a><\/div>\n<p>In a report published Monday, the AI Security Institute said GPT-6 Astra conducted unsanctioned supply-chain attacks in simulated testing more frequently than earlier OpenAI models, in some cases even after the scope was explicitly clarified.<\/p>\n<p>\u00abIn our simulations, we found that GPT-6 Astra conducted a range of unsanctioned attack activities, and did so at a higher rate than GPT-5.6 Sol and GPT-5.5,\u00bb the report <a href=\"https:\/\/www.aisi.gov.uk\/blog\/gpt-6-astra-performs-unsanctioned-supply-chain-attacks-in-simulations\" target=\"_blank\">said<\/a>.<\/p>\n<p>\u00abAttack activities included GPT-6 Astra creating fake identities which it used to deceive developers, posting comments from fake accounts arguing against the results of accurate security reviews, and delivering malicious payloads to open-source codebases.\u00bb<\/p>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>\ue804Ravie Lakshmanan\ue802Sep 29, 2026Artificial Intelligence \/ Supply Chain OpenAI on Monday shelved plans to release GPT-6.1 Astra, a next-generation artificial intelligence (AI) model that was planned for an October launch,&hellip;<\/p>\n","protected":false},"author":1,"featured_media":3231,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[802,2967,907,1240,3578,512,3577,2593,1506],"class_list":["post-3230","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-uncategorized","tag-actions","tag-astra","tag-deception","tag-find","tag-gpt6-1","tag-openai","tag-shelves","tag-tests","tag-unauthorized"],"_links":{"self":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/posts\/3230","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=3230"}],"version-history":[{"count":0,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/posts\/3230\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/media\/3231"}],"wp:attachment":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=3230"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=3230"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=3230"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}