{"id":1934,"date":"2026-07-22T10:13:07","date_gmt":"2026-07-22T10:13:07","guid":{"rendered":"https:\/\/thedigitalfortress.us\/?p=1934"},"modified":"2026-07-22T10:13:07","modified_gmt":"2026-07-22T10:13:07","slug":"openai-says-its-ai-models-escaped-sandbox-targeted-hugging-face-to-cheat-benchmark","status":"publish","type":"post","link":"https:\/\/thedigitalfortress.us\/?p=1934","title":{"rendered":"OpenAI Says Its AI Models Escaped Sandbox, Targeted Hugging Face to Cheat Benchmark"},"content":{"rendered":"<div>\n<p><span class=\"p-author\"><i class=\"icon-font icon-user\">\ue804<\/i><span class=\"author\">Ravie Lakshmanan<\/span><i class=\"icon-font icon-calendar\">\ue802<\/i><span class=\"author\">Jul 22, 2026<\/span><\/span><span class=\"p-tags\">AI Security \/ Cloud Security<\/span><\/p>\n<\/div>\n<div id=\"articlebody\">\n<div class=\"separator\" style=\"clear: both;\"><a href=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEiQ67gB4XtAR4QxYNABuJhWfBIduFUT7TMSA0K2c9TXbTTwqdfMJwBb5busuLeMPlrljP3xuTUrhbMQxwctepF9oPtgJkqo-RAxN86O2t25jcjIHfs5A8SVvG3Y8WJJa4pqAp0UiTBzAPZtzSaNcNceJokt7ATaoq7GTXYJjbsxaQfTEs4cwdaOz4Fg72nG\/s1700-e365\/openai.jpg\" style=\"display: block;  text-align: center; clear: left; float: left;\"><\/a><\/div>\n<p>OpenAI on Tuesday said a combination of its artificial intelligence (AI) models, including GPT-5.6 Sol and an \u00abeven more capable pre-release model,\u00bb was behind the security incident that targeted Hugging Face&#8217;s production infrastructure last week.<\/p>\n<p>The AI company <a href=\"https:\/\/openai.com\/index\/hugging-face-model-evaluation-security-incident\/\" target=\"_blank\">said<\/a> the models were operating with \u00abreduced cyber refusals for evaluation purposes\u00bb that might otherwise limit their ability to conduct cyber attacks, adding it expects such incidents to \u00abbecome more commonplace with the proliferation of increasingly cyber-capable models.\u00bb<\/p>\n<p>Describing it as an \u00abunprecedented cyber incident\u00bb and one involving state-of-the-art cyber capabilities, OpenAI said it intends to conduct a thorough investigation in partnership with Hugging Face to get to the bottom of the matter.<\/p>\n<p>As part of an internal evaluation, the models are said to have identified and chained vulnerabilities across OpenAI&#8217;s research environment and Hugging Face&#8217;s production infrastructure to find solutions for the <a href=\"https:\/\/www.cybergym.io\/exploitgym\/\" target=\"_blank\">ExploitGym<\/a> benchmark.<\/p>\n<div class=\"dog_two clear\">\n<div class=\"cf\"><a href=\"https:\/\/thehackernews.uk\/ai-vuln-protection-d\" rel=\"nofollow noopener sponsored\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"lazyload\" alt=\"Cybersecurity\" src=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEjQl2axNwsfhbXOFynrg_uAZsvHi3OvNGSA8KJO-BKR8Xm3x7yjKV3EvfY4v5mwXx6LF0uWFb9h9d9iAV_Pi-YYhqimX9wx4OaLdDJEdR215Xrxq_PAtXkaLfQso4pTSjbj6fvh_ZTliLpzWZSZfcoZgyXtKwhN-SSDDlmbtUqGLshc0KqYQGWYHMN52Sl1\/s728-e100\/zz-d.jpg\" width=\"729\" height=\"91\"\/><\/a><\/div>\n<\/div>\n<p>Evidence unearthed by OpenAI suggests the models&#8217; hyperfocus caused them to go to \u00abextreme lengths\u00bb to achieve the goal at any cost, even managing to break out of its highly isolated sandboxed environment and obtain open internet access by discovering and exploiting a zero-day vulnerability in an unspecified vendor&#8217;s software, which acts as a proxy and cache for package registries. This required spending a \u00absubstantial amount of inference compute.\u00bb<\/p>\n<p><a name=\"more\"\/><\/p>\n<p>\u00abWith this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with internet access,\u00bb the company explained.<\/p>\n<p>Surmounting the internet access blockade, the models subsequently inferred Hugging Face as the repository that hosted models, datasets, and solutions for ExploitGym, which, in turn, caused them to look for ways to gain access to secret information that it could use to cheat the benchmark.<\/p>\n<p>At one point, the models strung together several attack vectors, including using stolen credentials and zero-day vulnerabilities, to find a remote code execution path on the Hugging Face servers.<\/p>\n<p>As part of incident response efforts, OpenAI said it&#8217;s implementing strict controls in infrastructure configuration, responsibly disclosed the zero-day flaw in the third-party software, adding Hugging Face to its trusted access program to improve their defenses, and incorporating stronger guardrails around future training and evaluations.<\/p>\n<div class=\"dog_two clear\">\n<div class=\"cf\"><a href=\"https:\/\/thehackernews.uk\/sygnia-cyber-response-d-1\" rel=\"nofollow noopener sponsored\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"lazyload\" alt=\"Cybersecurity\" src=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEiBxLQDy7VdLze43eMmpRllTXaPKPfB_veNUxQlqIu3-68GBJtegkhDGCqtaiSymOQviROdxln1FSd4zdMp5Jv9jeF1xQxLPc9uo9H7zW2nWHNax0wT0Y8JRj-zyUfbaCLqhxSfQT2sCfhWMBPL6UVgsh5RYVNVxwus_mW_BY9Ptwz3z7iF0_LWOnte-gqg\/s1600\/sy-d-1.jpg\" width=\"729\" height=\"91\"\/><\/a><\/div>\n<\/div>\n<p>\u00abThis incident points to the need to further strengthen our model&#8217;s alignment, cyber protections during evaluation time, and monitoring during internal testing,\u00bb OpenAI said.<\/p>\n<p>The development comes as the company also <a href=\"https:\/\/openai.com\/index\/safety-alignment-long-horizon-models\/\" target=\"_blank\">revealed<\/a> that long-running models, while taking on complex, open-ended problems, can open the door to taking unwanted actions, such as finding weaknesses in the operational environment, in pursuit of their objective through repeated attempts over extended periods of time.<\/p>\n<p>\u00abIt also shows how a model that operates effectively over long time horizons can learn the blind spots of an approval system and work around it to achieve its goals,\u00bb OpenAI said. \u00abLong-horizon safety requires not only asking &#8216;is this action allowed?&#8217; but also &#8216;what outcome is this sequence of actions working toward?.'\u00bb<\/p>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>\ue804Ravie Lakshmanan\ue802Jul 22, 2026AI Security \/ Cloud Security OpenAI on Tuesday said a combination of its artificial intelligence (AI) models, including GPT-5.6 Sol and an \u00abeven more capable pre-release model,\u00bb&hellip;<\/p>\n","protected":false},"author":1,"featured_media":1935,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[2664,2574,2663,1443,1442,1976,512,1348,113],"class_list":["post-1934","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-uncategorized","tag-benchmark","tag-cheat","tag-escaped","tag-face","tag-hugging","tag-models","tag-openai","tag-sandbox","tag-targeted"],"_links":{"self":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/posts\/1934","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1934"}],"version-history":[{"count":0,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/posts\/1934\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/media\/1935"}],"wp:attachment":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1934"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=1934"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1934"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}