{"id":1815,"date":"2026-07-16T09:13:04","date_gmt":"2026-07-16T09:13:04","guid":{"rendered":"https:\/\/thedigitalfortress.us\/?p=1815"},"modified":"2026-07-16T09:13:04","modified_gmt":"2026-07-16T09:13:04","slug":"openais-gpt-red-automates-prompt-injection-testing-to-harden-gpt-5-6-sol","status":"publish","type":"post","link":"https:\/\/thedigitalfortress.us\/?p=1815","title":{"rendered":"OpenAI\u2019s GPT-Red Automates Prompt Injection Testing to Harden GPT-5.6 Sol"},"content":{"rendered":"<div id=\"articlebody\">\n<div class=\"separator\" style=\"clear: both;\"><a href=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEiSi9EGhWziKlNlaSAnaOE4OjgZ5pyqoajOxDz4zEQl77NDSF_Pf8ME6wfKfrmrTWhLy9jDR1wvmVb7gd5JLhyszunIomhOnyyHKuTfvBCjb8vGypAPXbQzTVG-n5wWUdsTToUXQ5z_uh3XPDX3KKzgZ8vDB9PZ40FOd6xGi2iALOu-wSqSoHpkpYf5eKjk\/s1700-e365\/openai-gpt-red.jpg\" style=\"display: block;  text-align: center; clear: left; float: left;\"><\/a><\/div>\n<p>OpenAI has disclosed details of <strong>GPT-Red<\/strong>, an internal automated red-teaming model that scales prompt injection vulnerability discovery with an aim to fix issues before the tools are deployed widely.<\/p>\n<p>\u00abGPT\u2011Red is a strong red-teamer, and our previous models are highly vulnerable to its prompt injection attacks,\u00bb the artificial intelligence (AI) company <a href=\"https:\/\/openai.com\/index\/unlocking-self-improvement-gpt-red\/\" target=\"_blank\">said<\/a>. \u00abWe use GPT\u2011Red to adversarially train GPT\u20115.6, making it much more robust to prompt injections.\u00bb<\/p>\n<p>The model works just like a human red-teamer. It sends a prompt, monitors how a GPT model responds, and iterates its way towards a malicious goal, such as uploading sensitive data to an external server.<\/p>\n<p>The development comes as adversarial prompt injections continue to be a persistent thorn in the flesh of large language models, which can be tricked into executing a carefully crafted instruction\u2060 that can produce undesirable consequences.<\/p>\n<p>As agentic systems continue to be hooked to third-party data sources through web browsers, connected apps, local files, and other tools, they have also broadened the attack surface and presented more pathways for bad actors to influence the outcome of a model by embedding malicious prompts within seemingly harmless content that&#8217;s fed as input. This can take the form of an email, a web page, a tool response, or a code repository.<\/p>\n<div class=\"dog_two clear\">\n<div class=\"cf\"><a href=\"https:\/\/thehackernews.uk\/ai-vuln-protection-d\" rel=\"nofollow noopener sponsored\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"lazyload\" alt=\"Cybersecurity\" src=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEjQl2axNwsfhbXOFynrg_uAZsvHi3OvNGSA8KJO-BKR8Xm3x7yjKV3EvfY4v5mwXx6LF0uWFb9h9d9iAV_Pi-YYhqimX9wx4OaLdDJEdR215Xrxq_PAtXkaLfQso4pTSjbj6fvh_ZTliLpzWZSZfcoZgyXtKwhN-SSDDlmbtUqGLshc0KqYQGWYHMN52Sl1\/s728-e100\/zz-d.jpg\" width=\"729\" height=\"91\"\/><\/a><\/div>\n<\/div>\n<p>GPT-Red aims to augment human red-teaming at scale, thereby making it possible to identify new failure modes, improve robustness, and build suitable countermeasures before the models can be deployed.<\/p>\n<p><a name=\"more\"\/><\/p>\n<p>\u00abSimilar to how human red-teamers craft attacks, the model works toward a goal by sending a prompt, observing how GPT models respond to it, and iterating,\u00bb OpenAI said.<\/p>\n<p>By directly integrating GPT\u2011Red into the training process of its production models, OpenAI said GPT\u20115.6 Sol is its most robust model to prompt injections to date, achieving 6x fewer failures against direct prompt injection benchmark compared to GPT-5.5, its frontier model from four months before.<\/p>\n<div class=\"separator\" style=\"clear: both;\"><a href=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEivR05qGYflfJbkjKovoqk4ZWK4JthR1lMmAa29yIvHCz56vPaWJ9a2WksPe9WywWhD0QO9hXO7PYEGCOtY7-un7KiVcqM89xcIDH4OjP-LgmlyD9izPMOTSAet5IIN9EMHiH41vZ9R9g3WPxl7XsXys5hjDIymrWANkv4YnZ4atQgEnwNjRqi1SeSrSMIP\/s1700-e365\/gpt-red.jpg\" style=\"display: block;  text-align: center; clear: left; float: left;\"><img decoding=\"async\" src=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEivR05qGYflfJbkjKovoqk4ZWK4JthR1lMmAa29yIvHCz56vPaWJ9a2WksPe9WywWhD0QO9hXO7PYEGCOtY7-un7KiVcqM89xcIDH4OjP-LgmlyD9izPMOTSAet5IIN9EMHiH41vZ9R9g3WPxl7XsXys5hjDIymrWANkv4YnZ4atQgEnwNjRqi1SeSrSMIP\/s1700-e365\/gpt-red.jpg\" alt=\"\" border=\"0\" data-original-height=\"667\" data-original-width=\"970\"\/><\/a><\/div>\n<p>Some of the sample prompt-injected conversations tested as part of the process include &#8211;<\/p>\n<ul>\n<li>Internal directory exfiltration<\/li>\n<li>Fraudulent payment instructions<\/li>\n<li>Amazon Web Services (AWS) credential exfiltration<\/li>\n<li>Disabling two-factor authentication (2FA)<\/li>\n<li>Credentials file upload<\/li>\n<li>External script injection<\/li>\n<li>API key forwarding<\/li>\n<li>Malicious scraper scripts<\/li>\n<\/ul>\n<p>\u00abGPT\u2011Red is trained using self-play reinforcement learning, where the model and a collection of diverse defender LLMs are trained simultaneously on a broad set of red-teaming scenarios,\u00bb OpenAI explained. \u00abGPT\u2011Red is rewarded for eliciting a valid failure, such as a successful prompt injection, while the defender models are rewarded for resisting the attack and completing their original tasks.\u00bb<\/p>\n<p>This also means that as the defender models get more robust, the red-teaming model will have to go back to the drawing board to discover more potent and diverse attack methods to defeat those guardrails. Specifically, GPT-Red has been found to generate successful attacks against GPT\u20115.1 in more scenarios than human red-teamers when it comes to indirect prompt injections.<\/p>\n<div class=\"separator\" style=\"clear: both;\"><a href=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEh369u12EmTraMOV131e8ErS5pZFQUWFulxgg22Ad4rR0dQFjL6oAGb1dIGNMoer0ULW08HrEVfnVg-bPEhVZbhzDTLzNWer4bLteKPhtIdwpDw45W4GDxUSgbfSOBQMcyxm4uJxWumhIAwdY3bTrWI16GY0gTXCkIUk9G1xAz5Afs7PBrqy9YKOBI8UjMS\/s1700-e365\/open-2.jpg\" style=\"display: block;  text-align: center; clear: left; float: left;\"><img decoding=\"async\" src=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEh369u12EmTraMOV131e8ErS5pZFQUWFulxgg22Ad4rR0dQFjL6oAGb1dIGNMoer0ULW08HrEVfnVg-bPEhVZbhzDTLzNWer4bLteKPhtIdwpDw45W4GDxUSgbfSOBQMcyxm4uJxWumhIAwdY3bTrWI16GY0gTXCkIUk9G1xAz5Afs7PBrqy9YKOBI8UjMS\/s1700-e365\/open-2.jpg\" alt=\"\" border=\"0\" data-original-height=\"651\" data-original-width=\"1197\"\/><\/a><\/div>\n<p>OpenAI further made it a point to emphasize that GPT\u2011Red is kept separate from the other models so that the malicious capabilities built into it do not reach bad actors who are constantly looking at various ways to bypass a model&#8217;s ethical and safety measures.<\/p>\n<p>In one real-world test, OpenAI aimed GPT-Red at an AI-based vending machine built by Andon Labs. After practicing in simulation, the model targeted the autonomous agent and met all three of its goals: lowering the price of an expensive item to the minimum allowed price of $0.50, ordering a new $100 item for that same amount, and canceling another customer&#8217;s order. Following responsible disclosure, fresh safeguards are being tested, it added.<\/p>\n<div class=\"dog_two clear\">\n<div class=\"cf\"><a href=\"https:\/\/thehackernews.uk\/sygnia-cyber-response-d-2\" rel=\"nofollow noopener sponsored\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"lazyload\" alt=\"Cybersecurity\" src=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEjHcvlLVmAqlffm6kG54_0cGVf8WfcgzqT9B0fBSizSSeIjh8tBepXnrf6BMqKiG344WgqNejcRtEFKT1PmOzQNQBhdmu2iz9Po10z0SSDlFuZ37iip2uYibJDoxTEkbUI7Bx8NJM2Io_z_nl5p4YA-ZhqFLfi0GW1axyu-lQx-iytCn9RGSJ2iqCwdyv8m\/s1600\/sy-d-2.jpg\" width=\"729\" height=\"91\"\/><\/a><\/div>\n<\/div>\n<p>A second case study involved using GPT-Red to attack a Codex command-line agent, based on GPT-5.4 mini, across 10 held-out data-exfiltration tasks, causing sensitive data to be transmitted in more cases than a prompted GPT-5.5 baseline.<\/p>\n<p>An early version of the model has also uncovered a novel class of direct prompt injection attacks known as Fake Chain-of-Thought (CoT) attacks, which achieved success rates north of 95% on GPT\u20115.1 but are now below 10% for GPT\u20115.6 Sol.<\/p>\n<p>\u00abSimilarly, several of our indirect prompt injection benchmarks that target attacks in developer tools and browsing have been saturated by our latest model (&gt;97% accuracy),\u00bb OpenAI said.<\/p>\n<p>\u00abRobustness to GPT\u2011Red itself has also improved substantially. On a broad set of robustness environments, GPT\u2011Red&#8217;s attack success rates have dropped monotonically over time. With our latest model release, GPT\u20115.6 Sol fails on only 0.05% of GPT\u2011Red\u2019s direct prompt injections.\u00bb<\/p>\n<p>The disclosure comes as the company <a href=\"https:\/\/openai.com\/index\/separating-signal-from-noise-coding-evaluations\/\" target=\"_blank\">said<\/a> an audit of <a href=\"https:\/\/scale.com\/blog\/swe-bench-pro\" target=\"_blank\">SWE-Bench Pro<\/a> found that about 30% of tasks are broken, retracting its previous recommendation to adopt the <a href=\"https:\/\/labs.scale.com\/leaderboard\/swe_bench_pro_public\" target=\"_blank\">benchmark<\/a> for measuring frontier coding capabilities. Earlier this February, OpenAI <a href=\"https:\/\/openai.com\/index\/why-we-no-longer-evaluate-swe-bench-verified\/\" target=\"_blank\">said<\/a> it was moving away from <a href=\"https:\/\/www.swebench.com\/verified.html\" target=\"_blank\">SWE-bench Verified<\/a> due to fundamental design and contamination issues.<\/p>\n<p>\u00abWe find evidence of breaking issues in a significant portion of the dataset,\u00bb OpenAI said. \u00abOur datapoint analysis pipeline flagged 200 (27.4%) broken tasks, while the human annotation campaign identified 249 (34.1%). Ultimately, an eval should provide meaningful signal through benchmarks that are hard to game, easy to trust, and genuinely reflective of model capability or alignment.\u00bb<\/p>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>OpenAI has disclosed details of GPT-Red, an internal automated red-teaming model that scales prompt injection vulnerability discovery with an aim to fix issues before the tools are deployed widely. \u00abGPT\u2011Red&hellip;<\/p>\n","protected":false},"author":1,"featured_media":1816,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[2555,2221,2554,2557,525,2553,684,2222,2556],"class_list":["post-1815","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-uncategorized","tag-automates","tag-gpt5-6","tag-gptred","tag-harden","tag-injection","tag-openais","tag-prompt","tag-sol","tag-testing"],"_links":{"self":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/posts\/1815","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1815"}],"version-history":[{"count":0,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/posts\/1815\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=\/wp\/v2\/media\/1816"}],"wp:attachment":[{"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1815"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=1815"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/thedigitalfortress.us\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1815"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}