OpenAI expanded the Daybreak program, splitting it into two tiers — Blue for defensive tasks and Red for offensive research. The new Red tier provides access to the GPT-5.6-Cyber model, specifically trained for exploit creation and authentication bypass, which demonstrates a 95 percent success rate on cybersecurity requests compared to 1.5 percent for the standard model.

image
image

What happened

OpenAI released two specialized variants based on GPT-5.6 through the expanded Daybreak program. The Daybreak Blue tier provides GPT-5.6 Sol with security filters removed for defense and pentesting tasks. The Daybreak Red tier provides access to the new GPT-5.6-Cyber model, trained for offensive research: exploit creation, authentication bypass, privilege escalation. In field trials, GPT-5.6-Cyber discovered a critical vulnerability CVE-2026-15903 in the V8 engine for Google Chrome, more than 400 privilege escalation vulnerabilities in operating system kernels, 5 vulnerabilities in mobile OSes, and 3 critical bugs in a popular database. From September 1, 2026, all Daybreak accounts are required to use hardware security keys, and high-risk calls via Codex require manual confirmation in auto-review mode.

Context

The split into Blue and Red tiers is technically justified — defensive and offensive tasks are optimized on different distributions with different loss functions and eval metrics. Opening Daybreak Blue without filters showed that the base GPT-5.6 Sol model has the capability for offensive tasks, but filtering was suppressing this potential. The expansion of Daybreak with the introduction of the Red tier is a direct competitive response to Anthropic's Project Glasswing (Claude Mythos Preview). Competition between providers is forming a new market for specialized cyber-LLMs and setting a standard: two-tier access with hardware keys and auto-review.

Why this matters for the industry

OpenAI is trying to get ahead of malicious actors by providing verified defenders with access to AI models with maximum cyber-offensive power before such models fall into the hands of adversaries. The 63-fold gap (95 percent versus 1.5 percent) between the specialized and universal model confirms that targeted fine-tuning of LLMs in the cybersecurity domain provides a qualitative leap. Companies already working with Daybreak (SpecterOps and partners) gain a competitive advantage: vulnerability research tasks have been reduced from weeks to less than one day. For startups, this means growing dependence on infrastructure giants — the AI-assisted pentesting market is moving out of the experimental phase.

Why this matters for users

Information security professionals can apply for access through OpenAI's partner program — access to Daybreak Blue for defensive tasks is already open. The Chrome V8 vulnerability (CVE-2026-15903) with RCE and sandbox escape is already being patched — it is worth ensuring that the browser is updated. For most teams, access to GPT-5.6-Cyber at the Red tier remains limited to the partner program. The mandatory transition to hardware security keys from September 1 means that program users must prepare the corresponding equipment.

What is still unknown / limitations

Access to GPT-5.6-Cyber is currently limited to the partner program and is not available to a wide range of specialists. The Blue/Red pattern is technically justified specifically for cybersecurity — the separation of offensive and defensive tasks has a fundamental basis, but it is not a universal template for other domains like biotech or finance. The risk of compromising access to Daybreak Red remains: malicious actors will get an automated state-of-the-art class tool for attacks. Independent verification of 400+ vulnerabilities in OS kernels and other findings has not yet been presented.

Sources

Author

Look at AI, editorial team