Skip to content
Automation Squad
A concrete service tunnel forking into two passages, one lit blue and one lit red
Security·2 min read·By the Automation Squad Research

OpenAI Split Its Cyber Models in Two

A 63-point gap in refusal rates, decided entirely by who you can prove you are.

Robert MacKelfresh

By Robert MacKelfresh

Founder, Automation Squad ·

The short answer

OpenAI launched GPT-5.6-Cyber on August 10, 2026 and split its Daybreak programme into Blue and Red tiers. On OpenAI's internal Advanced Cybersecurity Completion Rate evaluation, GPT-5.6-Cyber completes 95.0% of advanced requests against 1.5% for stock GPT-5.6 Sol and 2.0% for Sol under Daybreak Blue.

Refusal table

The completion rates, and what the split actually gates

The numbers are OpenAI's own, on OpenAI's own evaluation. Read them as a description of policy rather than of capability.

  1. Read the 1.5% before the 95%

    The stock model refuses almost every advanced security request. That is the status quo a lot of legitimate security work has been running into, and it is the reason this tier exists at all.

  2. Understand that the gap is policy, not capability

    The same underlying family produces 1.5% and 95% depending on which access tier you are in. What changed is permission and verification, not intelligence — which is a useful way to read most safety announcements.

  3. If you do authorised security work, find out whether you qualify

    Launch partners span consultancies and security vendors — Accenture, IBM, PwC, KPMG, EY, Capgemini, Cognizant, NCC Group, SpecterOps, plus Palo Alto Networks, CrowdStrike, Cisco, Sophos, Akamai, Fortinet and Cloudflare. Access runs through vetting rather than a signup form.

  4. If you don't, take the structure as the lesson

    Tiered access with refusal rates set by verified identity is a template. Expect it wherever a capability is genuinely dual-use, and expect the interesting question to be who does the vetting.

Model / tierAdvanced cyber request completion
GPT-5.6-Cyber95.0%
GPT-5.5-Cyber57.3%
GPT-5.6 Sol under Daybreak Blue2.0%
Stock GPT-5.6 Sol1.5%
Daybreak BlueFrontier models with defensive guardrails
Daybreak RedSeparately vetted offensive access
Model IDsgpt-5.6-cyber, daybreak-red-latest, daybreak-blue-latest

Take it with you

OPENAI GPT-5.6-CYBER + DAYBREAK BLUE/RED — Aug 10, 2026
Source: https://automationsquad.com/news/openai-daybreak-blue-red-split/

ADVANCED CYBER REQUEST COMPLETION RATE
(OpenAI's internal evaluation)
  GPT-5.6-Cyber ..................... 95.0%
  GPT-5.5-Cyber ..................... 57.3%
  GPT-5.6 Sol under Daybreak Blue ... 2.0%
  Stock GPT-5.6 Sol ................. 1.5%

THE SPLIT
  Daybreak BLUE .... frontier models with defensive guardrails
  Daybreak RED ..... separately vetted offensive access

  Model IDs: gpt-5.6-cyber · daybreak-red-latest · daybreak-blue-latest

THE THING TO NOTICE
1.5% -> 95% is the SAME model family. What changed is permission and
verified identity, not intelligence. Read most safety announcements
this way and they get clearer.

[ ] Do I do authorised security work? -> check qualification (vetting,
    not a signup form)
[ ] If not: the transferable lesson is the STRUCTURE — tiered access
    with refusal rates set by verified identity. Ask who does the
    vetting, what's logged, and what happens to the log.

OpenAI launched GPT-5.6-Cyber on August 10, 2026, built on GPT-5.6 Sol and explicitly trained for finding vulnerabilities and building exploit chains, and split its Daybreak programme into two tiers.

The facts: on OpenAI's internal Advanced Cybersecurity Completion Rate evaluation, GPT-5.6-Cyber completes 95.0% of advanced requests, against 1.5% for stock GPT-5.6 Sol, 2.0% for Sol under Daybreak Blue, and 57.3% for the previous GPT-5.5-Cyber. Daybreak Blue provides general-purpose frontier models with defensive guardrails; Daybreak Red is separately vetted offensive access. Model IDs are gpt-5.6-cyber, daybreak-red-latest and daybreak-blue-latest. Launch partners include Accenture, IBM, PwC, KPMG, EY, Capgemini, Cognizant, NCC Group and SpecterOps, alongside Palo Alto Networks, CrowdStrike, Cisco, Sophos, Akamai, Fortinet and Cloudflare.

Automation Squad's take: the number worth sitting with is 1.5%, not 95%. A stock frontier model refusing 98.5% of advanced security requests is the environment legitimate security professionals have been working in, and it is the reason a separate tier had to exist. The 63-point swing between Blue and Cyber is produced by verification rather than by a better model, which is a clarifying way to read the whole category — most of what gets described as a safety property is a policy applied to the same underlying capability. Whether identity verification is a durable answer to dual-use is genuinely unresolved. What is clear is that refusing everyone was not working, and this is the first serious attempt at the alternative.

Run this now: if your organisation does authorised offensive security work and has been quietly routing around refusals, this is the moment to find out whether you qualify for vetted access rather than continuing to fight the stock model. If that is not you, the transferable question is worth carrying into your next vendor conversation about any gated capability: who performs the vetting, what exactly gets logged, and who can later read that log.

Questions people are asking

What is the difference between Daybreak Blue and Daybreak Red?
Blue provides general-purpose frontier models with defensive guardrails. Red is separately vetted access for offensive security work. The two carry very different refusal thresholds, gated by verification rather than by capability.
How much more permissive is GPT-5.6-Cyber?
On OpenAI's internal Advanced Cybersecurity Completion Rate evaluation it completes 95.0% of advanced requests, against 1.5% for stock GPT-5.6 Sol and 2.0% for Sol under Daybreak Blue. Its predecessor GPT-5.5-Cyber sat at 57.3%.
Who can access these models?
Access runs through vetting rather than self-service. Launch partners include Accenture, IBM, PwC, KPMG, EY, Capgemini, Cognizant, NCC Group and SpecterOps, alongside security vendors including Palo Alto Networks, CrowdStrike, Cisco, Sophos, Akamai, Fortinet and Cloudflare.

Last checked August 10, 2026 against the primary sources above, by Automation Squad Research. Spot an error? [email protected].

Related artifacts

More news