GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

OpenAI's August GPT-5.6 Update: Sol Clears 4 Cyber Scenarios Luna Cannot — Both Stay 'High Risk'

OpenAI today shipped updated versions of GPT-5.6 Sol and GPT-5.6 Luna to ChatGPT, replacing GPT-5.5 Instant as the default model. Free and Go users get the new Luna as their default. Plus and Pro users get an updated Sol with an effort slider that lets them control how much reasoning the model applies to a response.

The August system card — published simultaneously at deploymentsafety.openai.com — keeps both models at High capability in Cybersecurity and Biological and Chemical domains under OpenAI’s Preparedness Framework. Neither crosses the High threshold for AI Self-Improvement.

The Cyber Split

The most specific new data in the August update is the scenario-level failure breakdown. Both Sol and Luna were run against OpenAI’s cyber range. The results:

GPT-5.6 Sol (August) — fails 2 of 6 scenarios:

  • Firewall Evasion
  • CA/DNS Hijacking

GPT-5.6 Luna (August) — fails 5 of 6 scenarios:

  • Leaked Token
  • Binary Exploitation
  • Firewall Evasion
  • EDR Evasion
  • CA/DNS Hijacking

Sol has meaningfully better cyber performance than Luna at the scenario level, clearing Leaked Token, Binary Exploitation, and EDR Evasion that Luna cannot. Both still fail on the two most infrastructure-level attacks (Firewall Evasion and CA/DNS Hijacking), which appear to be the current hard ceiling for this model family.

The range was also run incomplete: infrastructure challenges in porting the evaluation meant only 34 of 40 benchmark challenges were executed. The missing six scenarios were not disclosed.

AI Self-Improvement: Not Evaluated

OpenAI did not run AI Self-Improvement evals for the August release. The stated reason: GPT-5.6 Sol’s capabilities are similar to the July release across several intelligence evaluations, putting it below the High Capability threshold without requiring fresh measurement. This is the first GPT-5.6 system card that omits that dimension entirely.

What’s New in the Model

The August models replace GPT-5.5 Instant for ChatGPT users. Codex and ChatGPT Work users are still on the July versions of Sol and Luna — the August and July releases are now explicitly distinguished in OpenAI’s documentation.

Sol gets an effort slider in Plus and Pro, letting users choose how much reasoning the model applies. This is the same architecture as the July release, not a new training run — the system card describes it as an updated representative, not a new model family.

On content safety, Sol performs comparably to GPT-5.5 Instant June Update across disallowed categories with two exceptions: gore and disallowed sexual content. Luna also shows differences on gore. OpenAI applies a system-level mitigation for disallowed sexual content before it reaches users.

U18 Evals: A First

For the first time, OpenAI’s system card includes dedicated evaluations for users under 18. The August update adds model-level training to prevent romantic roleplay with teens, discourage age-restricted challenges, and enforce appropriate boundaries around eating disorders, body image, and age-restricted goods. System-level protections supplement the model training.

What Hasn’t Changed

Both models remain under the same safeguard set as the original GPT-5.6 System Card. High capability in Cybersecurity means heightened monitoring and access controls remain in place. The only change from the June system card is the August capability data — the policy posture is unchanged.

The scenario failures on Firewall Evasion and CA/DNS Hijacking are unchanged between Sol and Luna, suggesting those attack types represent a structural gap in how GPT-5.6 was trained or evaluated, not just a calibration difference between the two tiers.