The October system card reports stronger jailbreak resistance alongside teen-safety setbacks, additional content filters, and High cyber and biological capability ratings below Critical.
Listen
AI narration
12:46
0:00 / 12:46
AI SummaryGenerated from this article
OpenAI's October GPT-6 models show improved jailbreak resistance but report statistically significant regressions in teen-safety evaluations including age-restricted content, sexual content, emotional reliance, and gore. Both models rank High in cybersecurity and biological capabilities, below Critical thresholds. Additional system-level classifiers for self-harm, sexual content, and gore provide mitigation, though OpenAI notes manual review found violations were generally lower severity.
The safety evaluations test deliberately challenging, production-derived examples without system-level safeguards to assess underlying model behavior. OpenAI cautions against interpreting benchmark results as everyday failure rates, emphasizing that improved jailbreak resistance does not establish better teen-safety performance across all categories.
OpenAI reports statistically significant regressions in several teen-safety evaluations for the October versions of GPT-6 Sol and GPT-6 Luna, even as it describes improved resistance to jailbreaks. The same disclosure places both models in its High capability category for cybersecurity and biological/chemical domains, below its Critical thresholds.
The October 7, 2026 system card presents a more complicated safety picture than a single claim of improvement. OpenAI says some underlying-model behaviors deteriorated, reviewed failures were generally low severity, and additional system-level protections help mitigate risks.
These findings need two qualifications upfront. They concern the October models deployed in ChatGPT, not the September versions still used in Codex and ChatGPT Work. And they are OpenAI’s evaluation results on deliberately challenging tests, not estimates of how often ordinary ChatGPT conversations produce unsafe responses.
October ChatGPT Models Are Not the Codex Versions
OpenAI distinguishes the GPT-6 models by release month because the same model names cover different versions across its products.
According to the system card, October GPT-6 Sol and Luna replace GPT-5.6 Sol and Luna in ChatGPT. Users accessing GPT-6 through Codex and ChatGPT Work continue to use previously released September versions.
That distinction limits how far readers can generalize the disclosure. A reported regression in an October ChatGPT model does not establish the same regression in the September model used elsewhere. Likewise, an improvement reported for October should not automatically be credited to every product carrying the GPT-6 name.
The comparison baseline also matters. OpenAI describes the October models’ improvements and regressions relative to their respective GPT-5.6 counterparts. That is not necessarily a measurement of how the October GPT-6 models changed from September GPT-6.
For anyone assessing these AI products, the release month and deployment surface are therefore part of the model’s identity, not incidental details.
Teen-Safety Regressions Span Several Categories
OpenAI’s dedicated under-18 evaluations test whether models maintain age-appropriate boundaries in sensitive conversations. The company says these evaluations include particularly challenging, production-derived examples and adversarial prompts.
The reported statistically significant regressions cover:
Work with Zeniteq
Let’s work together
We’re open to thoughtful collaborations with teams building in AI. Explore the ways we can work together.
Both October models: age-restricted content, sexual content, and emotional reliance.
GPT-6 Luna additionally: gore.
These categories should not be collapsed into a single “harmful content” score. OpenAI says its teen-specific policies impose more restrictive thresholds in areas where younger users may face heightened risks, including sexual content, emotional reliance, eating disorders, and access to age-restricted goods and services.
The emotional-reliance finding is especially distinct from a conventional refusal failure. It concerns the boundaries of an interaction with a younger user, rather than simply whether the model supplies prohibited information. Stronger resistance to an explicit jailbreak would not, by itself, demonstrate better behavior in that category.
OpenAI also reports regressions outside the dedicated teen evaluations. In its standard safety tests, GPT-6 Sol regressed on self-harm, while GPT-6 Luna regressed on self-harm, gore, and sexual content.
The company says manual review found the violations were borderline and generally safe or lower severity. It describes the models as more willing to answer informational questions about self-harm while directing users toward professional resources. According to OpenAI’s evaluations, they did not comply with requests to facilitate self-harm.
OpenAI further says adversarial red teaming did not identify high-severity risks under its evaluation criteria. Those are important qualifications, but they remain the company’s assessment of the tested failures. They do not turn a statistically significant regression into an improvement, nor establish that every possible failure would be similarly mild.
Extra Classifiers Sit Outside the Model Scores
OpenAI reports applying an additional classifier block for self-harm, sexual content, and gore. That protection sits at the system level, separate from the underlying behavior measured in the model evaluations.
This separation is central to interpreting the disclosure. OpenAI says its challenging-prompt safety evaluations run without system-level safeguards so that they can assess the model itself. The resulting scores therefore do not include every intervention that may affect a response delivered through ChatGPT.
The card also identifies system-level measures such as Trusted Contact, localized crisis helplines, and parental controls for younger users. OpenAI presents these as protections that contribute to the assistant’s overall safety but are not represented in the underlying-model scores.
Consequently, two apparently conflicting statements can both be accurate: a model can perform worse on a safety category in isolation, while the deployed product applies additional controls intended to reduce harmful responses.
That does not establish how completely the extra controls compensate for the regressions. The presence of a classifier is evidence of a mitigation, not proof of its effectiveness across every relevant conversation.
There is also a category mismatch worth preserving. The disclosed additional classifier block covers self-harm, sexual content, and gore. That list alone does not explain how the reported emotional-reliance and age-restricted-content regressions are addressed. It would be inaccurate to assume one mitigation resolves every teen-safety finding.
Jailbreak Resistance Improved, According to OpenAI
The negative findings sit alongside reported gains in resistance to adversarial manipulation.
Compared with GPT-5.6 Sol and Luna, OpenAI says the October GPT-6 models better resisted jailbreaks, including attacks that adapt across multiple conversational turns. It also reports reductions in dishonesty, deception, and circumvention of guardrails.
Adaptive, multi-turn attacks are relevant because they test more than whether a model rejects one obviously prohibited prompt. They probe whether an attacker can change tactics as the conversation develops and push the model away from its intended boundaries.
OpenAI also reports that the October models are less likely to refuse harmless requests or add excessive, judgmental caveats. Its account is therefore not that the models became uniformly more restrictive. Rather, the company describes improved helpfulness on legitimate requests alongside stronger performance against some attacks and weaker results in certain safety categories.
Those dimensions should remain separate. Better jailbreak resistance does not establish better teen-safety behavior, and fewer unnecessary refusals do not establish that all newly answered requests are appropriate.
The card’s aggregate assessment depends partly on OpenAI’s judgment about the nature and severity of the failures, not solely on whether every individual benchmark moved upward.
High Cyber and Bio Capability Does Not Mean Critical
Under its Preparedness Framework, OpenAI classifies both October models as High capability in cybersecurity and biological and chemical domains. The company reports that neither reaches its Critical thresholds, and neither reaches High capability in AI self-improvement.
These classifications concern capability. They should not be read as estimates of the probability that a normal ChatGPT response will cause harm.
Capability assessments and content-safety assessments answer different questions. The former examine what a model can accomplish under testing conditions; the latter examine whether its responses follow safety policies. A High capability rating does not, by itself, establish frequent harmful behavior, while strong refusal behavior does not erase the significance of the underlying capability.
OpenAI says it evaluates these two areas differently. Safety evaluations run at the models’ lowest deployed reasoning settings to capture behavior relevant to the vast majority of usage. Capability assessments use maximum reasoning effort to estimate an upper bound.
The company says it has implemented the same set of safeguards for the October models as those detailed for GPT-5.6 Sol and Luna.
“Below Critical” should therefore be read precisely: OpenAI says the models have not crossed that higher threshold. It is not a declaration of negligible risk. The High classifications are the basis for safeguards, even though the company assesses the models as remaining below Critical.
OpenAI explicitly warns against interpreting these benchmark results as ordinary production failure rates.
Its Production Benchmarks use conversations representative of challenging examples drawn from production data. The company says they were deliberately built around cases where existing models were not yet producing ideal responses, after earlier evaluations became relatively saturated.
That selection makes them useful for exposing weaknesses. It also makes them unsuitable for estimating how frequently those weaknesses appear across all ChatGPT traffic. A test set concentrated on difficult cases is not a random sample of everyday usage.
The primary metric is safe completion rate: whether a model’s response and actions comply with OpenAI’s safety policies. A statistically significant regression indicates a detectable difference under the evaluation setup. Without considering the size and nature of that difference, it does not establish a corresponding increase in real-world harm.
Historical comparisons require similar care. OpenAI says comparison values may reflect later versions of previously launched models. Policies, graders, datasets, and other evaluation details also evolve, so scores in older system cards should not generally be treated as directly comparable unless included in the current comparison.
The consequential unanswered question is how the disclosed model-level regressions behave after ChatGPT’s system-level protections are applied. OpenAI says it continues monitoring these categories after launch and can adjust safeguards.
The strongest evidence for that next stage would connect the safeguards to the specific failures they are intended to prevent, particularly for younger users. Until then, the system card supports a bounded conclusion: OpenAI reports stronger resistance to some forms of manipulation, but that progress does not settle the separate questions raised by its teen-safety regressions.