GPT-6 Sol is two versions under one name: October's in ChatGPT chat, September's in Codex, 15 days apart
OpenAI's 7 October card says ChatGPT chat now runs the October GPT-6 Sol and Luna, while Codex and Work keep the 22 September versions. Both are rated High for cyber, October's prompt injection figure is lower than September's, and no page read shows how to choose between them.
By Parminder Kumar Sharma · · 25 min read

One name, two versions, 15 days apart
OpenAI's system card of 7 October 2026 says that GPT-6 Sol and GPT-6 Luna in ChatGPT chat are now the "October" versions, and that the GPT-6 Sol and GPT-6 Luna in Codex and ChatGPT Work are "September" versions of the same names. The September versions launched on 22 September, so the two are 15 days apart (derived: 8 days left in September plus 7 in October). OpenAI's cards rate both pairs High for cybersecurity under its Preparedness Framework, the same label GPT-5.6 carried. Two other models in the family, GPT-6 Astra and GPT-6.1 Sol, are rated Critical.
That does not show the October version is better or worse than September's, that anyone has misused either, or that OpenAI concealed the split: the card states it in its first paragraph. What it does not establish is the thing an organisation needs, which is which of the two a member of staff is using at a given moment. OpenAI's pages use the name GPT-6 Sol for the chat model, for a Codex and Work model and for an API identifier, and its help page says the chat menu entry is "not labeled Sol". A safety figure measured at the lowest reasoning setting is not the behaviour at the highest setting you may select. High is OpenAI's own category under its own framework. A speed claim is not your latency. A launch page that says "stronger resistance" gives no number, and the card's number depends on which model it is compared with.
The announcement states no cyber rating: it does not mention the Preparedness Framework, High capability or Critical capability. The rating sits in the card's introduction. OpenAI sells these models, wrote the card and ran every evaluation in it. Each figure below is attributed to the page it was read from, and a figure marked derived is our arithmetic. OpenAI says the model is "built for more than 1.2 billion people who use ChatGPT each week"; this briefing uses that only as a statement of reach, not as a count of who runs which version.
GPT-6 is not one thing, and "for everyone" is not one product
The announcement is titled "GPT-6 and Intelligent UI for everyone". OpenAI's own documentation shows five different things behind that phrase. The table says where each runs, what OpenAI's pages say, and what they do not.
Which GPT-6 runs where, as OpenAI's pages state it on 8 October 2026. Sources: the announcement, the October card, OpenAI's ChatGPT, Enterprise and Work and Codex help pages, and the API model pages and changelog.
- Where
- Chat tab: Plus, Pro, Business, Enterprise
- What OpenAI's pages say runs there
- GPT-6 Sol (October), shown in the menu as GPT-6. Existing "Latest" selections move to it.
- Not stated
- A menu label that says Sol or October. A way to stay on the September version.
- Where
- Chat tab: Free and Go
- What OpenAI's pages say runs there
- GPT-6 Luna (October) from 8 October per the announcement. The model help page still says GPT-5.6 Luna.
- Not stated
- Which of the two pages is current.
- Where
- Chat tab: Pro option
- What OpenAI's pages say runs there
- GPT-6 Pro, powered by GPT-6 Astra, a model OpenAI rates Critical. No Intelligent UI.
- Not stated
- Any mention of the rating on the model-selection pages.
- Where
- Work and Codex
- What OpenAI's pages say runs there
- GPT-6 Sol, GPT-6.1 Sol and GPT-6 Luna, described as "separate model options", and GPT-6 Astra. The card calls the Sol and Luna there the September versions.
- Not stated
- Which version a policy means when it says GPT-6 Sol.
- Where
- API
- What OpenAI's pages say runs there
- gpt-6-sol and gpt-6-luna. The chat-latest alias "points to the latest model available in ChatGPT" for paid tiers, updated 7 October.
- Not stated
- A dated snapshot or an October identifier. Whether gpt-6-sol is the September version.
| Where | What OpenAI's pages say runs there | Not stated |
|---|---|---|
| Chat tab: Plus, Pro, Business, Enterprise | GPT-6 Sol (October), shown in the menu as GPT-6. Existing "Latest" selections move to it. | A menu label that says Sol or October. A way to stay on the September version. |
| Chat tab: Free and Go | GPT-6 Luna (October) from 8 October per the announcement. The model help page still says GPT-5.6 Luna. | Which of the two pages is current. |
| Chat tab: Pro option | GPT-6 Pro, powered by GPT-6 Astra, a model OpenAI rates Critical. No Intelligent UI. | Any mention of the rating on the model-selection pages. |
| Work and Codex | GPT-6 Sol, GPT-6.1 Sol and GPT-6 Luna, described as "separate model options", and GPT-6 Astra. The card calls the Sol and Luna there the September versions. | Which version a policy means when it says GPT-6 Sol. |
| API | gpt-6-sol and gpt-6-luna. The chat-latest alias "points to the latest model available in ChatGPT" for paid tiers, updated 7 October. | A dated snapshot or an October identifier. Whether gpt-6-sol is the September version. |
The second row is the sharpest. The announcement says Free and Go get GPT-6 Luna from tomorrow, and the Intelligent UI help article says the same. The model-selection page, updated about 18 hours before it was read, says Free and Go "receive GPT-5.6 Luna" and that Think uses GPT-5.6 Luna. Either that page is behind or the rollout is. A reader cannot tell which from OpenAI's pages, and a policy that names "the free model" would be wrong in one of the two cases.
The same pattern has happened before. On 6 August OpenAI updated GPT-5.6 Sol in ChatGPT chat and said Work and Codex were not changing; its August card calls the result the "August" release. The October split is therefore the second time in 62 days (derived: 6 August to 7 October) that chat has had a changed model under an unchanged name. OpenAI is open about it each time. Being open is not the same as giving you a label you can write into a policy.
The cyber rating is High: the same as September, and lower than two models in the same family
Our earlier briefing on dots and GPT-6.1 Sol reported that OpenAI rates Astra and Sol Critical for cybersecurity and that no launch page said so. The Sol in that briefing was GPT-6.1 Sol, launched on 29 September. The Sol in ChatGPT chat is GPT-6 Sol, and OpenAI's cards give it a different rating. The briefing on GPT-6.1 Astra covers the other Critical model.
So the answer to "is High a lower rating than before?" is no: the September Sol and Luna were High, and so was GPT-5.6. The answer to "is it lower than the Sol OpenAI launched a week earlier?" is yes. The October card treats both as High but below Critical, as it did for the August GPT-5.6 models. Its stated reason is one sentence about capability and one about process: GPT-6 Sol "performed comparably to GPT-5.6 Sol without a clear improvement in capabilities", and the Safety Advisory Group "had sufficient evidence to recommend that both models were below the critical threshold". The paragraph giving that reason is identical in the September appendix (checked). The October card never mentions GPT-6.1 Sol.
What each card prints for GPT-6 Sol and GPT-6 Luna on OpenAI's cyber evaluations, Sol then Luna. September: the appendix to the Astra card, 22 September 2026. October: the card of 7 October 2026. All run by OpenAI. ExploitBench is at maximum reasoning effort.
- Evaluation
- ExploitBench
- September card
- 81.7%, 43.4%
- October card
- 82.62%, 44.66%
- Evaluation
- SEC-Bench Pro (Astra: 85.4%)
- September card
- 66.3%, 34.2%
- October card
- 68.85%, 48.50%
- Evaluation
- ExploitBench internal port, flaws disclosed June to August (Astra: 31.5%)
- September card
- 5.5%, 0%
- October card
- Not reported
- Evaluation
- ExploitGym
- September card
- 22.1%, 11.6%
- October card
- Not reported
- Evaluation
- Sandbox Bench, targets of 22 (Astra: 10)
- September card
- 1, 1
- October card
- Not reported
- Evaluation
- Static jailbreak test, cyber: share of attacks refused
- September card
- 79.8%, 87.5%
- October card
- Not reported
| Evaluation | September card | October card |
|---|---|---|
| ExploitBench | 81.7%, 43.4% | 82.62%, 44.66% |
| SEC-Bench Pro (Astra: 85.4%) | 66.3%, 34.2% | 68.85%, 48.50% |
| ExploitBench internal port, flaws disclosed June to August (Astra: 31.5%) | 5.5%, 0% | Not reported |
| ExploitGym | 22.1%, 11.6% | Not reported |
| Sandbox Bench, targets of 22 (Astra: 10) | 1, 1 | Not reported |
| Static jailbreak test, cyber: share of attacks refused | 79.8%, 87.5% | Not reported |
By the printed figures ExploitBench rose 0.92 points for Sol and 1.26 for Luna, and SEC-Bench Pro rose 2.55 points for Sol and 14.30 for Luna (derived). The card says Luna "underperformed" Sol and does not remark on that jump. It prints two of the six cyber results in the table again, and does not say whether the other four were re-run. The one it drops that matters most is the test built from flaws disclosed after the models' knowledge cutoff: Sol scored 5.5% on it against Astra's 31.5%, which is 17% of Astra's rate (derived), and GPT-6.1 Sol scored 21.5%. That test separates the Critical models from the High ones. The October card does not say what the October Sol would score.
Which framework. The card links OpenAI's Preparedness Framework, version 2, and OpenAI's page for it is dated 15 April 2025. On 18 August 2026 OpenAI wrote that it needs an approach that "extends beyond the current Preparedness Framework" and that it will evolve the framework. The October card still links version 2, and we found no later version on OpenAI's safety pages on 8 October. High and Critical are therefore labels from a framework OpenAI has said is being superseded, applied by OpenAI to its own models.
High is a treatment. The card's wording is that OpenAI is "treating" these models as High. In biology its own Table 8 shows Sol below the indicative High threshold on two of four evaluations (ProtocolQA 25.62% against 54%, tacit knowledge 72.45% against 80%) and Luna below it on one. For cyber the card leans on the determination being the same as GPT-5.6's. Our reading, which is inference, is that High tells you which safeguards OpenAI applies. It is not a measured score you can compare across vendors.
What High implies for how the model is gated. The October card says it applies "the same set of safeguards" as the GPT-5.6 card. That card describes refusals; real-time monitors that restrict "scaled agentic vulnerability research and chained exploit development for users outside of our trusted access program"; account-level enforcement that can end in a ban; and a Trusted Access for Cyber programme, Daybreak, for verified defenders. It also says its activation classifiers "are trained and tuned separately for each model", and it names Sol and Terra for them, not Luna. The October card does not say whether they were retuned for the October models.
The Critical models run on a different stack. GPT-6.1 Sol "uses the same safeguards stack as GPT-6 Astra", and the Astra card adds misalignment monitoring to all tool-using inference. The same Chat tab reaches both kinds: GPT-6 by default, and the Pro option, GPT-6 Pro on Astra, for Pro, Business and Enterprise plans. For Enterprise, OpenAI says Astra has its own access control and was off by default at launch.
Safety was measured at the lowest setting, capability at the maximum
The card's introduction sets two rules. For all safety evaluations it measures models "at their lowest reasoning deployment settings", to capture "the vast majority of usage". For capability assessments it evaluates them at "maximum reasoning effort to get an upper bound of capabilities". The rule matters because the two ends of the dial behave differently, and a figure is only a figure of the setting it was measured at.
Of the card's 18 results sections (our count), five state the setting in the section itself: four say maximum effort and one says its figure is averaged across reasoning levels. The other 13 state none and rely on the introduction. The alignment sections say maximum, so whatever rule covers them is not the introduction's lowest-setting rule, which names disallowed content and mental health as its examples. And the card never says how its settings map to what a person can choose: ChatGPT offers Instant to Extra High, the API none to max.
Where the October card states a reasoning setting, and what each figure is therefore a figure of. Section numbers are the card's.
- Section
- 4.2 Prompt injection
- Setting the card states
- Averaged equally across available reasoning levels
- What the figure is a figure of
- A mean over levels. No per-level figure is given.
- Section
- 7.1.1 Auto-review, 7.1.2 Warnings, 7.2.2 Broken search
- Setting the card states
- Maximum reasoning effort
- What the figure is a figure of
- The top setting, in simulated tasks. Auto-review and Warnings say they ran without system-level safeguards.
- Section
- 8.1.2.1 ExploitBench
- Setting the card states
- Maximum reasoning effort
- What the figure is a figure of
- An upper bound of capability, not the Chat default.
- Section
- 13 other sections, including safe completions, under-18, vision, jailbreaks, health, hallucinations, blocker deception, biology and SEC-Bench Pro
- Setting the card states
- None in the section
- What the figure is a figure of
- The introduction's rule, if it holds. Several say they ran without production safeguards.
- Section
- Nowhere
- Setting the card states
- How the card's settings map to Instant through Extra High
- What the figure is a figure of
- No printed figure can be placed on a slider position.
| Section | Setting the card states | What the figure is a figure of |
|---|---|---|
| 4.2 Prompt injection | Averaged equally across available reasoning levels | A mean over levels. No per-level figure is given. |
| 7.1.1 Auto-review, 7.1.2 Warnings, 7.2.2 Broken search | Maximum reasoning effort | The top setting, in simulated tasks. Auto-review and Warnings say they ran without system-level safeguards. |
| 8.1.2.1 ExploitBench | Maximum reasoning effort | An upper bound of capability, not the Chat default. |
| 13 other sections, including safe completions, under-18, vision, jailbreaks, health, hallucinations, blocker deception, biology and SEC-Bench Pro | None in the section | The introduction's rule, if it holds. Several say they ran without production safeguards. |
| Nowhere | How the card's settings map to Instant through Extra High | No printed figure can be placed on a slider position. |
A worked example. The card's Respecting Warnings chart says that in 28.1% of rollouts at maximum effort Sol (October) hit a barrier and found a way round it, against 33.9% for GPT-5.6 Sol (August). That is behaviour at the top setting in a simulated workplace with no system controls, which the card says "would plausibly stop the observed circumventions". It is not a rate for your staff.
Prompt injection: three cards, three numbers for the same names
The card reports prompt injection in one table. GPT-6 Sol (October) scores 97.13% and GPT-6 Luna (October) 95.80% on the GPT-Red indirect test, averaged equally across available reasoning levels. On instruction hierarchy it says both "saturate" the test, at 99.99% and 99.79%. Read as the share of attacks that get through (100 minus robustness, derived), that is 2.87% for Sol, about 1 in 35, and 4.20% for Luna, about 1 in 24.
Prompt injection robustness for the same names on three cards, higher is better. Indirect means a malicious instruction in third-party content; instruction hierarchy means a user overriding system or developer instructions. The GPT-5.6 column is 100 minus the attack-attempt success rates printed in the GPT-5.6 card, for GPT-5.6 Sol and Luna (derived; for Sol, 100 minus 3.77 is the 96.23% the Astra card prints).
- Test
- Indirect, Sol
- September card
- 99.050%
- October card
- 97.13%
- GPT-5.6 card
- 96.23%
- Test
- Indirect, Luna
- September card
- 98.605%
- October card
- 95.80%
- GPT-5.6 card
- 97.06%
- Test
- Instruction hierarchy, Sol
- September card
- 99.97%
- October card
- 99.99%
- GPT-5.6 card
- 99.949%
- Test
- Instruction hierarchy, Luna
- September card
- 99.97%
- October card
- 99.79%
- GPT-5.6 card
- 99.89%
| Test | September card | October card | GPT-5.6 card |
|---|---|---|---|
| Indirect, Sol | 99.050% | 97.13% | 96.23% |
| Indirect, Luna | 98.605% | 95.80% | 97.06% |
| Instruction hierarchy, Sol | 99.97% | 99.99% | 99.949% |
| Instruction hierarchy, Luna | 99.97% | 99.79% | 99.89% |
Against GPT-5.6 Sol, October Sol is better on the indirect test: 2.87% getting through against 3.77%, a 24% reduction (derived). Against the September GPT-6 Sol it is not: 2.87% against 0.95%, about three times as many (3.02, derived). Luna has the same shape: 4.20% against 1.395%, 3.01 times as many, and 4.20% against the 2.94% printed for GPT-5.6 Luna. That can be a change in the models, a change in how the figure is averaged, or both. The September card averaged "per defender query"; the October card averages equally across reasoning levels. The October card does not compare itself with September on prompt injection, as it does for multi-turn jailbreaks, and does not say which explanation applies.
What the number is made of. OpenAI ran it, with GPT-Red, its own automated attacker. The GPT-5.6 card, which the October card cites for the method, says the attacker controls a single message as the final attack and counts successful attacks against defender queries. The October card gives no count of trials or queries and no confidence interval for this table, although it gives intervals for the multi-turn jailbreak chart, and it names no outside evaluator. The Astra card had one: Gray Swan's IPI Arena put Astra's attack success at 8.5% against 27.0% for GPT-5.6 Sol, on 1,810 curated attacks with 15 attempts per scenario. It is not repeated for the October models. A rate per attempt and a rate over 15 attempts are different things.
What the card does not cover. The words Intelligent UI, interactive, browsing and connector do not appear in it (searched). The Astra card says OpenAI retired its earlier connector and search-and-function-calling tests as saturated, and the October card does not restore them. So the card reports no prompt injection result for an interactive answer, for web search results flowing into one, or for a connected mailbox. OpenAI's report on self-replicating prompt injection, covered in our earlier briefing, is the other place it discusses the connector case, and it gives no rate.
The NCSC wrote on 8 December 2025 that prompt injection attacks "may never be totally mitigated" and that those who own the risk should know it "will remain a residual risk". A high percentage is a reason to worry less about one attempt. It is not a control.
Intelligent UI: what OpenAI says is produced, and what no page says
Intelligent UI is the headline feature. GPT-6 is trained "to compose responses using text, visuals, and interactive elements", including "tappable buttons, forms, charts, and interactive experiences you can use directly in your conversation", and no special prompt is required (help page). OpenAI describes the machinery in two phrases: "a library of native, streamable components" and "a compiler that processes the interface as the model generates it". Interactive content was already in ChatGPT: the release notes list interactive code-block previews in February, interactive learning modules in March and interactive flashcards on 22 September. What changes is that the model now chooses to use it.
What OpenAI's pages state about Intelligent UI and what they leave out. Sources: the announcement, the Intelligent UI help article, the GPT-6 and Enterprise model pages and the October card, read on 8 October 2026.
- Question
- How the interface is made
- Stated
- A library of native components and a compiler that builds the interface as the model generates it.
- Not stated
- Whether the model writes code or a description of components, and whether any generated code is executed.
- Question
- Where it runs
- Stated
- Web and updated mobile apps, Instant to Extra High. Not in Work, Voice, the Pro option or older desktop apps.
- Not stated
- Any sandbox, origin or isolation between a component and the rest of the page.
- Question
- What a component can reach
- Stated
- Some components, such as checklists, keep state when the chat is refreshed, not across chats. Connected plugins such as Gmail, Health and Finance "continue to work".
- Not stated
- Whether a component can make network requests, load remote content, submit a form to a third party or call a connected plugin.
- Question
- What an action is
- Stated
- OpenAI says ChatGPT can "combine UI, data, and actions".
- Not stated
- What an action can do, and whether a person approves it first.
- Question
- Controls
- Stated
- A web setting, Layout and visuals, "reduces" interactive elements, "though some may still appear". In Enterprise, GPT-6 follows the existing GPT-6 Sol access setting.
- Not stated
- A switch for interactivity alone. Whether turning off Sol access in Chat also turns it off in Work and Codex.
- Question
- Security testing
- Stated
- Nothing in the October card.
- Not stated
- Any prompt injection, output handling or red-team result for the feature.
| Question | Stated | Not stated |
|---|---|---|
| How the interface is made | A library of native components and a compiler that builds the interface as the model generates it. | Whether the model writes code or a description of components, and whether any generated code is executed. |
| Where it runs | Web and updated mobile apps, Instant to Extra High. Not in Work, Voice, the Pro option or older desktop apps. | Any sandbox, origin or isolation between a component and the rest of the page. |
| What a component can reach | Some components, such as checklists, keep state when the chat is refreshed, not across chats. Connected plugins such as Gmail, Health and Finance "continue to work". | Whether a component can make network requests, load remote content, submit a form to a third party or call a connected plugin. |
| What an action is | OpenAI says ChatGPT can "combine UI, data, and actions". | What an action can do, and whether a person approves it first. |
| Controls | A web setting, Layout and visuals, "reduces" interactive elements, "though some may still appear". In Enterprise, GPT-6 follows the existing GPT-6 Sol access setting. | A switch for interactivity alone. Whether turning off Sol access in Chat also turns it off in Work and Codex. |
| Security testing | Nothing in the October card. | Any prompt injection, output handling or red-team result for the feature. |
None of this says Intelligent UI is unsafe. It says the security properties of this surface are not described. The site's pattern PIP-014, rendered-output exfiltration covers the class of risk that applies whenever an interface renders what a model writes: "The vulnerability lives in the renderer, not the model", so a model-level robustness percentage does not measure it. That is the site's own text about a class of weakness, not a finding about ChatGPT. The pattern's three controls (do not auto-load model-authored references, apply a content security policy to the rendering surface, treat model output as untrusted input) are for anyone who embeds or captures this output in their own systems.
The NCSC's interim advice on agentic AI, 20 August 2026, says a model's built-in safeguards "should not be treated as holistic". OpenAI's own 2 October guide to the GPT-6 family tells builders to test before deploying and "measure task success, latency, and cost per successful task".
The speed claim is one number with no stated method
For questions that need web search, the announcement says "GPT-6 Instant starts answering 44% sooner, on average, than GPT-5.6 Instant". The paragraph around it describes two internal evaluations. This sentence is not labelled as one, and nothing says how it was measured. Not stated: what "starts answering" counts (first token or first visible content), how many questions, which tier or region, what load, the spread. The page also says GPT-6 builds answers "across multiple partial responses", so an earlier start is not an earlier complete answer, and no time to complete is given. The page's other speed claim, from an internal evaluation of agentic tasks, is that Extra High begins answering "in the same amount of time as GPT-5.6 Medium".
For scale, Artificial Analysis measured GPT-6 Sol at maximum effort through OpenAI's API and gives a time to first token of 137.69 seconds, against a median of 3.76 seconds for similar reasoning models (36.6 times, derived). That is the API at the top setting, not ChatGPT, so it does not test OpenAI's claim. It shows how far time to first answer moves with the reasoning setting, which is why a percentage at Instant says nothing about Extra High, and nothing about your latency.
Independent evidence: none yet for the October versions
Artificial Analysis and Vals AI publish their own runs, with cost beside the score. Both ran OpenAI's API models at maximum reasoning (Artificial Analysis labels each model max; Vals states max on each page, with OpenAI as the default provider). The pages fetched list no GPT-6 Sol or Luna labelled October and no entry for chat-latest. Artificial Analysis marks its September Sol page as deprecated and points to GPT-6.1 Sol. Both sites were fetched between about 13:42 and 13:44 BST on 8 October. Vals dated its CyberBench ranking 1 October, six days before the October versions.
Independent scores for OpenAI's API models, with cost per task, read on 8 October 2026. Both sites ran the models at maximum reasoning effort. None is a measure of the October chat versions or of cost per task on your workload.
- Source and measure
- Artificial Analysis Intelligence Index
- Model
- GPT-6 Sol (max)
- Score
- 48
- Cost per task
- $1.04
- Source and measure
- Artificial Analysis Intelligence Index
- Model
- GPT-6 Luna (max)
- Score
- 38
- Cost per task
- $0.07
- Source and measure
- Artificial Analysis Intelligence Index
- Model
- GPT-6.1 Sol (max)
- Score
- 52
- Cost per task
- $0.72
- Source and measure
- Vals Index
- Model
- GPT-6 Sol
- Score
- 57.54%, plus or minus 1.01
- Cost per task
- $7.581 per test
- Source and measure
- Vals Index
- Model
- GPT-6 Luna
- Score
- 51.22%, plus or minus 1.07
- Cost per task
- $0.431 per test
- Source and measure
- Vals CyberBench v1.1
- Model
- GPT-6 Sol, GPT-6 Luna
- Score
- 77.98%, 76.25%
- Cost per task
- $1.34, $0.13 per test
- Source and measure
- Vals CyberBench v1.1
- Model
- GPT-6.1 Sol, GPT-6 Astra
- Score
- 39.29%, 41.07%
- Cost per task
- $0.29, $1.84 per test
| Source and measure | Model | Score | Cost per task |
|---|---|---|---|
| Artificial Analysis Intelligence Index | GPT-6 Sol (max) | 48 | $1.04 |
| Artificial Analysis Intelligence Index | GPT-6 Luna (max) | 38 | $0.07 |
| Artificial Analysis Intelligence Index | GPT-6.1 Sol (max) | 52 | $0.72 |
| Vals Index | GPT-6 Sol | 57.54%, plus or minus 1.01 | $7.581 per test |
| Vals Index | GPT-6 Luna | 51.22%, plus or minus 1.07 | $0.431 per test |
| Vals CyberBench v1.1 | GPT-6 Sol, GPT-6 Luna | 77.98%, 76.25% | $1.34, $0.13 per test |
| Vals CyberBench v1.1 | GPT-6.1 Sol, GPT-6 Astra | 39.29%, 41.07% | $0.29, $1.84 per test |
Cost changes the reading (derived). On Vals, Luna costs 5.7% of Sol per test and scores 89% of Sol's score. On Artificial Analysis it costs 6.7% of Sol per task for 79% of the score. Luna is the model Free and Go accounts get, if the announcement rather than the help page is right. Independent numbers also move: Vals's own note of 22 September put GPT-6 Sol at 62.57% on its index, and its page now shows 57.54%.
CyberBench is a defender task, finding and patching memory-safety bugs in open-source projects, and Vals states that "provider refusals count as failures". The two Critical-rated models score lowest of the OpenAI models there, 39.29% and 41.07% against 77.98% for GPT-6 Sol, and their runs were shorter: about 3 minutes 42 seconds and 4 minutes 10 seconds against 12 minutes 14 seconds. The page text read does not say why. Stronger safeguards stopping defensive work would fit, but that is our inference, not Vals's finding. Vals lists Daybreak Blue separately (72.98%), so these entries are not the trusted-access route. A security team should ask about trusted access before assuming a Critical-rated model is the better tool.
UK: a vendor changed the model behind a name your policy may use
A UK organisation whose staff already use ChatGPT is, in the language of the DSIT Code of Practice for the Cyber Security of AI (published 31 January 2025, voluntary), likely to be a System Operator, because it only uses third-party AI components. The Code asks for things that OpenAI's pages do not show for this change. The provisions are quoted from the Code's page; the right-hand column is what we found.
DSIT Code of Practice for the Cyber Security of AI, voluntary, provisions on model change, against what OpenAI's pages show on 8 October 2026.
- Provision
- 7.3
- What it asks
- Developers and System Operators "shall re-run evaluations on released models that they intend on using".
- What the pages show
- No dated snapshot for the October versions. gpt-6-sol and gpt-6-luna kept their names through a 25 September fix to image encoding. chat-latest is "regularly updated".
- Provision
- 7.4 and 10.2.2
- What it asks
- System Operators tell their end-users before a model is updated, and of security-relevant updates. That needs notice first.
- What the pages show
- Card, announcement and Enterprise notes appeared on the day of rollout. The Enterprise notes say the release "does not introduce a new admin preview period".
- Provision
- 11.2 and 11.3
- What it asks
- Developers treat a major update as a new version and test it. They should support operators, for example with preview access and versioned APIs.
- What the pages show
- The card is a new evaluation, which is what 11.2 asks. A versioned identifier or a preview for October was not found.
| Provision | What it asks | What the pages show |
|---|---|---|
| 7.3 | Developers and System Operators "shall re-run evaluations on released models that they intend on using". | No dated snapshot for the October versions. gpt-6-sol and gpt-6-luna kept their names through a 25 September fix to image encoding. chat-latest is "regularly updated". |
| 7.4 and 10.2.2 | System Operators tell their end-users before a model is updated, and of security-relevant updates. That needs notice first. | Card, announcement and Enterprise notes appeared on the day of rollout. The Enterprise notes say the release "does not introduce a new admin preview period". |
| 11.2 and 11.3 | Developers treat a major update as a new version and test it. They should support operators, for example with preview access and versioned APIs. | The card is a new evaluation, which is what 11.2 asks. A versioned identifier or a preview for October was not found. |
One setting, two models. OpenAI says GPT-6 in Chat "uses your workspace's existing GPT-6 Sol access setting", that admins can turn it off, and that already enabled models "do not need to be enabled again for this Chat rollout". An administrator who enabled GPT-6 Sol for Work and Codex in September therefore receives the October chat model, where the rollout has reached the workspace, without a new decision. The pages do not say whether the chat and Work and Codex uses can be separated. In the Admin Console, Models > Test shows which models a member can use and the settings behind it; its description does not mention versions. For contrast, OpenAI lets Enterprise admins turn off in-app updates for the desktop app so that organisations can review releases first (6 August). The model behind a name has no stated equivalent.
Where a model swap belongs in an AI management system. In the site's ISO/IEC 42001 readiness tool, clause 6.3 covers planning of changes and clause 8.4 covers impact assessment when a system is introduced or materially changed. Those descriptions come from the site's own reference data. The standard itself is paywalled and was not read, so this briefing makes no claim about its wording.
What OpenAI's pages say for each tier on 8 October 2026: the chat model, and the default for using your content to train models. The data statements are from OpenAI's help article on model improvement. It speaks of "services for individuals" and does not list tiers by name, so treating Free, Go, Plus and Pro as individual accounts is our reading.
- Tier
- Free and Go
- Chat model, per OpenAI
- GPT-6 Luna (October) from 8 October; the model help page still says GPT-5.6 Luna
- Content used for training by default
- May be used unless the person opts out
- Tier
- Plus and Pro
- Chat model, per OpenAI
- GPT-6 Sol (October). The Pro option is Astra, rated Critical.
- Content used for training by default
- May be used unless the person opts out
- Tier
- Business
- Chat model, per OpenAI
- GPT-6 Sol (October). Admins may control which models members can use.
- Content used for training by default
- Not used by default
- Tier
- Enterprise
- Chat model, per OpenAI
- GPT-6 Sol (October) under the existing GPT-6 Sol access setting
- Content used for training by default
- Not used by default
| Tier | Chat model, per OpenAI | Content used for training by default |
|---|---|---|
| Free and Go | GPT-6 Luna (October) from 8 October; the model help page still says GPT-5.6 Luna | May be used unless the person opts out |
| Plus and Pro | GPT-6 Sol (October). The Pro option is Astra, rated Critical. | May be used unless the person opts out |
| Business | GPT-6 Sol (October). Admins may control which models members can use. | Not used by default |
| Enterprise | GPT-6 Sol (October) under the existing GPT-6 Sol access setting | Not used by default |
Shadow AI. A member of staff on a personal Free account gets Luna, rated High for cyber in either version, on individual terms. OpenAI's 5 October post says a visual ad format will be tested during image generation for Free and Go users in the US later this month, and that advertising "does not influence the answers"; it says nothing about the UK, and our earlier briefing covers the ads. The point for a policy is that the same name now spans a consumer account and an Enterprise tenant with different data terms.
Region. The announcement says the rollout is global, and neither it, the card nor the Intelligent UI help article lists a country exclusion. OpenAI does state UK exclusions where they exist: dots on Pro exclude the UK, and public publishing for Sites was not available in the UK at launch. The only model page that names the UK is the GPT-5.6 section of the model help page. Silence is not a statement of availability for a UK tenant, so check the model menu. OpenAI's 5 October post says an invisible text watermark will be added over the coming weeks to eligible ChatGPT and Codex text in the EU, with no UK mention; whether GPT-6 chat text counts as eligible is not stated (earlier briefing).
What to do, in order
Take this with you
Actions, in order
- Find which tier and which model your staff reach today. Ask for a screenshot of the model menu from a Free, a Plus and an Enterprise seat. The menu says GPT-6, not GPT-6 Sol, so record it as GPT-6 (Sol, October, per OpenAI) with today's date.
- Check the admin setting that controls rollout. In Enterprise it is the existing GPT-6 Sol access setting, and OpenAI says admins can turn it off. Find out whether it also governs Work and Codex, check the separate Astra control, and use Models > Test in the Admin Console on a sample of members.
- Update your AI acceptable use policy so it names a capability and a rating, not a model name. For example: chat assistants whose vendor rates them High for cyber, with no Critical-rated model used without written approval.
- Re-run your own evaluation before relying on the new default: your prompts, your documents, your refusals. Record the version label and the date, and repeat it when OpenAI's pages change. The DSIT Code, provision 7.3, asks operators to re-run evaluations on released models they intend to use.
- Treat interactive answers as untrusted output wherever you embed, capture or archive them. Apply the three controls in PIP-014: no automatic loading of model-authored references, a content security policy on the rendering surface, and validation of model output as with any external input.
- Decide whether staff may use connected plugins such as mail, finance and health alongside interactive answers until OpenAI says what a component can do.
- Tell security staff that the rating is High and what follows: expect blocks on scaled vulnerability research and exploit chaining for accounts outside trusted access, apply to Trusted Access for Cyber if the team needs it, and note that the Pro option and GPT-6.1 Sol are Critical-rated models with a heavier safeguard stack.
- Ask OpenAI in writing for what its pages do not give: a dated identifier for the October versions, which API identifier maps to which version, a prompt injection figure for Intelligent UI, and what an interactive component can do.
- Put a date in the diary to re-read the model help page once the Free and Go rollout lands.
The question the card cannot answer for you
OpenAI has been open that the October and September models differ. Its card names the split, its release notes call the Work and Codex models "separate", and its Enterprise page says the access setting is shared. What the pages do not give is a way to tell, from a user's seat, which one answered. The card is a careful document, and it tells you in its own words where its numbers stop: lowest setting for safety, maximum for capability, simulated tasks, no system controls. Those limits are the part to carry into your own records.
If OpenAI changed the model behind a name in your acceptable use policy on a Wednesday afternoon, which of your own records would show it by Friday?
Key facts
Sources
- PrimaryGPT-6 and Intelligent UI for everyone, 7 October 2026, read in full in a browser: the 1.2 billion weekly users, Intelligent UI, the 44% claim and its neighbouring sentences, the safety paragraph, the availability section (Plus, Pro, Business and Enterprise today, Free and Go tomorrow, Sol and Luna, Work and Codex unchanged). The page uses none of the words High, Critical or PreparednessOpenAIaccessed 2026-10-08
- PrimaryGPT-6 Sol and GPT-6 Luna: October 2026 update, system card published 7 October 2026, read in full as HTML with all seven figures viewed: the October and September versions, the High rating in cyber and biology, the lowest and maximum reasoning settings, Table 4 prompt injection, the cyber capability results, the safeguards section. The PDF was not downloadedOpenAIaccessed 2026-10-08
- PrimaryGPT-6 Astra system card, with change log and the 22 September appendix for GPT-6 Sol and GPT-6 Luna: the Critical rating for Astra, the September Sol and Luna cyber results, static jailbreak and prompt injection tables, Figure 59 (99.050% and 98.605%), the Gray Swan result, Trusted Access for CyberOpenAIaccessed 2026-10-08
- PrimaryAddendum to the GPT-6 Astra system card: GPT-6.1 Sol, 29 September 2026: the Critical cyber rating, the 21.5% recent-flaws result, the Astra safeguards stackOpenAIaccessed 2026-10-08
- PrimaryGPT-5.6 system card, 9 July 2026: the High rating, the GPT-Red prompt injection method and attack success rates, the monitor design and activation classifiers, actor-level enforcement, Trusted Access for CyberOpenAIaccessed 2026-10-08
- PrimaryGPT-5.6 August Updates card, 6 August 2026: the August release in ChatGPT, the same High designations, the lowest-setting rule for safety evaluationsOpenAIaccessed 2026-10-08
- PrimaryDeployment Safety Hub index, read 8 October 2026: the list of cards and dates, and the absence of a separate September card for GPT-6 Sol and LunaOpenAIaccessed 2026-10-08
- PrimaryOur updated Preparedness Framework, 15 April 2025: the page for version 2 of the framework that the October card links. The PDF was not openedOpenAIaccessed 2026-10-08
- PrimaryPacing model development in an era of cyber-critical capabilities, 18 August 2026: the statement that OpenAI needs an approach beyond the current Preparedness Framework and will evolve itOpenAIaccessed 2026-10-08
- PrimaryOpenAI news index, read 8 October 2026, for the posts of 2 to 7 October, including the 5 October posts on ads and EU text provenanceOpenAIaccessed 2026-10-08
- PrimaryA model guide for the GPT-6 family, 2 October 2026: the three API models it lists (Astra, GPT-6.1 Sol, Luna), no October version, and the advice to measure cost per successful taskOpenAIaccessed 2026-10-08
- PrimaryIntelligent UI in ChatGPT, updated about 19 hours before it was read: Sol for paid tiers and Luna for Free and Go, Chat only, no Work or Voice, the Layout and visuals setting, state kept across a refresh, plugins continue to workOpenAI Help Centeraccessed 2026-10-08
- PrimaryGPT-6 and other models in ChatGPT, updated about 18 hours before it was read: GPT-6 not labelled Sol, Latest selections move to GPT-6, GPT-6 Pro on Astra, Free and Go described as on GPT-5.6 Luna, GPT-6.1 Sol separate from Chat, no mention of any ratingOpenAI Help Centeraccessed 2026-10-08
- PrimaryChatGPT Enterprise and Edu models and limits, updated about 18 hours before it was read: the existing GPT-6 Sol access setting, models already enabled need no re-enabling, Astra off by default, Models > TestOpenAI Help Centeraccessed 2026-10-08
- PrimaryChatGPT Enterprise and Edu release notes, read in full: 7 October entry (no new admin preview period), 29 September GPT-6.1 Sol entry, 6 August desktop app update controls, earlier early-access practiceOpenAI Help Centeraccessed 2026-10-08
- PrimaryChatGPT release notes, read in full: 7 October, 22 September (Sol and Luna in Work and Codex are separate from Chat), 3 September (Astra), 6 August (GPT-5.6 update in chat), earlier interactive featuresOpenAI Help Centeraccessed 2026-10-08
- PrimaryChatGPT Work and Codex, updated 7 days before it was read: GPT-6.1 Sol, GPT-6 Sol and GPT-6 Luna described as not available in regular chat conversations, a statement superseded by 7 OctoberOpenAI Help Centeraccessed 2026-10-08
- PrimaryHow your data is used to improve model performance, updated 10 days before it was read: services for individuals may use content to train models unless opted out; Business, Enterprise, Edu and API are not used by defaultOpenAI Help Centeraccessed 2026-10-08
- PrimaryAPI model catalog, read 8 October 2026: the flagship list (Astra, GPT-6.1 Sol, Luna), GPT-6 Sol under More models, Chat Latest under ChatGPT models, Daybreak modelsOpenAIaccessed 2026-10-08
- PrimaryGPT-6 Sol API model page: single identifier gpt-6-sol with no dated snapshot, a pointer to GPT-6.1 Sol as the newer Sol, April 2026 knowledge cutoff, $2 and $10 per million tokensOpenAIaccessed 2026-10-08
- PrimaryChat Latest API model page: points to the latest model in ChatGPT, snapshot regularly updated, $5 and $30 per million tokens, August 2025 knowledge cutoff, which do not match the gpt-6-sol pageOpenAIaccessed 2026-10-08
- PrimaryOpenAI API changelog, read in full: 7 October chat-latest update, 25 September image-encoding fix to gpt-6-sol and gpt-6-luna, 22 September launch, 29 September GPT-6.1 Sol, 3 September Astra, 7 August DaybreakOpenAIaccessed 2026-10-08
- PrimaryBuilding advertising for the way people use AI, 5 October 2026: a visual ad format tested during image generation for Free and Go users in the US later in October, and the statement that advertising does not influence answersOpenAIaccessed 2026-10-08
- PrimaryOur approach to EU text provenance rules, 5 October 2026: an invisible watermark to be added over the coming weeks to eligible ChatGPT and Codex text in the EUOpenAIaccessed 2026-10-08
- PrimaryGPT-6 Sol (max) page, fetched about 13:42 BST on 8 October 2026: Intelligence Index 48, $1.04 per task, time to first token 137.69 seconds against a median of 3.76, marked deprecated in favour of GPT-6.1 SolArtificial Analysisaccessed 2026-10-08
- PrimaryGPT-6 Luna (max) page, fetched 8 October 2026: Intelligence Index 38, $0.07 per taskArtificial Analysisaccessed 2026-10-08
- PrimaryGPT-6.1 Sol (max) page, fetched 8 October 2026: Intelligence Index 52, $0.72 per taskArtificial Analysisaccessed 2026-10-08
- PrimaryGPT-6 Sol page, fetched about 13:42 BST on 8 October 2026: Vals Index 57.54% plus or minus 1.01 at $7.581 per test, reasoning effort max, the 22 September note that gave 62.57%Vals AIaccessed 2026-10-08
- PrimaryGPT-6 Luna page, fetched 8 October 2026: Vals Index 51.22% plus or minus 1.07 at $0.431 per test, CyberBench 76.25%Vals AIaccessed 2026-10-08
- PrimaryCyberBench v1.1 leaderboard, updated 1 October 2026: GPT-6 Sol 77.98% at $1.34, GPT-6 Luna 76.25% at $0.13, GPT-6.1 Sol 39.29%, GPT-6 Astra 41.07%, Daybreak Blue 72.98%, and the statement that provider refusals count as failuresVals AIaccessed 2026-10-08
- PrimaryCode of Practice for the Cyber Security of AI, published 31 January 2025, read in full: the System Operator definition and provisions 7.3, 7.4, 10.2.2, 11.2 and 11.3. A voluntary codeUK Department for Science, Innovation and Technologyaccessed 2026-10-08
- PrimaryManaging the cyber risk of agentic AI, 20 August 2026, read in full: built-in model safeguards should not be treated as holisticNational Cyber Security Centreaccessed 2026-10-08
- PrimaryPrompt injection is not SQL injection (it may be worse), 8 December 2025, read in full: prompt injection may never be totally mitigated and will remain a residual riskNational Cyber Security Centreaccessed 2026-10-08


