OpenAI halves GPT-6 Sol and Luna prices. Five of its six charts compare against numbers it did not measure
OpenAI cut GPT-6 Sol and Luna API prices by half on 22 September and published six benchmark charts to support the move. Five of them carry a Claude series, and OpenAI ran none of those Claude numbers itself.
By Parminder Kumar Sharma · · 14 min read

The cut is real. One of the four numbers is bigger than the label beside it
GPT-6 Luna's output price fell from $1.20 to $0.50 per million tokens on 22 September 2026. That is 58.3 per cent off. OpenAI's own pricing table prints 50% cheaper in the cell beside it.
The other three reductions announced that day are exactly half: Sol input $4.00 to $2.00, Sol output $20.00 to $10.00, Luna input $0.20 to $0.10. One label is carried across two different arithmetic results, and on the row where it is wrong it is wrong in OpenAI's disfavour. This is a rounding convention rather than a claim, and it is the only figure in the announcement that understates itself.
What the cut does not establish is what anything now costs you. A price per million tokens is an input to a bill, not a bill. The number that decides whether this is good news for your organisation is cost per completed task, and that depends on how many tokens your workload actually spends, which depends in turn on the reasoning effort you select and on how much of your context is served from cache.
OpenAI knows this, which is why every chart in the announcement is plotted against cost per task rather than against price. It is the right axis. It is also an axis on which almost no organisation can locate itself, because almost nobody measures it.
GPT-6 API prices per million tokens, from OpenAI's launch post on 22 September 2026. Astra pricing from OpenAI's published model cards.
| Model | Input | Output | Stated positioning |
|---|---|---|---|
| GPT-6 Astra | $10.00 | $50.00 | Best results, unchanged by this announcement |
| GPT-6 Sol | $2.00 | $10.00 | Complex coding and professional work |
| GPT-6 Luna | $0.10 | $0.50 | Fast everyday work at scale |
Astra costs 100 times Luna on input and 100 times Luna on output. That spread is the announcement's actual subject. OpenAI's phrase for it is the cost to intelligence curve, and the six charts exist to argue that GPT-6 now occupies the good corner of it.
Five of the six charts carry a Claude series. OpenAI ran none of those Claude numbers
The launch post publishes six benchmark charts: AutomationBench, Agents' Last Exam, a factual error rate chart, FrontierCode, DeepSWE and OSWorld 2.0. Every one of them plots an outcome measure against cost per task, on a logarithmic cost axis. Five of the six include at least one Claude model.
At the foot of the page, after the charts and after the availability section, sits this:
Evaluations of competitor models were taken from publicly available reports. Scores for Claude Fable 5 were reported when scores for Claude Fable 5.1 were unavailable.
So in every comparison in the announcement, one side is a measurement OpenAI performed and the other side is a figure OpenAI read somewhere. That is disclosed, it is normal industry practice, and it is not the same thing as a head to head test.
The six charts in OpenAI's launch post, and what each one puts on the other side of the comparison. Compiled from the charts and their captions.
| Chart | Competitor shown | What OpenAI states about it |
|---|---|---|
| AutomationBench 1.0.6 | Claude Opus 5, Claude Fable 5.1 | Fable cost understated, omits Opus 5 fallbacks on about 40% of tasks |
| Agents' Last Exam V1 | Claude Opus 5, Claude Fable 5 with Opus 4.8 fallback | A previous Claude generation, per the sourcing footnote |
| Factual error rate | None | Five OpenAI models only, no competitor plotted |
| FrontierCode 1.1 Main | Claude Opus 5, Claude Fable 5.1 | No numeric comparison given in the prose |
| DeepSWE 1.1 | Claude Opus 5, Claude Fable 5 | Sol is 1.1 points behind Fable 5, not Fable 5.1 |
| OSWorld 2.0 offline | Claude Opus 5 only | No Fable series on the chart at all |
Two of those rows matter more than the rest. The closest coding comparison in the announcement, DeepSWE, puts GPT-6 Sol at 68.8 per cent within 1.1 percentage points of a Claude model, and that model is Fable 5, not the current Fable 5.1. Agents' Last Exam does the same. Where the current competitor's number was not available, the previous one was used, and OpenAI says so.
Note what the redraw makes visible that the original does not. OpenAI published one absolute cost, Sol's $0.27 per task. Every other cost in that chart reaches the reader as a multiple of Sol: 3.9 times, 8.9 times, 11.1 times. The chart has a dollar axis, so the values can be read off it, but the prose is written entirely in Sol-relative terms. The cheapest model is the unit of account.
Every headline comparison sets the effort dial differently on each side
Reasoning effort is a per request setting that trades tokens for score. It moves both axes of every chart in the announcement at once. So a comparison is only as meaningful as the effort settings it names, and the announcement names them.
The four numeric comparisons in OpenAI's launch post, with the effort setting stated for each side.
| Benchmark | GPT-6 and its effort | Competitor and its effort | Result |
|---|---|---|---|
| AutomationBench | Sol, xhigh | Claude Opus 5, max | 33.2% against 26.9% |
| Agents' Last Exam | Sol, max | Claude Opus 5, its highest score | 56.4%, 60% lower cost per task |
| DeepSWE 1.1 | Sol, max | Claude Fable 5, xhigh | 68.8% against 69.9% |
| OSWorld 2.0 | Sol, xhigh | Claude Opus 5, medium | 60.5% against 60.3% |
The OSWorld row is the one to read twice. GPT-6 Sol at its second highest effort matches Claude Opus 5 at its middle effort, 60.5 against 60.3, and OpenAI reports roughly 80 per cent lower cost per task. Both halves of that sentence are true and they are not the same claim. A near tie on score against a model that was not run at its maximum is a cost finding, not a capability finding, and OpenAI presents it as the cost finding it is.
The factuality chart is the only one with no competitor on it
The factual error rate chart plots five series: GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, GPT-5.6 Sol and GPT-5.6 Luna. No Claude, no Gemini, nothing outside OpenAI. Its y axis is labelled answers with any factual error, running from zero to 40 per cent.
The headline claim from it is that GPT-6 Sol makes about half as many mistakes as its predecessor, and that Luna at higher effort matches GPT-5.6 Sol at about a hundredth of the cost. Both are improvements against OpenAI's own prior models, which is what an internal evaluation can honestly demonstrate.
OpenAI adds one thing most vendors omit: the scores are not controlled for answer length, but verbosity sweeps showed almost no dependence on it. Anticipating the obvious confound and reporting the check is the behaviour you want from a model card, and it is worth saying so.
The cache is a thirty minute retention window that the launch post calls a discount
Alongside the price cut, OpenAI shipped an improved prompt caching system with higher cache hit rates by default and discounts of up to 90 per cent on cached input tokens. The launch post frames this entirely as savings. The companion engineering post, published the same day, carries the detail that a governance function actually needs:
We now give cache discounts for eligible shared prefixes reused within a 30-minute window.
That sentence is a retention statement written as a pricing rule. Your system prompt, your tool definitions and your accumulated conversation context are held server side and eligible for reuse for thirty minutes, by default, at a higher hit rate than before.
None of this is new in kind and none of it is hidden. Prompt caching has worked this way across the industry for two years, the window is documented, and OpenAI has now added a dashboard and a diagnostics endpoint so developers can see their own cache behaviour. The point is narrower and it is about paperwork: if your data protection impact assessment or your ISO/IEC 42001 records describe what leaves your estate when you call this API, the answer changed on 22 September, in the direction of more context retained more often, and the change was announced as a discount.
{
"prompt_cache_diagnostics": {
"type": "cache_miss",
"reason": "tools_changed",
"comparison_reusable_tokens": 5629,
"cache_missed_tokens": 5629
}
}
Two controls in the release are worth knowing about because they change agent design rather than agent cost. Reasoning effort can now be raised or lowered between responses without breaking cache, and tools can be enabled or disabled without breaking it, provided definitions and ordering stay stable. GitHub reports that caching work across the past several months cut the share of prompt tokens needing fresh processing by more than 50 per cent across billions of requests. That is a real engineering result, reported by the customer rather than the vendor.
Deception, measured as detected deception
The alignment section reports five evaluations: coding deception, broken search, reviewer bypass, warning circumvention and unauthorized interaction. Both new models improve over their GPT-5.6 counterparts, including on what OpenAI calls lower rates of misleading claims about their coding work.
The measurement definition is stated in the caption, and it is the sentence to keep:
Deception rate measures the fraction of answers with any detected deception. Effort was set to maximum.
Detected. Every number in that section is bounded above by what the graders caught, which makes it a floor on the true rate rather than an estimate of it. OpenAI also states that the evaluations deliberately test challenging situations and do not measure failure rates in typical use.
None of that is a criticism of the method. A deception rate can only ever be a detected deception rate, and saying so in the caption is better practice than the alternative. It matters because of where these numbers travel: an assurance questionnaire that asks whether your model provider tests for deceptive behaviour will be answered yes, with a percentage, and the percentage will be read as a rate rather than as a floor produced by adversarial prompts at maximum effort.
Today, two people in the same organisation are on different models
The rollout is deliberately uneven, and OpenAI says so. GPT-6 Sol and Luna are in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users. Free and Go users get Luna, in the desktop app. Neither model is in Chat yet. The API identifiers are gpt-6-sol and gpt-6-luna. OpenAI is rolling out gradually through the day to keep the service stable, and tells users who do not see the models to try later.
Availability as stated in OpenAI's launch post on 22 September 2026.
| Surface | Sol | Luna |
|---|---|---|
| ChatGPT Work and Codex, paid tiers | Available | Available |
| Desktop app, Free and Go | Not stated | Available |
| Chat | Not yet available | Not yet available |
| API | gpt-6-sol | gpt-6-luna |
For anyone operating an AI acceptable use policy, this is the practical consequence: for at least the length of the rollout, the question which model answered this is not answerable from the user's side, and the answer differs by plan, by surface and by hour. If your policy names approved models, it named a set that changed under it without a change request. If your AI inventory records model versions in production, the entry for ChatGPT is stale today.
What to check this week
Take this with you
Five checks, in the order worth doing
- Find out whether anyone in your organisation can state your cost per completed task for any AI workload. If nobody can, a 50 per cent price cut is not measurable as a saving and should not be recorded as one.
- Pull the model identifier out of your API logs for the last seven days and confirm which models your applications actually called. Pinned identifiers do not move on their own; defaults do.
- Open your AI system inventory and check the ChatGPT entry against the availability table above. Record that the model set changed on 22 September without a change request from you.
- Read the prompt caching guide before enabling explicit breakpoints, and record the thirty minute reuse window wherever your DPIA or AIMS documents what is retained outside your estate.
- If a supplier sends you a benchmark comparison this quarter, ask one question in writing: did you run both sides, and at what effort setting. Both halves matter and the second is usually missing.
The question the charts cannot answer for you
OpenAI has made a strong, well-documented case that GPT-6 Sol sits at a better point on the cost to intelligence curve than the models it is compared with. The charts are honest about their sourcing, the sample definitions are stated rather than buried, and the one cost figure that flatters OpenAI is flagged by OpenAI as incomplete. That is a higher standard of disclosure than most launches manage, and it deserves to be said plainly before the criticism.
The criticism is narrow. Five of the six charts compare a measurement with a citation. Two of them cite a previous generation of the competitor. Every headline comparison sets the effort dial differently on each side. None of that makes the conclusion wrong. It makes the conclusion OpenAI's reading of somebody else's published numbers, which is a different object from a test.
And the number that decides whether any of this matters to you is not on any of the six charts, because OpenAI does not have it.
Can anyone in your organisation say what a completed task costs you today? If not, the price of the model halving changes your invoice and nothing else you can prove.
Sources
- PrimaryIntroducing GPT-6 Sol and Luna: the launch post carrying the pricing table, the six benchmark charts, the availability list and the evaluation footnoteOpenAIaccessed 2026-09-23
- PrimaryBetter prompt caching for GPT-6: the 30 minute reuse window, the 90 per cent cached input discount, prewarming and the diagnostics payloadOpenAIaccessed 2026-09-23
- PrimaryIntroducing Claude Opus 5.5: used to cross-check the AutomationBench figures OpenAI attributes to Claude Opus 5 and Claude Fable 5.1Anthropicaccessed 2026-09-23


