Ten cents. That’s what a million input tokens now costs on GPT-6 Luna, the smaller of two models OpenAI dropped on September 22, 2026. Output runs $0.50 per million. Its bigger sibling, GPT-6 Sol, sits at $2 in and $10 out. Both are half of what the GPT-5.6 versions charged.
I’ve been reviewing toolkits long enough to know that a 50% price cut usually comes with an asterisk somewhere. So far I haven’t found one in the spec sheet, which is genuinely interesting. Let’s go through what’s actually on the table and what I’d want to test before moving a production workload.
What’s confirmed
These are positioned as the cheaper, faster siblings to GPT-6 Astra, OpenAI’s flagship that shipped earlier in September. Sol and Luna are direct replacements for the GPT-5.6 lineup, not additions alongside it.
- GPT-6 Sol — $2 per million input tokens, $10 per million output
- GPT-6 Luna — $0.10 per million input, $0.50 per million output
- Context window — 1.05M tokens on both, with up to 922K usable as input and 128K max output
- Knowledge cutoff — April 20, 2026 for Sol, May 18, 2026 for Luna
That third bullet is the one I’d underline. Both models share the same context ceiling. The cheap one isn’t getting a smaller window as a consolation prize, which is how tiered model families usually work. If you’ve been paying flagship rates purely because your pipeline stuffs large documents into the prompt, Luna at a dime per million input tokens changes that math considerably.
The cutoff thing is backwards
Luna’s knowledge cutoff is May 18, 2026. Sol’s is April 20, 2026. The cheaper model knows about roughly a month more of the world than the pricier one.
I don’t have an explanation for that, and I’m not going to invent one. It probably reflects separate training runs rather than any deliberate tiering decision. But it matters practically: if your app asks questions about anything from late April or early May 2026, the twenty-times-cheaper model is the one with the fresher information. That’s an unusual sentence to write about a model lineup.
What I can’t tell you yet
Price and specs are the easy part. Here’s what the announcement doesn’t settle.
Quality per dollar. A 50% price cut is only a cut if output quality holds. Half price on a model that needs two attempts per task is not half price. I want to run the same eval suite across GPT-5.6 Sol and GPT-6 Sol before I call this a savings rather than a swap.
Where Luna actually breaks. Small models tend to fail in specific, predictable ways — instruction following under long context, structured output formatting, multi-step reasoning. A 922K input window is impressive on paper. The question is whether Luna reasons across the last 800K tokens as reliably as it handles the first 50K. Long-context benchmarks and long-context reality have historically been different animals.
The competitive claims. Some coverage frames this launch as undercutting Anthropic’s pricing. That may be true, but I haven’t verified the comparison numbers myself, so I’d rather you check current Claude pricing directly than take a headline’s word for it. Model pricing moves fast enough that any comparison written today has a short shelf life.
How I’d approach it
If you’re already running GPT-5.6 Sol, the migration is close to forced, since these are replacements rather than parallel options. Budget time for re-running your evals, not just a config change. Prompts tuned against a previous model version are not guaranteed to transfer cleanly, and “the spec sheet says it’s better” has burned me before.
If you’re on the flagship Astra for everything, this is a good moment to audit which calls actually need it. My rough heuristic: classification, extraction, routing, summarization, and tagging jobs rarely justify flagship pricing. Try Luna on those, measure the failure rate, and decide whether the error budget is acceptable. At ten cents per million input tokens, you can afford a verification pass on top and still come out ahead.
If you’re building something new, start on Luna and move up only where you see it fail. Starting cheap and escalating gives you a real map of where your task gets hard. Starting expensive teaches you nothing about your own cost floor.
The honest summary: the pricing is real, the shared context window is the most useful detail in the announcement, and the cutoff inversion is a small oddity worth remembering. Everything about quality is still an open question until somebody runs the tests. I’m running mine this week.
đź•’ Published: