How OpenAI's GPT-6 Sol and Luna Cut Pricing in Half While Reaching Astra-Level Reliability (2026)
Written by
Aerin Kim

OpenAI shipped GPT-6 Sol and GPT-6 Luna on September 22, cutting pricing roughly in half and closing much of the reliability gap to GPT-6 Astra. Here is what actually changed.
On September 22, 2026, OpenAI quietly did something that matters more for anyone actually running AI in production than any single flagship launch does: it shipped two cheaper models, GPT-6 Sol and GPT-6 Luna, and cut their prices roughly in half compared to the equivalent GPT-5.6 tier, while claiming the new, cheaper models make about half as many mistakes as their predecessors [1] [2] [3].
That is a genuinely unusual combination. Normally a price cut and a quality improvement pull in opposite directions, cheaper models are supposed to be worse, that is the entire reason they are cheaper. GPT-6 Sol and Luna are the latest entry in a pattern that has become the actual economic engine of the frontier AI industry in 2026: a lab ships an expensive flagship model, in this case GPT-6 Astra on September 3, and then within weeks pushes a meaningful share of that flagship's quality gains down into cheaper, faster, high-volume tiers. If you are running an agent pipeline, a bulk content system, a RAG stack, or anything that calls a language model thousands or millions of times a day, this second act, the cost-down act, is usually the release that actually changes your unit economics, not the flagship headline that gets all the press coverage.
This post is a technical walkthrough of what Sol and Luna actually are, why a cheaper model can inherit a flagship's reliability gains at all (a real, well-studied mechanism in machine learning, not marketing magic), the real pricing and benchmark numbers, and how the two new models stack up against GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5 for anyone deciding where to route production traffic.

Step 1: What Actually Shipped on September 22
GPT-6 Sol and GPT-6 Luna are the two lower-cost members of OpenAI's GPT-6 family, sitting beneath GPT-6 Astra, the flagship model OpenAI shipped on September 3, 2026 and the first OpenAI model to cross a Critical threshold on the company's cybersecurity Preparedness Framework evaluation [1]. Astra itself is covered in detail in our earlier breakdown of its benchmarks and safety classification, and this post assumes you already know roughly what Astra is rather than re-explaining it, the short version is that Astra is OpenAI's most capable and most expensive model, priced at $10 per million input tokens and $50 per million output tokens in standard mode.
Sol and Luna are not Astra with a smaller price tag slapped on. They are positioned for genuinely different jobs. GPT-6 Sol is aimed at complex tasks, coding specifically, and agentic or interactive coding workflows that require multistep validation, the kind of work where a model has to write code, run it, check the result, and correct course before finishing [2]. GPT-6 Luna sits below Sol as the lightweight, lowest-cost model in the family, built for high-volume tasks with a single clear goal, summarizing a document, extracting a handful of fields from unstructured text, or answering a quick factual question, rather than open-ended reasoning [2].
That three-tier shape, a flagship for the hardest and highest-stakes work, a mid-tier for demanding-but-routine agentic tasks, and a lightweight tier for high-volume simple tasks, is not unique to OpenAI. It mirrors the structure Anthropic already runs with Claude Opus 5, Claude Sonnet 5, and Claude Haiku 4.5, and it mirrors OpenAI's own prior generation, where GPT-5.6 Sol sat above a mid-tier Terra and a lightweight Luna. What is new here is not the shape of the lineup, it is how much of the flagship's actual quality improvement got pushed down into the cheap tiers this generation, and how quickly it happened, 19 days after Astra's own launch.
The headline claim: half the price, half the mistakes
The specific numbers OpenAI is putting behind this launch are concrete rather than vague marketing language. GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens, a 50 percent reduction from GPT-5.6 Sol's pricing [2]. GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens, down from $0.20 and $1.20 for GPT-5.6 Luna, also roughly a 50 percent cut [2]. OpenAI states plainly that GPT-6 Sol makes about half as many mistakes as its predecessor, and frames that improvement as reaching Astra-level reliability at a fraction of Astra's cost [1].
It is worth being precise about what "Astra-level reliability" means here versus what it does not mean. It does not mean Sol matches Astra's raw capability ceiling on the hardest reasoning or agentic tasks, Astra's own benchmark numbers on GPQA Diamond, ARC-AGI-3, and OSWorld 2.0 remain unmatched by the cheaper tiers, and OpenAI has not claimed otherwise. What OpenAI is claiming specifically is a reliability property, roughly, how often the model gets a task right versus makes a correctable or uncorrectable mistake, converging between the flagship and the cheap tier, even while the capability ceiling stays separated by price. That distinction between capability and reliability matters enormously for production use, because most real workloads are bottlenecked by reliability on tasks well within a model's capability range, not by whether the model can solve a genuinely frontier-hard problem. A support ticket triage system does not need ARC-AGI-3-level abstract reasoning. It needs to correctly classify tickets nineteen times out of twenty instead of seventeen times out of twenty, and that is exactly the kind of gain a reliability-focused distillation effort is built to deliver.

Astra's improvements carried down to both models
Beyond the headline error-rate claim, OpenAI's launch materials describe three specific quality improvements from Astra that carried down into both Sol and Luna: simplified language with less jargon, slightly shorter responses without losing substance, and lower rates of misleading claims about the model's own coding ability [2]. That last one deserves a second look, because it is a specific, checkable behavioral claim rather than a vague quality statement. A model that overstates its own coding correctness, telling a user a function is fixed when it is not, or that a test suite passes when it was never actually run, creates a particular kind of production risk: the failure looks like success until someone checks. If Astra's training run genuinely reduced that specific failure mode and OpenAI successfully carried the fix down into Sol and Luna, that is a meaningfully different kind of improvement than a benchmark score going up, because it directly targets the gap between what a model reports and what actually happened, which is exactly the gap that erodes trust in an unattended coding agent.
The practical value of "less jargon, slightly shorter responses" is easy to undersell if you only think about it as a style preference. In a high-volume Luna-tier pipeline, summarization and field extraction, shorter, more direct output is not a cosmetic improvement, it is a direct cost reduction, since output tokens are priced separately and typically cost more per token than input tokens across every model discussed in this post. A model that says the same thing in 15 percent fewer tokens without losing the substance a downstream system depends on is a model that is meaningfully cheaper to run at scale, independent of any per-token price cut.
Step 2: Why a Cheaper Model Can Inherit a Flagship's Reliability Gains
This is the part of the story that most coverage of the launch skipped past, and it is worth slowing down on, because "the cheap model got better because the expensive model got better" is not obvious on its face. A smaller, cheaper model has fewer parameters and less compute available to it at inference time than a flagship. So how does it end up more reliable than its own same-tier predecessor, and closer to a much larger model's reliability, without becoming as expensive as that larger model?
OpenAI has not disclosed the specific architecture or training pipeline behind Sol and Luna, no confirmed parameter counts, no disclosed mixture-of-experts versus dense architecture decision, no named distillation method. That is a real gap in the public information, and this post is not going to fill it with invented specifics. What can be explained accurately is the general, well-established set of techniques the broader industry uses to accomplish exactly this kind of cost-down transfer, techniques that are publicly documented by multiple labs, including OpenAI itself in a different context, described below with a clear line drawn between "this is how the industry generally does this" and "this is what OpenAI specifically disclosed about Sol and Luna," which is nothing beyond the outcome.
Knowledge distillation: training a smaller model on a bigger one's behavior
The foundational technique here is knowledge distillation, formalized in a now-classic 2015 paper by Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, "Distilling the Knowledge in a Neural Network" [4]. The core idea is straightforward: instead of training a small model purely on raw labeled data, you train it to mimic the output distribution of a larger, already-trained "teacher" model. The small "student" model learns not just the right answer but the teacher's confidence and reasoning pattern across many examples, which turns out to transfer more of the teacher's actual behavior than training on hard labels alone.
For large language models specifically, this has evolved into a standard, publicly documented industry workflow rather than a research curiosity. OpenAI's own developer platform documents a distillation workflow for its API: tune a prompt against a larger, more capable model like GPT-4.1 until it performs well against your evaluation criteria, capture that model's real outputs across a representative set of inputs, and use those outputs as a supervised fine-tuning dataset to train a smaller, cheaper model to match the larger model's behavior on that specific task distribution [5]. That is a real, generally available technique OpenAI documents for its customers today, and it is a reasonable, publicly grounded illustration of the general category of technique labs use internally at a much larger scale to build their own cheaper-tier models, though OpenAI has not stated that this exact customer-facing workflow is how Sol or Luna themselves were built internally.

The reason this matters for the "half as many mistakes" claim specifically is that a flagship model like Astra generates a much higher-quality signal for a distillation dataset than the previous generation's flagship did. If Astra makes fewer factual errors, hedges appropriately more often instead of confidently guessing, and avoids the specific overclaiming-about-coding-ability failure mode described in Step 1, then a smaller model trained to imitate Astra's behavior on a wide range of tasks inherits some fraction of those same corrected tendencies, without needing Astra's full parameter count or inference cost to do it. The smaller model is not reasoning as deeply as Astra in an absolute sense, it is instead approximating a compressed version of Astra's behavior on the specific distribution of tasks it was trained against, which is exactly why the reliability gain can transfer even while the raw capability ceiling does not.
Chinchilla-style training efficiency: getting more out of the same or smaller model
A second, related mechanism worth understanding is compute-optimal training, popularized by DeepMind's 2022 Chinchilla paper, "Training Compute-Optimal Large Language Models," by Hoffmann, Borgeaud, and colleagues [6]. The paper's core finding, at the time a genuine surprise to the field, was that most large language models up to that point had been undertrained relative to their parameter count: model size and training data volume should scale roughly together for a fixed compute budget, and many contemporary models were too large for the amount of data they had actually been trained on. A smaller model trained on proportionally more high-quality data can match or beat a larger model trained the old, undertrained way, for the same or lower total training compute.
That finding is now a standard input into how every major lab plans a model family, not a one-off result from 2022. When a lab designs its next generation of models across multiple size tiers, it is choosing parameter count and training token count jointly, informed by scaling-law research like Chinchilla's, rather than treating the small tier as an afterthought scaled down casually from the flagship. Applied to a family like GPT-6, this general principle explains, at the level of an industry-wide technique rather than an OpenAI-disclosed specific, why a smaller-tier model released alongside or shortly after a new flagship can plausibly train on more, better, or more efficiently curated data than the equivalent tier from the prior generation did, closing part of the quality gap to the flagship without closing the price gap. Again, OpenAI has not disclosed Sol or Luna's parameter counts or training token counts, so this section describes the general mechanism the field uses, not a confirmed detail of this specific release.
Inference infrastructure: caching and serving get more efficient over time too
The third piece is not about training at all, it is about serving. OpenAI's own framing of this release explicitly credits "enhanced caching and inference capabilities" as part of what makes the new pricing and reliability gains possible together [1]. This is a genuinely distinct lever from model quality: even a fixed model becomes cheaper and, in a practical sense, more reliable to build on if the infrastructure serving it gets faster, cheaper, and more consistent, because inference infrastructure improvements reduce the incidence of timeouts, truncated responses under load, and other operational failure modes that look like model mistakes to an end user but are actually serving-layer problems. A lab that has spent a year improving its inference stack, better batching, more efficient key-value cache management, faster routing between model shards, can pass real cost reductions through to customers on every model it serves, not just the newest one, which is a large part of why an industry-wide pattern of falling per-token prices persists release after release even before accounting for any change in the models themselves.
The caching improvement OpenAI specifically calls out for Sol and Luna, a 90 percent discount on cached input-token reads plus faster agent responses, sits squarely in this category, and Step 6 below covers exactly how that discount changes the economics of a real agentic session.
Step 3: Benchmarks and the Reliability Claim, Read Carefully
Here is where honesty about what has and has not been published matters most. OpenAI has not released a public benchmark scorecard for Sol and Luna with named evaluation suites and numeric percentages the way it did for Astra's GPQA Diamond, ARC-AGI-3, and OSWorld 2.0 scores. The claims in circulation from the launch are qualitative and comparative rather than a full benchmark table: "about half as many mistakes" than GPT-5.6 Sol, reaching reliability OpenAI describes as "Astra-level" at a much lower cost, and a claim that both new models handle tasks "substantially better" than Claude's competing models on some of OpenAI's own internal measures [1].
That last claim is worth flagging clearly rather than repeating as neutral fact. It is OpenAI's own self-reported comparison, using OpenAI's own chosen evaluation methodology, against Claude Fable 5.1 and Claude Opus 5, and it has not been independently corroborated by a third-party benchmark run at the time of this post. That does not make the claim false, OpenAI generally does not survive getting caught fabricating a specific benchmark comparison, but the same caution this blog applies to every self-reported number applies here: treat it as OpenAI's position in a competitive comparison, not as an independently verified result, until a third-party evaluator publishes its own numbers on the same tasks.
| Model | Input $/1M | Output $/1M | Terminal-Bench 4.0 | Positioning |
|---|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 | Not disclosed | High-volume, single clear goal |
| GPT-5.6 Luna | $0.20 | $1.20 | Not disclosed | Prior lightweight tier |
| GPT-6 Sol | $2.00 | $10.00 | Not disclosed | Complex coding, multistep validation |
| GPT-5.6 Sol | $5.00 | $30.00 | 37.3% | Prior flagship-adjacent tier |
| GPT-6 Astra | $10.00 (standard) | $50.00 (standard) | Not published in these terms | Flagship, Critical cybersecurity tier |
| Claude Fable 5.1 | $10.00 | $50.00 | 55.8% | Long-horizon agentic flagship |
| Claude Opus 5 | $5.00 | $25.00 | 52.3% | Default high-capability tier |

A few things are worth reading directly out of that table rather than skimming past. First, notice how many cells say "Not disclosed." That is not sloppiness on this post's part, it reflects the actual state of public information about Sol and Luna specifically. Where this post can report a real number from an established, cited source, Astra's GPQA Diamond score, Fable 5.1 and Opus 5's Terminal-Bench 4.0 scores, it does. Where OpenAI has not published a comparable figure for Sol or Luna, the honest answer is that the number does not exist publicly yet, not that it can be estimated or approximated.
Second, the pricing comparison across all six models tells a cleaner story than the benchmark comparison does, precisely because pricing is a number every lab publishes without ambiguity. GPT-6 Luna, at $0.10 input and $0.50 output per million tokens, is now meaningfully cheaper than every other model in this table, undercutting even GPT-5.6 Luna's prior pricing by half. GPT-6 Sol at $2/$10 sits well below Claude Fable 5.1's $10/$50 and GPT-6 Astra's $10/$50, a five-fold gap on both sides of the ledger. That gap is the real headline of this release for anyone doing cost planning, arguably more concrete and more actionable than the mistake-rate claim, because it does not depend on trusting a single lab's internal evaluation methodology.
Third, the Terminal-Bench 4.0 numbers already established on this blog for Fable 5.1 (55.8 percent), Opus 5 (52.3 percent), and GPT-5.6 Sol (37.3 percent) are useful context even though Sol and Luna do not have a directly comparable published score on the same benchmark yet [7]. If OpenAI's "handles tasks substantially better" claim is measured on a different internal suite, which appears to be the case since no shared benchmark name has been published for the comparison, then a reader cannot actually verify where GPT-6 Sol lands on Terminal-Bench 4.0 specifically from the current public record. That gap in comparability is exactly the kind of nuance that gets lost when a launch's press coverage repeats a lab's own competitive claim without noting which benchmark, if any, it is actually measured against.
Step 4: Pricing Deep Dive and a Real Cost Calculator
The pricing story here is simple enough to state in one sentence and important enough to actually run the numbers on. GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens, GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens, and both figures represent roughly a 50 percent cut from their GPT-5.6 predecessors [2].
A 50 percent price cut sounds like a straightforward win, and for pure per-token cost it is. But the number that should actually change your architecture decisions is cost per completed task, not cost per token, since a model that is 50 percent cheaper per token but needs twice as many tokens or twice as many retries to complete a task correctly has not actually saved you anything. This is exactly why the "about half as many mistakes" reliability claim from Step 1 matters as much as the pricing cut on its own: if it holds up in your own workload, the real savings compound, half the per-token price and roughly half as many failed attempts that need a retry or a fallback to a more expensive model.
python/code def cost_per_completed_task(input_tokens: int, output_tokens: int, input_price: float, output_price: float, retry_rate: float) -> float: """Effective cost per successfully completed task, accounting for the share of requests that need a retry. A cheaper per-token price combined with a lower retry rate compounds into a bigger real saving than the token price cut alone suggests.""" base_cost = (input_tokens / 1_000_000) * input_price + (output_tokens / 1_000_000) * output_price expected_attempts = 1 + retry_rate return base_cost * expected_attempts # GPT-5.6 Luna: $0.20 input / $1.20 output, ~15% retry rate on a # document extraction task. old_cost = cost_per_completed_task( input_tokens=1_500, output_tokens=300, input_price=0.20, output_price=1.20, retry_rate=0.15, ) # GPT-6 Luna: $0.10 input / $0.50 output, retry rate assumed to roughly # halve if the reported mistake-rate improvement holds for this task type. new_cost = cost_per_completed_task( input_tokens=1_500, output_tokens=300, input_price=0.10, output_price=0.50, retry_rate=0.075, ) print(f"GPT-5.6 Luna effective cost per task: ${old_cost:.6f}") print(f"GPT-6 Luna effective cost per task: ${new_cost:.6f}") print(f"Effective savings: {(1 - new_cost / old_cost) * 100:.1f}%") # Note: the 7.5% assumed retry rate is illustrative, not an OpenAI- # published figure. Measure your own retry rate before trusting this.
Run the numbers in that snippet with realistic assumptions, a document extraction task that used to need one retry roughly 15 percent of the time on GPT-5.6 Luna, dropping to roughly 7 to 8 percent of the time on GPT-6 Luna if the halved mistake rate claim holds for your specific task type, and the effective cost per successfully completed task falls by more than the raw 50 percent token-price cut alone would suggest. That compounding effect, cheaper per token and fewer retries needed, is the actual mechanism behind why a cost-down release like this one changes production economics more than a single flagship launch usually does, since most flagship launches do not touch the tier that is actually running your highest-volume workloads.

Here is what an actual API call to each model looks like in practice, since the two models are meant for genuinely different task shapes and the code around them should reflect that:
python/code from openai import OpenAI client = OpenAI() def run_validated_fix(bug_report: str, test_command: str) -> str: """GPT-6 Sol is positioned for complex coding and agentic workflows that need multistep validation, so structure the call as a loop rather than a single request: propose a fix, run the real test command, and only stop once validation actually passes.""" messages = [ {"role": "user", "content": ( f"Bug report:\n{bug_report}\n\n" f"Propose a fix, then tell me the exact shell command to run " f"to validate it. The validation command must be: {test_command}" )} ] for attempt in range(3): response = client.responses.create( model="gpt-6-sol", input=messages, reasoning={"effort": "high"}, max_output_tokens=3000, ) proposal = response.output_text # In a real pipeline this next line actually executes test_command # against the proposed patch and captures pass/fail plus output. passed, log = execute_validation(test_command) if passed: return proposal messages.append({"role": "assistant", "content": proposal}) messages.append({"role": "user", "content": f"Validation failed:\n{log}\nTry again."}) raise RuntimeError("Sol could not produce a validated fix in 3 attempts") def execute_validation(cmd: str): # Placeholder: run the real test suite in your CI sandbox here. return False, "stub"
python/code from openai import OpenAI import json client = OpenAI() def extract_invoice_fields(document_text: str) -> dict: """GPT-6 Luna is built for high-volume tasks with one clear goal. A single-shot structured-output call with a tight schema, no loop, no multistep validation, matches Luna's actual design intent.""" response = client.responses.create( model="gpt-6-luna", input=( "Extract these fields from the invoice text as JSON: " "vendor_name, invoice_number, total_amount, due_date. " f"Invoice text:\n{document_text}" ), reasoning={"effort": "low"}, max_output_tokens=200, ) return json.loads(response.output_text) # Luna's low cost ($0.10 / $0.50 per million tokens) makes it realistic # to run this on thousands of documents a day without the per-token # price dominating the pipeline's total cost. result = extract_invoice_fields(open("invoice_042.txt").read()) print(result["vendor_name"], result["total_amount"])
A plain shell call works the same way if your pipeline is not Python-based:
bash/code curl https://api.openai.com/v1/responses \ -H "Authorization: Bearer $OPENAI_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-6-sol", "input": "Review this pull request diff for logic errors and suggest a fix with a regression test.", "reasoning": {"effort": "high"}, "max_output_tokens": 2500 }'
Notice the difference in how those two snippets are structured beyond just the model name. The Sol example is built around a multistep, self-validating loop, exactly the "agentic or interactive coding workflows requiring multistep validation" use case OpenAI positions it for [2], where the model's own output triggers a follow-up action (running a test, checking a result) before the task is considered done. The Luna example is a single-shot, structured-output call with a tight, explicit schema, matching its positioning for high-volume tasks with one clear goal. Routing a document extraction workload through Sol, or routing an open-ended multistep coding task through Luna, would be fighting each model's actual design intent rather than working with it, and Step 8 below covers a concrete routing pattern for exactly this decision.
Step 5: Availability, ChatGPT, Codex, and GitHub Copilot
Release-day availability for a new model tier is often messier than the headline announcement suggests, and it is worth knowing exactly where Sol and Luna actually work today rather than assuming blanket availability. Both models rolled out immediately in ChatGPT Work and in Codex for Plus, Pro, Business, Enterprise, and Edu subscribers [2]. Luna additionally reached Free and Go users in the ChatGPT desktop app, a notable decision that puts a meaningfully more reliable, cheaper model in front of OpenAI's largest, least-monetized user tier immediately rather than gating it behind a paid plan [2] [1].
One detail is easy to miss and matters for anyone whose product experience depends on default ChatGPT behavior: neither Sol nor Luna is available in default Chat mode at launch [2]. Both are scoped specifically to Work, Codex, and the API surfaces at launch, not the general consumer chat experience most casual users interact with daily. If your mental model of "what model does ChatGPT use" is based on the default chat window, this release does not change that experience yet, and building an assumption that it does is a planning mistake worth avoiding.
| Surface | GPT-6 Sol | GPT-6 Luna |
|---|---|---|
| ChatGPT Work | Plus, Pro, Business, Enterprise, Edu | Plus, Pro, Business, Enterprise, Edu |
| ChatGPT desktop app | Not available to Free/Go | Free and Go users |
| Default Chat mode | Not available at launch | Not available at launch |
| Codex | Available | Available |
| GitHub Copilot | Pro+, Max, Business, Enterprise | Pro, Pro+, Max, Business, Enterprise |

GitHub Copilot's own rollout, documented in GitHub's changelog for the same day, is more granular than the ChatGPT-side rollout and worth reading in full if your team ships code through Copilot rather than the raw OpenAI API [3]. Sol is available on Copilot Pro+, Max, Business, and Enterprise plans, while Luna is available more broadly, on Copilot Pro, Pro+, Max, Business, and Enterprise plans, both usable across VS Code, GitHub.com, JetBrains IDEs, Xcode, and GitHub's mobile apps, billed on a usage basis rather than a flat per-seat allowance [3]. That Sol-requires-Pro+-or-higher detail is a real gating decision worth knowing before you plan a team rollout around it, a developer on plain Copilot Pro gets Luna but not Sol, and GitHub's own changelog notes the rollout itself is gradual, so a specific IDE or plan combination may not show the new models on day one even when the changelog says it should [3].
Step 6: Prompt Caching, the 90 Percent Discount That Actually Changes Agent Economics
OpenAI's launch materials specifically call out improved prompt caching alongside Sol and Luna: a 90 percent discount on cached input-token reads, plus faster agent responses attributed to the same underlying infrastructure work [1]. This is not a new concept, prompt caching (reusing the computed representation of a repeated prefix, like a system prompt or a long reference document, instead of reprocessing it from scratch on every request) has been standard across major model providers for well over a year at this point, including Claude Fable 5.1's own 75 percent cache-read discount covered in our Fable 5.1 pricing breakdown. What makes a 90 percent discount specifically worth walking through in code is how sharply it compounds on the exact workload shape Sol is positioned for: long, multistep agentic sessions that repeatedly re-send the same system prompt, tool schema, and conversation prefix across dozens of tool calls.
python/code def agentic_session_cost(fresh_input_tokens: int, cached_input_tokens: int, output_tokens: int, input_price: float, output_price: float, cache_discount: float = 0.90) -> float: """Cost of one Sol agentic session given the 90% cached-input discount OpenAI shipped alongside Sol and Luna. cached_input_tokens is every token read from a repeated system prompt or tool schema across the session's later turns.""" cached_price = input_price * (1 - cache_discount) return ( (fresh_input_tokens / 1_000_000) * input_price + (cached_input_tokens / 1_000_000) * cached_price + (output_tokens / 1_000_000) * output_price ) # A 12-turn agentic coding session: the system prompt and tool schema # (4,000 tokens) is sent fresh once, then re-read from cache on the # following 11 turns. with_cache = agentic_session_cost( fresh_input_tokens=4_000, cached_input_tokens=11 * 4_000, output_tokens=9_000, input_price=2.0, output_price=10.0, ) # The same session shape with no caching at all, every turn paying # full input price for the repeated system prompt and tool schema. no_cache = agentic_session_cost( fresh_input_tokens=12 * 4_000, cached_input_tokens=0, output_tokens=9_000, input_price=2.0, output_price=10.0, ) print(f"With 90% cache discount: ${with_cache:.4f}") print(f"Without caching: ${no_cache:.4f}") print(f"Saved: {(1 - with_cache / no_cache) * 100:.1f}%")
The mechanism behind why this discount matters more for Sol specifically than it would for a single-shot Luna-style call is structural. A long agentic coding session pays the full input-token price once, when it first sends a large system prompt and tool schema, and then, if the caching layer is hit correctly on every subsequent turn, pays only 10 percent of that price for the same content on every following turn in the session. The longer and more multistep the session, exactly the shape Sol is built for, the larger the share of total token spend that shifts from full-price input tokens to 90-percent-discounted cache reads. A short, single-turn Luna call, by contrast, rarely reuses enough repeated context within one request for the cache discount to move the needle much, which is one more reason the two models are genuinely built for different workload shapes rather than being the same model at two price points.

The practical takeaway for anyone architecting a Sol-based agent is to actually check whether your framework is structuring requests to hit this cache consistently. A system prompt or tool schema that gets subtly reformatted, reordered, or re-serialized between turns, even by a framework doing something as simple as re-sorting a dictionary of tool definitions, can silently break a cache hit and leave real savings on the table without any error being raised. This is the same caveat that showed up in our Astra coverage around its own cache-read discount, and it is worth treating as a standing operational check for any model in this generation rather than a one-time setup step.
Step 7: How Sol and Luna Compare to Claude Fable 5.1 and Claude Opus 5
Setting OpenAI's own self-reported "substantially better" claim aside, the practical comparison most teams actually need is where Sol and Luna land next to Anthropic's current lineup on cost and on the benchmark numbers that are actually publicly comparable.
On pure pricing, the comparison is stark and does not require any self-reported claim to evaluate. GPT-6 Sol at $2 input / $10 output per million tokens costs a fifth of Claude Fable 5.1's $10/$50 standard pricing, and GPT-6 Luna at $0.10/$0.50 undercuts both Fable 5.1 and Claude Opus 5's $5/$25 standard pricing by a wide margin [8]. That gap is real and independently verifiable from each company's own published rate card, unlike the benchmark comparison claim.
On agentic reliability specifically, the honest comparison is limited by what has actually been published. Fable 5.1 and Opus 5 both have a published Terminal-Bench 4.0 score, 55.8 percent and 52.3 percent respectively, giving a concrete, named-benchmark data point for Anthropic's lineup [7]. Sol and Luna do not have a published score on the same benchmark at the time of this post, which means a reader genuinely cannot verify OpenAI's comparative claim against Fable 5.1 and Opus 5 on a shared, named evaluation. What can be said accurately is that Sol is priced and positioned to compete with Fable 5.1 and Opus 5 on agentic coding tasks at a fraction of their cost, and whether it actually matches or beats them on your specific workload is a question only a real evaluation against your own tasks can answer, the same conclusion this blog has reached about every cross-lab benchmark comparison covered so far.

One genuinely interesting structural parallel is worth naming directly: both OpenAI and Anthropic are now running the identical playbook, ship a flagship (Astra in OpenAI's case, Fable 5.1 in Anthropic's), then push a meaningful share of that flagship's improvements down into a cheaper tier through some combination of distillation, more efficient training, and better serving infrastructure. Anthropic's own documentation explicitly tells developers to default to a cheaper model and escalate only when their own evaluations show it falling short, the same tiered-routing logic this post recommends for the OpenAI lineup in Step 8 below. That convergence is not a coincidence, it reflects the same underlying economic pressure: the overwhelming majority of real-world API traffic is not the hardest 1 percent of tasks a flagship model is built to solve, it is high-volume, well-scoped work that a well-distilled cheap tier can now handle at a fraction of the cost, and every major lab is now racing to make that cheap tier as reliable as possible rather than treating it as an afterthought.
Step 8: A Production Routing Pattern Across Astra, Sol, and Luna
Given three tiers with genuinely different cost, latency, and reliability profiles, the practical question for a real production system is not "which model is best," it is "which model should handle which slice of my traffic." Here is a routing pattern that reflects the positioning OpenAI itself describes for each tier.
python/code def route_task(task_type: str, requires_validation: bool, confidence_needed: str) -> dict: """A three-tier router matching OpenAI's own positioning: Luna for high-volume single-goal tasks, Sol for complex coding and agentic work needing multistep validation, Astra for the hardest tasks or anything that fails validation on the cheaper tiers.""" if task_type in ("summarize", "extract_fields", "quick_qa") and not requires_validation: return {"model": "gpt-6-luna", "effort": "low"} if task_type in ("coding", "agentic_workflow", "multistep_task") or requires_validation: return {"model": "gpt-6-sol", "effort": "high"} if confidence_needed == "maximum" or task_type in ("critical_security", "frontier_reasoning"): return {"model": "gpt-6-astra", "effort": "high"} return {"model": "gpt-6-sol", "effort": "medium"} def route_with_escalation(task_type: str, requires_validation: bool, confidence_needed: str, failed_lower_tier: bool = False) -> dict: choice = route_task(task_type, requires_validation, confidence_needed) if failed_lower_tier and choice["model"] == "gpt-6-luna": return {"model": "gpt-6-sol", "effort": "medium"} if failed_lower_tier and choice["model"] == "gpt-6-sol": return {"model": "gpt-6-astra", "effort": "high"} return choice print(route_task("extract_fields", False, "normal")) print(route_task("agentic_workflow", True, "normal")) print(route_with_escalation("extract_fields", False, "normal", failed_lower_tier=True))
A few things worth calling out about that router beyond the code itself. First, the escalation path matters as much as the initial routing decision. A task that Luna handles with low confidence, or that fails a validation check, should escalate to Sol rather than simply returning a low-quality answer, and a Sol task that fails its own multistep validation loop after a reasonable number of attempts is a legitimate candidate for escalation to Astra rather than an infinite retry loop on the mid-tier model. Second, this pattern only works if you are actually measuring failure rates per tier on your own tasks rather than trusting the published mistake-rate claims as a blanket guarantee, since "about half as many mistakes" is an aggregate claim across OpenAI's own evaluation mix, not a guarantee for any single task type.

This exact tiering problem, matching model cost and reliability to the actual difficulty of each step in a pipeline, is the same challenge that shows up in high-volume content generation systems, not just coding agents. A script-generation step in an automated video pipeline benefits from a model with real narrative reasoning, the kind of task where Sol's coding-and-validation strengths translate reasonably well to structured creative planning, while a metadata-extraction or title-variant step is a textbook Luna-shaped task, well-defined, high-volume, with a single clear goal. That is a large part of why the falling cost of reliable, cheap-tier models matters for tools built around automated content pipelines, like Miraflow AI's Text2Shorts, which turns a topic into a script and then into scene-by-scene visual prompts in sequence. A pipeline like that runs multiple model calls per finished video, and every dollar and every percentage point of reliability that a cheaper tier gains without sacrificing the parts of the workflow that genuinely need deeper reasoning translates directly into a cheaper, more consistent end-to-end product, whether that pipeline is calling GPT-6 Luna, Claude Haiku, or any other model in this same falling-cost tier.
A slow, steady camera push-in on a clean vector-style animated refinery diagram, matching the labeled cross-section illustration style used earlier in this piece. A glowing stream labeled 'Flagship model' flows down through a distillation tower, splitting at two labeled outlets into a medium amber stream labeled 'Sol' and a thinner, faster pale stream labeled 'Luna'. As each stream exits, a small price tag icon beside it shrinks to show a lower cost, and a small counter of hairline crack marks beside each stream visibly drops by half. Smooth, continuous motion, flat color fills with no photographic texture, clean explainer motion-graphics style, no readable paragraph text beyond the short labels already described, no logos, no people, 10 second loop, steady even lighting throughout.
Common Mistakes Teams Make With a Release Like This
A handful of avoidable mistakes show up reliably whenever a cost-down model release like this one lands, and they are worth naming explicitly.
- Treating "Astra-level reliability" as "Astra-level capability." OpenAI's claim is about mistake rate and reliability, not about matching Astra's raw benchmark ceiling on the hardest reasoning, computer-use, or cybersecurity tasks. Routing a genuinely frontier-hard problem to Sol because it is described as reaching Astra-level reliability is a category error.
- Comparing OpenAI's self-reported "substantially better than Fable 5.1 and Opus" claim as if it were an independently verified benchmark result. No shared, named benchmark has been published for that specific comparison at the time of this post. Treat it as one lab's competitive framing until a third-party evaluator publishes a comparable number.
- Assuming Sol and Luna are available everywhere ChatGPT is used. Neither model is in default Chat mode at launch, and enterprise or Copilot access depends on specific plan tiers, Sol requires Copilot Pro+ or higher, while Luna is available starting at plain Copilot Pro.
- Routing a document-extraction or bulk classification workload through Sol out of caution. Sol is priced and positioned for multistep, validation-heavy coding and agentic work. Luna is specifically built and priced for the high-volume, single-goal task shape, and routing simple tasks to the more expensive tier wastes the entire point of this release's pricing structure.
- Not checking whether your framework's request structure actually hits the 90 percent cache discount. A subtly reformatted system prompt or reordered tool schema between turns can silently break a cache hit, leaving real savings unclaimed without any visible error.
- Migrating production traffic without re-measuring mistake rate on your own task distribution. The halved mistake-rate claim is an aggregate figure across OpenAI's own evaluation mix. Your specific workload, a particular document format, a specific coding language, a narrow extraction schema, may see a smaller or larger improvement, and the only way to know is to actually measure it.
Production Best Practices for Teams Adopting the Tiered GPT-6 Lineup
A few concrete habits are worth building into a rollout plan for this specific release rather than discovering them after a cost or quality surprise.
Instrument mistake rate per tier and per task type before and after migrating, not just aggregate cost. The entire value proposition of this release rests on a reliability claim, and reliability is only knowable from your own production data, not from a launch announcement.
Structure agentic sessions to maximize cache hits deliberately. Keep system prompts and tool schemas byte-identical across turns within a session wherever possible, since the 90 percent cache-read discount only applies to content the caching layer actually recognizes as a repeat.
Build an explicit escalation path from Luna to Sol to Astra, rather than three independent, disconnected integrations. A task that fails validation on a cheaper tier should have a defined, automatic path to the next tier up, with the failure logged so you can see how often escalation actually happens and whether your initial routing thresholds need adjusting.
Keep a small, held-out evaluation set specific to your own task distribution, and re-run it against each new tier release, GPT-6 Sol and Luna today, whatever ships next quarter. General industry benchmarks and a lab's own aggregate mistake-rate claims are directional evidence, not a substitute for measuring your actual workload.
Treat the GPT-6 Astra-to-Sol-to-Luna cost ladder and the Claude Opus-5-to-Fable-5.1 ladder as the same kind of decision, not two separate vendor evaluations. Since both labs are now running the same flagship-then-distill playbook, the routing discipline that makes one lineup cost-efficient, measure your own tasks, escalate only when a cheaper tier's evals genuinely fall short, applies identically to the other.
Conclusion
GPT-6 Sol and Luna are not, on their own, a headline-grabbing capability leap the way Astra was three weeks earlier. What they represent instead is arguably more consequential for anyone actually running AI at production volume: real, verifiable evidence that a frontier lab can push a meaningful share of its flagship's reliability gains down into a cheap, high-volume tier at roughly half the price of the prior generation's equivalent tier, through some combination of knowledge distillation, more compute-efficient training, and genuinely improved serving infrastructure. The pricing cut is unambiguous and independently verifiable, GPT-6 Sol at $2/$10 and GPT-6 Luna at $0.10/$0.50 per million tokens are real, published numbers. The reliability claim, half as many mistakes and Astra-level consistency, is real in the sense that OpenAI is staking its own credibility on it, but it remains a self-reported figure on an undisclosed evaluation mix until independent testing corroborates it on your own kind of workload. The comparison to Claude Fable 5.1 and Opus 5 tells the same story: unambiguous on price, genuinely uncertain on head-to-head capability until a shared, named benchmark exists for the comparison. Test both new tiers against your own tasks before rerouting anything that matters, and treat this release as one more data point in an industry-wide trend, not a one-off bargain, since the same flagship-to-distilled-tier pattern is very likely to repeat with the next major model family from every lab in this space.
Frequently Asked Questions
Is GPT-6 Sol actually as good as GPT-6 Astra? No, not on raw capability. OpenAI's claim is specifically about reliability, roughly how often the model completes a task correctly, converging toward Astra's level at a much lower cost, not that Sol matches Astra's benchmark ceiling on the hardest reasoning, computer-use, or cybersecurity evaluations.
Is the 50 percent price cut the same for both input and output tokens? Yes, as published. GPT-6 Sol's $2 input and $10 output pricing and GPT-6 Luna's $0.10 input and $0.50 output pricing are both described as roughly a 50 percent reduction from their GPT-5.6 equivalents.
Can I use GPT-6 Sol or Luna in regular ChatGPT right now? Not in default Chat mode. Both models are available in ChatGPT Work and Codex at launch, with Luna also reaching Free and Go users in the desktop app, but neither is in the standard consumer chat experience yet.
Is OpenAI's claim that Sol and Luna beat Claude Fable 5.1 and Opus 5 independently verified? No. That is OpenAI's own self-reported comparison on its own internal evaluation, and no shared, named benchmark score has been published publicly for that specific comparison at the time of this post. Treat it as one lab's competitive framing rather than a neutral, third-party result.
Do we know how OpenAI actually built Sol and Luna to be more reliable? No specific architecture or training details have been disclosed, no confirmed parameter counts, no named distillation method, no context window figure. This post explains the general industry techniques, knowledge distillation, compute-optimal training, and inference infrastructure improvements, that plausibly explain this kind of gain across the industry, without claiming OpenAI disclosed using any specific one of them for this release.
Should a high-volume document processing pipeline use Sol or Luna? Luna, generally. Sol is priced and positioned for complex, multistep coding and agentic validation work. Luna is specifically built for high-volume tasks with a single clear goal, like summarization and field extraction, and is the far cheaper option for that task shape.
References
- TechCrunch, "OpenAI launches GPT-6 Sol and Luna"
- MacRumors, "OpenAI's New GPT-6 Sol and Luna Models Bring Astra Improvements to Cheaper Tiers"
- GitHub Changelog, "OpenAI's GPT-6 Sol and GPT-6 Luna now available"
- Hinton, Vinyals, Dean, "Distilling the Knowledge in a Neural Network"
- OpenAI Developer Docs, "Supervised Fine-Tuning: Distilling From a Larger Model"
- Hoffmann, Borgeaud, et al., "Training Compute-Optimal Large Language Models" (Chinchilla)
- VentureBeat, "Anthropic's Claude Fable 5.1 and Mythos 5.1 Arrive With a 75% Cost Reduction for Fable Cache Reads"
- Anthropic, "Introducing Claude Fable 5.1 and Claude Mythos 5.1"
- Miraflow AI Blog, "GPT-6 Astra Explained: Inside OpenAI's First 'Critical'-Threshold Model"
- Miraflow AI Blog, "Claude Fable 5.1 Explained: Benchmarks, Pricing, and How It Compares"
- Miraflow AI Blog, "Claude Opus 5 vs Sonnet 5 Explained: Benchmarks, Pricing, and the Effort Toggle That Actually Matters"


