Back to blog

GPT-6 Sol and Luna: How to Choose the Right Model and Control API Costs in 2026

Hello HaWkers, OpenAI introduced GPT-6 Sol and GPT-6 Luna on September 22, 2026, and made them available in the API as gpt-6-sol and gpt-6-luna. Their arrival after GPT-6 Astra raises a practical question for product builders: how much capability is worth paying for on each task? The official launch announcement says API prices are 50% lower than the promotional prices of the corresponding GPT-5.6 models.

Should you move every workflow to a new model, or reserve each option for the work it does best? In this guide, we will separate token prices from cost per task, build a reproducible evaluation method, and outline a simple model selection policy. You can then decide from your own product data and keep a clear way out if quality drops.

What Changed with Sol and Luna

The GPT-6 family now offers Astra, Sol, and Luna for different kinds of work. The official model documentation presents Astra as the highest capability option, Sol for demanding coding and agent workflows, and Luna for focused tasks at high volume. Their API identifiers matter: gpt-6-astra, gpt-6-sol, and gpt-6-luna. Use those identifiers in your configuration, and check availability in your account before planning a broad migration.

Those descriptions do not give any model exclusive ownership of a task. A short classification may need Sol when the ambiguous cases are expensive. A research workflow may use Luna to sort material and Sol to write the final synthesis. The best design depends on the cost of an error, how often the task runs, and how easily you can verify its answer. A polished response in one demonstration says little about the exceptions your customers actually send.

The API changelog records both releases on September 22. It lists text and image input, text output, and access through Responses and Chat Completions. For workflows that combine tools and reasoning, the model guide recommends Responses. That distinction matters for an agent that checks external systems: copying an old request without reviewing its parameters may create an integration that behaves differently from what you expected.

The blog has already covered the wider family in its article on GPT-6 Astra and robotics applications. Here the question is a product decision: when does better output justify a higher bill, and when is a less expensive option enough? The answer affects user experience, margins, and reliability as much as engineering.

Price per Token Is Not Cost per Task

At launch, OpenAI listed, for the standard tier up to 272,000 input tokens, $2 per million input tokens and $10 per million output tokens for Sol. Luna was listed at $0.10 for input and $0.50 for output. The changelog lists cached input at $0.20 for Sol and $0.01 for Luna. These are dollar prices for the documented category, not a promise that every request costs exactly that amount. The pricing page also distinguishes context size and processing mode. Check the current table before approving a budget.

Imagine a task with extensive instructions and a short answer. Input price carries more weight there. Another task might produce a long answer, making output price the dominant factor. Tools, extra attempts, long context, caching, and services used alongside the model call all change the total. That is why comparing only dollars per million tokens often leads to a bad decision: a cheap model that needs three attempts can cost more than a stronger one that succeeds on the first.

The function below estimates only the text portion using rates you provide. Its example rates come from the launch changelog for the pricing category just described. It does not add tool charges, long-context pricing, taxes, or other discounts. Passing prices explicitly makes it easier to update a spreadsheet or dashboard when the official table changes.

const rates = {
  sol: { input: 2, cachedInput: 0.20, output: 10 },
  luna: { input: 0.10, cachedInput: 0.01, output: 0.50 },
}

function textCost({ input, cachedInput = 0, output }, rate) {
  // Separate cached input to avoid counting it twice in the estimate.
  const freshInput = Math.max(0, input - cachedInput)
  return (
    freshInput * rate.input +
    cachedInput * rate.cachedInput +
    output * rate.output
  ) / 1_000_000
}

const usage = { input: 12_000, cachedInput: 8_000, output: 1_000 }
console.log({ sol: textCost(usage, rates.sol), luna: textCost(usage, rates.luna) })

This calculation is an estimate. The billed amount depends on actual usage, service tier, and the rules of the applicable price table. Record fresh input tokens, cached input tokens, output tokens, attempts, and tool costs for each completed task. That gives you the number your product needs: the cost of delivering an acceptable answer, rather than the cost of running a single call.

When to Choose Luna, Sol, or Astra

Start with the risk of the result. Luna is a natural candidate for classification, field extraction, short drafts, and other frequent tasks with an objective acceptance criterion. The official Luna description positions it for focused, high-volume work. Even so, a task that looks simple still needs evaluation. Routing a cancellation request to the wrong department can increase handling time and frustrate a customer, even if the classifier returns only one word.

Sol makes more sense when a task has several constraints, extensive context, tools, or decisions with greater consequences. The official guide recommends it for strong reasoning on demanding tasks. Contract review, a code fix, or a support plan based on a long customer history deserves a direct comparison with Luna. If the quality gap is small on your cases, the less expensive model may be sufficient. If Sol prevents rework or catches exceptions that matter, the extra cost may pay for itself.

According to OpenAI, Astra remains the reference point for the hardest work in the family. That does not mean it belongs on every complicated request. First define success: a correct answer, traceable source, valid format, safe action, and acceptable response time. Then compare the three models on examples from your own workflow.

An initial policy might send routine tasks to Luna, exceptions to Sol, and critical cases to Astra or a person. Treat that as an operating hypothesis and revise it when measurements disagree. Human review may still be necessary at any tier when material is sensitive. For an irreversible action, require confirmation from the responsible system before executing it, regardless of how confident the generated text sounds. Model routing cannot serve as permission to act.

Build an Evaluation That Reflects Your Product

A useful evaluation starts with real examples and appropriate permission to use them. Remove personal data where possible, include easy and difficult situations, and record the expected answer or a criterion that can be checked. Do not select only demonstrations that already went well. Rare failures deserve special attention when they can cause financial loss, incorrect publication, or poor service.

Separate work by type: extraction, decision, explanation, generation, and tool use. Within each group, measure correctness, format, time, attempts, and human intervention. One aggregate pass rate can hide a model that is good at drafts but weak at decisions. The official announcement reports its own benchmark results; these show general capability, but they cannot replace an evaluation using your data and rules.

Give every model the same inputs, instructions, and acceptance criteria. If you change the prompt for one and leave another untouched, you are comparing two different systems. Record the prompt version, model, reasoning effort, and test date. When the provider updates a service, you will be able to tell whether quality improved or whether previously solved cases regressed. Keep the failures too: they reveal where a routing rule, better context, or a human check may have more value than a model upgrade.

The next example summarizes results people have already reviewed. Set passed only after applying a criterion written in advance. The function does not ask another AI to judge an answer; it calculates metrics from a test set so your team can make a transparent decision.

function summarize(results) {
  const total = results.length
  if (total === 0) throw new Error("Include evaluated cases")
  const passed = results.filter(item => item.passed).length
  const cost = results.reduce((sum, item) => sum + item.costUsd, 0)
  const latency = results.reduce((sum, item) => sum + item.latencyMs, 0)
  return {
    passRate: passed / total,
    costPerAcceptedTask: passed ? cost / passed : Infinity,
    meanLatencyMs: latency / total,
  }
}

// Each row represents one full task, including all its attempts.
const lunaResults = [
  { passed: true, costUsd: 0.002, latencyMs: 800 },
  { passed: false, costUsd: 0.003, latencyMs: 900 },
]
console.log(summarize(lunaResults))

The numbers in this code are fictional data to demonstrate the calculation, not measured results from OpenAI or this blog. In practice, inspect the distribution too: an average latency may conceal very long delays on the cases that matter most. Read samples of failed answers and identify patterns before moving every request to another model. A well designed routing rule may address the problem at lower cost.

Design Routing and an Emergency Exit

After evaluation, turn the findings into a small rule that anyone on the team can audit. It should consider task type, risk, and the need for tools, rather than magic words in a user's request. Keep the logic deterministic where you can, so you can explain why a request went to a more expensive model.

function chooseModel(task) {
  // This is a product policy; adjust its criteria using your real tests.
  if (task.irreversibleAction || task.highImpact) return "gpt-6-astra"
  if (task.requiresTools || task.longContext || task.ambiguous) return "gpt-6-sol"
  return "gpt-6-luna"
}

const tasks = [
  { kind: "classify", highImpact: false },
  { kind: "research", requiresTools: true },
  { kind: "approve-transfer", irreversibleAction: true },
]
console.log(tasks.map(task => ({ kind: task.kind, model: chooseModel(task) })))

The return value of chooseModel is a routing proposal, not authorization to act. In the transfer example, a person or approval service must still confirm the effect. Separate model selection, text generation, result validation, and action execution. This division helps reduce errors and makes incident analysis easier.

Define an emergency exit as well. If an answer fails its schema, lacks a required source, or exceeds a time limit, try an appropriate model again or send the task for review. Record why the request took that path. Without this record, a system designed to be economical can gradually send nearly everything to its most expensive option without anyone noticing.

Fallback logic needs a limit. Repeating a bad request indefinitely spends money and may amplify an error. One additional attempt after correcting a format may be reasonable; repeated attempts without a diagnosis deserve investigation. Measure fallback rate together with cost per accepted task and user satisfaction. Cutting the bill while worsening the experience is only an apparent saving.

Include Cache, Context, and Tools in the Final Bill

The official announcement highlights cache improvements for conversations and agents. Reusing an instruction prefix can reduce reading cost and response time, but do not treat that as a guaranteed discount on every call. Cache hit rate depends on how you structure inputs. If variable data appears at the start of a prompt, even small changes may prevent the reuse you expected. Observe actual cached-input metrics before forecasting savings, and compare them across the kinds of tasks you run most often.

Long context deserves care too. Room for many tokens does not mean the system should send an entire history. Remove duplicates, retrieve only relevant documents, and summarize what can safely be summarized with verification. Larger inputs increase the chance of carrying stale or contradictory information. A quality problem may begin in context selection, not in the model. Inspect what the model was given before concluding that more model capability is the answer.

Tools change an agent's economics. A search, database call, or operation in another service may cost more than the tokens and add delay. If an agent calls the same tool several times to answer a simple question, revisit the workflow. Separate steps that require current information from those that can use data already available, and keep the origins of answers traceable.

In the API, review the parameters accepted by the model you select. The official guide says Sol and Luna support none reasoning effort, while Astra does not offer that level. It also recommends Responses for reasoning with tools. A safe migration tests format, tools, latency, and cost together. Merely changing the model string can alter the behavior of a critical step in a way that a superficial comparison misses.

Perspective: Model Choice Is Product Management

The launch of Sol and Luna provides more options, but it does not eliminate the need to measure. Price per token tells you what an input costs. Cost per accepted task tells you what it costs to deliver value. The difference includes failures, reviews, time, tools, and the damage from a wrong answer. When a team can see these parts separately, it can spend more capability where users actually benefit.

Start with a small set of representative cases, compare Luna and Sol, add Astra where risk or quality warrants it, and record the routing policy. Reevaluate after changes to your product or to the models. The best arrangement today may change as volume, rates, and request types change. For current availability and prices, the final source remains the official documentation at the time you buy.

Let's go! 🦅

📚 Want to Keep Up With What Is Coming?

This article covered GPT-6 Sol and Luna, but the ecosystem changes every week, and not everything becomes a post here.

On X, I share what I am testing, what happens behind the scenes in my projects, and news I find before it becomes an article.

Follow Me There

👉 Follow @jeffbruchado on X

💡 Daily content about development, careers, and the tools I actually use

Comments (0)

This article has no comments yet 😢. Be the first! 🚀🦅

Add comments