GPT-6 Sol and Luna: Route Traffic, Cut Spend 50%+

Your OpenAI bill went up again last month. You're paying flagship prices for tasks a cheaper model handles fine — email classification, invoice parsing, "does this ticket need a human?" — because rewiring the router feels like a Saturday you don't have. OpenAI just made that Saturday worth it: GPT-6 Sol (mid-tier) and GPT-6 Luna (low-tier) land at roughly half or less of the predecessor API prices for the same slot.
This post is the practical version of what to actually do about it. Not the launch summary. What to route where, how to A/B without shipping regressions, how to think about the eval before you flip traffic, and where the savings are real versus imaginary.
What Sol and Luna actually change
Two things matter for your stack: tier positioning and per-token price. Sol slots in as the mid-priced workhorse — think of it as the "default model" for structured extraction, classification, routing, summarization, and short reasoning. Luna is the low-tier, the model you throw at high-volume, low-ambiguity work: log tagging, first-pass moderation, keyword-style intent detection, deduplication.
The pricing headline — "50% or more cheaper than predecessors" — is the useful part. Even if you never touched a flagship model, if you've been running everything on the previous mid-tier for six months, moving the mid-tier eligible traffic to Sol is a straight halving of that line item. Everything else in this post assumes you actually verify current pricing on OpenAI's pricing page before you refactor — model prices shift, and the specific per-million-token numbers change more often than blog posts do.
What doesn't change: your prompts still need to be good, your evals still need to exist, and cheap tokens on a badly-routed task still lose you money in retries and human review.
The router is the product
Most teams don't have an "OpenAI bill problem." They have a routing problem. One model, one endpoint, every task — the flagship handles a 20-token classification with the same firepower as a 4,000-token synthesis. That's the leak.
A workable router has three layers:
- Task classification. What is this call actually doing? Extraction, classification, reasoning, generation, tool use.
- Tier assignment. Given the task class and stakes, which tier is the default?
- Escalation. When does a low-tier response trigger a re-run on a higher tier?
Here's the minimum viable shape in Python. Not fancy — this is what I actually ship for clients on day one:
from openai import OpenAI
from enum import Enum
client = OpenAI()
class Tier(Enum):
LOW = "gpt-6-luna"
MID = "gpt-6-sol"
HIGH = "gpt-6" # placeholder for whichever flagship you use
ROUTING = {
"classify_intent": Tier.LOW,
"extract_fields": Tier.LOW,
"route_email": Tier.MID,
"summarize_thread": Tier.MID,
"draft_reply": Tier.MID,
"resolve_dispute": Tier.HIGH,
"multi_step_reasoning": Tier.HIGH,
}
def call(task: str, messages: list, escalate_on=None):
tier = ROUTING[task]
resp = client.chat.completions.create(
model=tier.value,
messages=messages,
temperature=0,
)
content = resp.choices[0].message.content
if escalate_on and escalate_on(content):
resp = client.chat.completions.create(
model=Tier.HIGH.value,
messages=messages,
temperature=0,
)
content = resp.choices[0].message.content
return content
The escalate_on callback is where the real work is. For classification, escalate when confidence is below a threshold or when the model returns "unclear." For extraction, escalate when required fields come back null. For drafting, escalate when the output fails a downstream validator (schema, length, tone check).
The point isn't the code. The point is that "which model" is now a config decision, not a code decision, and you can move traffic between tiers without shipping.
Where Luna earns its keep
Luna is aggressively cheap, which makes it tempting to dump everything there. Don't. The tasks where Luna actually wins are narrow but common:
- Boolean and small-enum classification. "Is this email a receipt? yes/no." "Priority: low/med/high." These do not need a reasoning model.
- Field extraction from structured or semi-structured input. Invoice line items, form data, log entries with a known schema.
- Deduplication and clustering hints. "Are these two support tickets about the same issue?"
- First-pass moderation. Filtering the obvious 90% so the mid-tier only sees ambiguous edges.
- Embedding companions. Cheap enrichment on top of vector search — expanding a query, generating alt phrasings, tagging retrieved chunks.
Where Luna loses money even at half the price:
- Anything requiring multi-step reasoning across a document
- Tone-sensitive drafting (customer replies, marketing copy)
- Tool-use loops with more than 2-3 steps
- Anything where a wrong answer costs more than the token savings
A useful heuristic: if you can write the eval as "did it return exactly this value," Luna is a candidate. If the eval is "does a human think this reads well," it isn't.
Sol as the new default
For most SMB automation work, Sol becomes the new default model. Email routing, ticket triage, meeting-note summarization, CRM enrichment, first-draft replies with a human in the loop, RAG over a modest corpus — this is Sol territory.
The mental model I use with clients: Sol handles anything where the task has structure but the input doesn't. A customer email is unstructured. "Extract sender intent, urgency, and any mentioned order numbers, then route to a queue" is structured. Sol handles that all day.
Traffic mix in a typical SMB automation stack after tuning:
| Tier | Share of calls | Share of spend | Typical work |
|---|---|---|---|
| Luna | 60-75% | 15-25% | Classification, extraction, first-pass filters |
| Sol | 20-35% | 45-60% | Routing, summarization, drafting, RAG |
| Flagship | 3-8% | 20-35% | Reasoning, escalations, agent tool loops |
Those are ranges from real client workloads, not benchmarks. Your mix depends entirely on what your pipeline does. But if 100% of your calls are hitting one tier, you're leaving money on the table in one direction or the other.
Don't cut costs without an eval
This is the part every "route to a cheaper model" post skips. If you don't have an eval set, moving traffic to Luna is a coin flip — you might save 50% and lose 8% accuracy, and only find out three weeks later when a customer complains.
The minimum eval for a routing decision is 50-200 labeled examples per task, held constant across model swaps. Build it once, run it before every routing change.
import json
from collections import defaultdict
def run_eval(model: str, examples: list, task_fn):
results = defaultdict(int)
for ex in examples:
pred = task_fn(model, ex["input"])
results["total"] += 1
if pred == ex["expected"]:
results["correct"] += 1
else:
results["errors"] += 1
return {
"model": model,
"accuracy": results["correct"] / results["total"],
"n": results["total"],
}
with open("eval_set.jsonl") as f:
examples = [json.loads(l) for l in f]
for model in ["gpt-6-luna", "gpt-6-sol", "gpt-6"]:
print(run_eval(model, examples, classify_intent))
You are looking for the tier where accuracy plateaus. If Luna hits 94% and Sol hits 95%, Luna wins on cost. If Luna hits 82% and Sol hits 96%, Sol wins on total cost including error recovery. Cost per correct answer is the metric that matters, not cost per token.
Two rules I don't break:
- Eval before rollout, not after. Run the eval on shadow traffic or a labeled set before you flip production.
- Keep the eval set adversarial. Include the weird cases: empty inputs, non-English content, prompt injection attempts, edge formats. If Luna handles those, it earns production. If it doesn't, keep them on Sol.
The escalation pattern that pays for itself
The highest-leverage pattern with a multi-tier lineup isn't "use cheap model everywhere." It's cheap-first with escalation. Every request starts on Luna. If Luna's output fails a validator, it re-runs on Sol. If Sol fails, it escalates to flagship or a human queue.
For a task where 85% of inputs are easy and 15% are hard, the math works out like this:
- Old: 100% on Sol → cost = 100 × Sol_price
- New: 100% on Luna + 15% retry on Sol → cost = 100 × Luna_price + 15 × Sol_price
At Luna being roughly half of Sol, that's about 57.5% of the old bill, for the same accuracy — assuming Luna's easy-case output matches Sol's. That "assuming" is what the eval proves.
The validator is where engineers get lazy. A real validator checks:
- Schema compliance (JSON parses, required fields present, types correct)
- Semantic guards (extracted email is a plausible email, extracted date is a plausible date)
- Confidence signals (did the model return "unknown" or hedge?)
- Length and format (reply isn't 4,000 characters when the template calls for 3 sentences)
def validate_extraction(output: dict) -> bool:
required = ["sender_intent", "urgency", "order_ids"]
if not all(k in output for k in required):
return False
if output["urgency"] not in ("low", "med", "high"):
return False
if output["sender_intent"] == "unclear":
return False
return True
If validate_extraction returns false, escalate. That's the whole pattern. It looks trivial. It cuts bills in half.
What breaks when you rewire routing
Three things go wrong in production when teams move to multi-tier routing, in this order:
Latency variance grows. Different tiers have different response times. If a downstream system has a tight timeout, escalations to a slower flagship can start failing under load. Fix: set per-tier timeouts, and treat escalation-under-load as a metric you alarm on.
Prompts don't transfer cleanly. A prompt tuned for a flagship model often over-specifies for Sol and confuses Luna. Luna in particular does better with tighter, more structured prompts — fewer clauses, more explicit output format, examples that show exactly one shape. Re-tune per tier. It's boring work but it's the difference between Luna at 78% and Luna at 93%.
Tool-use loops degrade quietly. If you're running agent loops (function calling, MCP tools, retrieval steps), lower tiers can get stuck or invoke the wrong tool. This is where you don't route down — keep agent loops on the flagship or on Sol at the highest end, and put the cost savings elsewhere in the pipeline (the pre-processing, post-processing, and enrichment steps around the loop).
The OpenAI Cookbook has decent production patterns worth reading before you ship.
A concrete migration plan
If you're running everything on one model today, here's the 5-day plan I use:
Day 1: Inventory. List every OpenAI call in your codebase. For each: task description, average input tokens, average output tokens, monthly volume. Most teams have somewhere between 4 and 20 distinct call sites.
Day 2: Label. Classify each call site into: extraction, classification, routing, summarization, drafting, reasoning, agent-tool. Assign a proposed tier per the table above.
Day 3: Eval sets. Build a 50-example labeled set for each call site that gets a proposed downgrade. Pull from real production logs. Anonymize where required.
Day 4: Shadow test. Run the proposed cheaper tier in shadow mode — call both models, log both, compare. Don't ship yet. Look at cost per correct answer, not raw accuracy.
Day 5: Ship the winners. For every call site where the cheaper tier holds up, ship the change behind a feature flag with escalation-to-higher-tier on validator failure. Monitor for a week before killing the flag.
The savings you unlock in week one usually cover the engineering time in the same month, and every subsequent month is pure margin recovered.
How BizFlowAI approaches this
I do this cost review for clients on a regular cadence — usually monthly on stacks running more than a few thousand dollars a month in model spend. The work is boring on purpose: read every prompt, look at every call site, build eval sets from real traffic, shadow-test the cheaper tier, ship the router, keep the validator honest. The pattern I described above isn't theory; it's the same playbook I ran when the previous mid-tier launched, and I'll run it again for Sol and Luna.
If you want a look at your own OpenAI bill through this lens — what routes to Luna, what stays on Sol, where escalation actually pays off — book a discovery call. I'll walk your call sites with you and give you a written recommendation. No obligation to have me implement it.
What to do this week
Three actions, in order:
- Pull your last 30 days of OpenAI usage by model. If you can't break it out by call site, add basic metadata (a
metadatafield on each call) and give it a week. - Pick the highest-volume call site and ask: is this task Luna-eligible? If yes, build the eval and shadow-test.
- Add validators to your existing calls if you don't have them. Even without changing tiers, validators catch silent failures you're paying for today.
Model launches like Sol and Luna are only savings on paper. The savings become real when the router, the evals, and the validators exist. That's the whole job.
Work with BizFlowAI
If you'd rather have this built for you, that's what we do: production AI automation for solo founders and small teams — agents, integrations, and document pipelines that actually ship.
Book a free discovery call — 30 minutes, we map the highest-ROI automation in your workflow. No pitch deck, just engineering.
More guides like this on the BizFlowAI blog.
Frequently asked questions
What are GPT-6 Sol and GPT-6 Luna?
GPT-6 Sol and GPT-6 Luna are two new OpenAI models positioned as mid-tier and low-tier options in the GPT-6 lineup. Sol is the mid-priced workhorse aimed at structured extraction, classification, routing, summarization, and short reasoning tasks. Luna is the low-tier model built for high-volume, low-ambiguity work like log tagging, moderation, and deduplication. Both are priced roughly 50% or more below their predecessor tiers.
How do I cut OpenAI API costs with tiered model routing?
Instead of sending every request to one model, classify each task, assign a default tier, and add escalation rules. Send classification and extraction to Luna, routing and summarization to Sol, and reserve the flagship for complex reasoning. Use a validator to re-run failed cheap-tier calls on a stronger model. Typical result: 40-60% cost reduction with no accuracy loss when backed by an eval set.
When should I use GPT-6 Luna versus GPT-6 Sol?
Use Luna for boolean or small-enum classification, field extraction from structured input, deduplication, first-pass moderation, and query enrichment for RAG. Use Sol when the task has structure but the input does not — email routing, ticket triage, summarization, drafting, and RAG over modest corpora. A simple heuristic: if the eval is exact-match, Luna is a candidate; if it needs human judgment on quality, use Sol.
What is the cheap-first escalation pattern for LLMs?
Cheap-first escalation sends every request to the cheapest model first, then re-runs it on a stronger model only if a validator rejects the output. For example, run 100% on Luna and escalate the 15% that fail schema or confidence checks to Sol. This typically costs around 57% of a Sol-only setup at similar accuracy. The validator must check schema compliance, required fields, and confidence, not just presence of a response.
How do I evaluate whether a cheaper LLM is safe to deploy?
Build a labeled eval set of 50-200 examples per task and run it against each candidate model before changing production routing. Measure cost per correct answer, not cost per token, since errors trigger retries or human review. Include adversarial cases like empty inputs, non-English text, and prompt injection. Deploy the cheaper tier only when accuracy plateaus at an acceptable level.