Sol vs Terra vs Luna: GPT-5.6's 3 Models Explained

GPT-5.6's three models — Sol, Terra, Luna — are live in ChatGPT as of July 9. What each does, which plan gets which, and when ultra mode is worth it.

When OpenAI named its new GPT-5.6 models Sol, Terra, and Luna, a chunk of the internet had the same reaction: wait, aren’t those crypto coins? (Terra and Luna were, famously — a pairing that imploded back in 2022.) The jokes wrote themselves.

But the names aren’t a finance thing. They’re the sun, the earth, and the moon — and they’re OpenAI’s new way of telling you, at a glance, how powerful a model is. Once you get the system, it’s actually simpler than the alphabet soup of “4o, o3, 5.5” we’ve been living with. Here’s what each one is, and which you’d reach for.

Update (July 10, 2026): All three models are now live for everyone — GPT-5.6 went public in ChatGPT, Codex and the API on July 9, after two weeks gated behind a government review. This guide is updated with the real plan-by-plan mapping and the final launch numbers from OpenAI’s release announcement.

The one idea that makes it click

GPT-5.6 isn’t one model. It’s a family of three, and the names are a ladder from “fast and cheap” to “slow and brilliant”:

🌙 Luna
Fast, low-cost. Quick questions, short drafts, simple rewrites, high-volume routine work.
🌍 Terra
The balanced default. As good as today's ChatGPT, cheaper. Most everyday work lands here.
☀️ Sol
The flagship. Deep reasoning, big coding jobs, research, multi-step projects.
quick & everyday reach for… hard & important

That’s the whole mental model: Sol = sun = brightest/strongest, Luna = moon = light and quick, Terra = earth = the solid middle. The number “5.6” tells you the generation; the name tells you the tier.

And here’s the clever part for the long run: OpenAI can upgrade each tier on its own. When the cheap model gets better next time, it’s still “Luna” — you don’t have to relearn a new name. It’s “good / better / best” that’s built to last.

How to read the name
GPT-5.6 the generation
+ Sol / Terra / Luna the tier (strong → light)
= which model, and how powerful
The number is the generation; the name is the power tier.

Sol — the heavy hitter

Sol is the one built for hard problems. OpenAI aims it at serious software engineering, scientific research, cybersecurity analysis, and long “agent” tasks where the model has to plan and work through many steps without losing the thread.

It also introduces two new “effort” settings you’ll hear about:

  • “max” — lets Sol take more time to think deeply on a tough problem instead of answering fast. Useful for a knotty proof, a tricky bug, or a decision with lots of moving parts.
  • “ultra” — goes a step further: it coordinates four AI agents working in parallel on your task by default (developers can push that to 16 through the API), then combines their results. Think of it as Sol delegating pieces of a big job to a small team it manages itself. On OpenAI’s launch numbers, ultra lifted Sol’s score on Terminal-Bench (a hard command-line benchmark) from 88.8% to 91.9%. It also burns usage dramatically faster — one early user reported a few minutes on ultra consuming most of a five-hour allowance — so save it for work that genuinely deserves a team.

There’s also a Sol Pro variant — a highest-quality version reserved for Pro and Enterprise plans in regular chat.

For the leaderboard-watchers, the honest picture from OpenAI’s own published table: Sol sets new records on agent-style and command-line work (Terminal-Bench, computer-use tests), while Anthropic’s Claude models still lead elsewhere — Claude Fable 5 edges Sol on two broad intelligence indexes, and Claude’s coding specialists remain well ahead on the SWE-Bench Pro software benchmark. Nobody swept the board. For everyone else, the takeaway is simpler — Sol is the one you’d point at your most demanding work.

Terra — the one you’d actually use most

Terra is the quiet hero. OpenAI describes it as a “balanced model for everyday work” with performance competitive with GPT-5.5, at about half the cost. For the vast majority of normal tasks — writing, summarizing, planning, answering questions — Terra is plenty, and it’s the tier most everyday use settles on.

It’s also the free tier’s ticket in: Free and Go users get Terra inside ChatGPT Work and Codex — the first time a brand-new frontier generation has reached free accounts on launch week.

If Sol is the specialist you call in for the hard case, Terra is the reliable generalist you talk to every day.

Luna — fast and cheap

Luna is built for speed and volume. It’s the lightest, least expensive option, meant for quick interactions and high-throughput, routine work where you’d rather have an instant answer than the deepest possible one. Short rewrites, quick lookups, simple drafts — Luna’s lane.

The launch benchmarks — and how much to trust them

OpenAI published a full comparison table at release. Here are the headline rows, kept honest by including the ones OpenAI didn’t win:

Benchmark (what it measures)SolTerraLunaGPT-5.5Best Claude
Agents’ Last Exam (long professional workflows)52.7%50.4%50.3%46.9%45.2% (Opus 4.8)
Coding Agent Index (agentic coding, independent)8077.474.676.477.2 (Fable 5)
Terminal-Bench 2.1 (command-line work)88.8% (91.9% ultra)87.4%84.7%85.6%88% (Mythos 5)
OSWorld 2.0 (using a computer)62.6%50.2%45.6%47.5%54.8% (Opus 4.8)
SWE-Bench Pro (real software fixes)64.6%63.4%62.7%59.4%80.3% (Mythos 5)
GDPval-AA (expert-rated work quality)1,747.81,5931,591.81,493.71,759.6 (Fable 5)
Artificial Analysis Intelligence Index58.95551.254.859.9 (Fable 5)

Read as a whole: Sol leads on agent-style work and computer use, Claude still leads on deep software engineering and holds a slight edge on two broad quality measures. Luna beating last generation’s flagship on several rows is the quiet story — that’s the free-tier model.

There’s a second dimension the score columns hide, and the independent measurement firm Artificial Analysis put numbers on it: efficiency. On their runs, Sol at max reasoning took the Coding Agent Index lead while using less than half the output tokens and less than half the time of Claude Fable 5, at roughly a third less cost — and came within one point of Fable 5 on their broad Intelligence Index in 61% less time at about half the estimated cost. For anyone paying per token, near-parity quality at half the cost and speed is arguably the real launch story — bigger than any single score gap.

Now the disclaimer, because it matters more than the table. These numbers come from OpenAI’s launch post. Vendors choose which benchmarks to publish, which competitors to include, and which settings to run — that 91.9% is Sol on ultra, a mode that burns usage four agents at a time, not the Sol you’ll use on a Tuesday. And benchmarks are contest conditions: clean inputs, well-defined goals, no office politics. Your actual work is ambiguous, half-documented, and full of context no leaderboard measures. A model that scores three points higher on a chart can absolutely be worse for you — because of speed, tone, how it handles your files, or how often it makes things up in your field.

One early piece of independent evidence is worth adding, because it cuts both ways. Cursor — the AI code editor, with no dog in this fight — runs its own benchmark (CursorBench) built from real user sessions: ambiguous, multi-file tasks, the opposite of contest conditions. On the current version, Claude Fable 5 Max scores highest (70.5%), with GPT-5.6 Sol Max close behind (67.2%) — but Sol did it at $5.22 in model costs against Fable’s $17.32. So on one of the few real-world-flavored evals available, Claude keeps a small quality lead and Sol wins decisively on price. Which of those matters more depends entirely on whether you’re paying per token.

The only benchmark that settles anything is a week of your own tasks. Run the same three real jobs on Sol and on whatever you use today, and keep the one that needed less babysitting.

How this maps to ChatGPT (live since July 9)

The tiers are no longer theoretical — here’s the actual mapping, straight from OpenAI’s availability notes:

WhereFree / GoPlusPro / Enterprise
Regular chatGPT-5.5 Instant (no 5.6)Sol (medium+ effort)Sol, plus Sol Pro
ChatGPT Work & CodexTerraSol, Terra, Luna + effort dialSol, Terra, Luna + effort dial
max setting✓ (togglable)✓ (togglable)
ultra settingCodex only✓ in Work and Codex

Two things worth noticing in that table. Free users’ only door to GPT-5.6 is the agent surfaces — ChatGPT Work and Codex — not normal chat. And Plus sits one notch below Pro twice: no Sol Pro, and no ultra inside ChatGPT Work.

On timing: OpenAI’s official plan was a global rollout “continuing over the next 24 hours” from the July 9 launch — so by July 10–11, every account should see the full mapping. During that window some Plus users reported seeing only Terra and Luna; if your picker still looks short after that, sign out and back in (or update the app) before concluding your plan is missing something.

Day to day, ChatGPT still routes most tier-picking automatically. Knowing the ladder matters for the moments you’d override it: nudging a hard job to Sol, or noticing a quick question doesn’t need it. (Our guide to which ChatGPT model to use covers the picker in full.)

What this means for you

  • If you’re a casual user: you mostly won’t think about this. The system picks the right tier for you. Just know “Sol” = the powerful one if you ever get to choose — and that on a free account, GPT-5.6 lives in ChatGPT Work, not regular chat.
  • If you do serious work in ChatGPT (coding, analysis, research): Sol at “max” is now available on every paid plan — that’s your default for hard tasks. Reserve “ultra” for the rare job that’s worth four agents and the usage bill that comes with them.
  • If you build apps or watch costs: the three tiers are a price/quality dial — Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per million tokens, with Sol priced identically to GPT-5.5. Terra at half the flagship’s cost is the headline; Luna is for cheap, high-volume calls. And per Artificial Analysis’s independent runs, Sol’s effective cost gap vs the top Claude models is bigger than the price sheet shows — it finishes comparable work in fewer tokens and less time.
  • If you just want the names to stop confusing you: sun, earth, moon — brightest to lightest. That’s it.

What the tiers won’t do

  • They won’t make a weak prompt smart. Pointing Sol at a vague question still gets you a vague (if eloquent) answer. Clear asks matter more than tier choice.
  • They won’t all feel different for simple chat. For everyday questions, you’d struggle to tell Terra from Sol. The gap shows up on hard, multi-step work.
  • They won’t be free across the board. The mapping is set now: free accounts get Terra (in ChatGPT Work and Codex only), Sol needs a paid plan, Sol Pro and ultra-in-Work need Pro or Enterprise.
  • They won’t perform like the launch chart on your desk. Benchmark scores are contest conditions at chosen settings. Judge the tiers on your own work, not the leaderboard — see the disclaimer above.
  • They won’t replace knowing your field. A brighter model still doesn’t know your job. It drafts; you decide.

The bottom line

Sol, Terra, and Luna aren’t crypto coins and they aren’t a gimmick — they’re a clean “good / better / best” ladder, with the generation number (5.6) separate from the tier name. Sol for hard problems, Terra for everyday work, Luna for fast and cheap. As of July 9 they’re live for everyone: check the plan table above for where each one lives on your account, treat the launch benchmarks as a menu rather than a verdict, and let your own work pick the winner.

Want to actually get good at picking and prompting models — the skill that outlasts every naming change? Start with ChatGPT vs Claude: Which Should You Use? or AI Fundamentals. And if the new agent side is what tempts you, our plain-English guide to ChatGPT Work is where Terra earns its keep.

Sources

Build Real AI Skills

Step-by-step courses with quizzes and certificates for your resume