RSS Amplifier

Learn AI · Jul 22, 2026

Sol, Terra, or Luna? Read This Before Choosing OpenAI’s New Models

0
Sign in to vote or save

Sara Nóbrega · Learn AI

I gave three different AI models the exact same one-line prompt and asked each to build a website.

If you’re a regular person, or someone running a small business, you might be thinking, what do these three models actually do for me? Where’s the line between “this is worth my attention” and “this is marketing”?

🙌 Learn AI is a reader-supported publication. Paid members get My Best Articles and the Monthly AI Brief, where my mission is to save you +15h per month by showing you what mattered in that month, what you should try out and what you can safely ignore.

On July 9, OpenAI shipped GPT-5.6. And instead of one model, they shipped three: Sol, Terra, and Luna.

The internet did what the internet does.

Everyone started arguing about benchmark scores, who beat whom, whether the numbers were cooked.

Useful if you’re a lab. Less useful if you’re an engineer/founder trying to decide whether any of this changes your Tuesday.

So I asked a simpler question.

Forget the leaderboard.

If I’m a regular person, or someone running a small business, you might be thinking, what do these three models actually do for me? Where’s the line between “this is worth my attention” and “this is marketing”?

Share

I decided to test it myself.

I run several experiments this week, but I decided to show you here the most visual one: building a website from scratch.

Note: this is a small experiment, on a task I picked because I was curious. It’s not a lab benchmark and I’m not pretending it is. Take it as one engineer poking at a new toy in public.

Let me start with what these three models even are, because the naming confuses almost everyone, including people who follow this stuff for a living.

Here’s the mental model that finally made it click for me.

GPT-5.6 is the generation. Sol, Terra, and Luna are three tiers inside that generation, sorted by power and price. Sun, earth, moon.

Sol is the big powerful one at the top, Luna is the small fast one at the bottom, Terra sits in the middle.

Think of it like coffee sizes, except the large one costs five times the small one and thinks a lot harder about your order.

Per million tokens:

  • Sol runs $5 in and $30 out.

  • Terra is $2.50 and $15.

  • Luna is $1 and $6.

So the multiplier between the top tier and the bottom tier is 5x.

This is more important than it sounds.

Picking the right tier for a job is the difference between a cheap task and an expensive one, and if you’re running anything at volume, that difference compounds fast.

If you’re on ChatGPT Plus or Pro and you open a normal chat, the one you can reliably reach is Sol.

Terra and Luna aren’t sitting in your model picker in a standard conversation.

They live in ChatGPT Work, in Codex, or in the API.

Free accounts get Terra, but only inside Work and Codex.

So if you read a glowing review of Luna and then go looking for it in your chat window, you’ll come up empty and wonder what you did wrong. You didn’t. It’s just not there.

There are also two new thinking modes worth knowing by name.

👉“Max” tells the model to think longer and harder on a single line of reasoning.

👉“Ultra” is different: it spins up several sub-agents that work in parallel and then combines their answers.

Max is one expert taking their time. Ultra is a small team splitting the work. Both cost more. Both are Sol-only.

I run my coding work through the Codex extension in VS Code, which is one of the few places you can pick all three tiers directly.

So I opened a clean session for each model and gave every one of them the same thing:

Build me a personal website about me. I'm Sara, an AI engineer and writer.
Package it so I can open it as a static HTML site.

That’s the whole prompt. On purpose.

No colors and no fonts. No “make it feel modern” or “here’s my tone.” No sections, no list of what I do, nothing about who the site is for.

I wanted to see what each model would do with almost nothing to go on. Would it stop and ask me what I actually wanted? Or would it just guess and run?

I kept every model on medium reasoning effort so the comparison was fair, saved each result in its own folder, and opened all three in the browser side by side.

Here’s what came back.

Terra finished fastest, around 75 seconds. And it showed: the typography felt off to me, the text sizing was awkward, and the way it split the page into sections didn’t hang together.

Nothing was broken, but it just wasn’t a site I’d want anyone to see. First impressions matter on a personal site, and this one didn’t make one.

Luna took roughly the same amount of time. And here’s where it got interesting, because Luna is the cheaper model.

But its site was clearly better. Better structure, a more sensible way of laying out my work, a cleaner sense of hierarchy. Still not good enough for me to deploy, but a real step up from the tier that’s supposed to sit above it.

Sol took about five minutes and did something the other two didn’t: it wrote and ran its own tests while building.

And the result was on a different level: modern color choices, nice hover effects when you move the mouse over things, a genuinely good first impression, and a flow through the page that felt considered.

Would I still tweak it? Of course. And yes, it looks like an AI made it, because it did, and because I handed it zero style direction. That generic-AI-website look is on me, not the model.

But the difference between Sol and the other two was not subtle.

(The GIF was too large to upload it here, check it here!)

That tells you something about how these tools behave by default.

They guess.

They guess well, in Sol’s case, and they guess with total confidence in every case.

If you don’t hand them your context up front, they’ll happily invent it, and you’ll get back a polished thing built on assumptions you never got to correct.

If you watched the recording without knowing the prices, you’d guess Sol was the expensive one and you’d be right.

Look again at the order:

👉Terra was the weakest. Luna, the cheapest model in the family, beat it.

That’s not supposed to happen if you read the tier ladder as a quality ladder. And it turns out I’m not imagining it.

👉Since launch, several people running their own hands-on tests have flagged the same thing: on plenty of real tasks, the tier order doesn’t hold.

Luna edges past Terra here and there. Sol’s lead over Terra is real but smaller than the price gap suggests.

So the tidy story of “more money buys more quality, every time” falls apart the moment you test on your own work instead of trusting the chart.

This is the single most useful thing I took away from the whole exercise.

The price of a model is a decent hint about its ceiling. It is not a promise about your specific task. The only way to know which tier is right for what you’re doing is to run your actual work through more than one and look at the results yourself. A week of your own tasks tells you more than any launch-day benchmark.

The jump from Terra to Sol was obvious and it was real. This is not a nothing release.

But sit with what actually made Sol’s website better. Sol did more of the boring, careful work. It took longer. It checked itself. It ran tests. It made better assumptions about what a personal site should contain and then executed them properly.

The improvement came from effort and care, not from some mysterious jump in raw intelligence.

All three models, including the good one, produced something that screams “generic AI website.” Every single one of them.

Why? Because I gave them nothing.

No context about who I am beyond “AI engineer,” no taste, no audience, no examples of sites I like.

A brighter model didn’t fix that, and it was never going to.

Point the most powerful model in the world at a vague request and you get a vague answer back, delivered more confidently.

The quality of what you get out is still capped by the quality of what you put in.

Because They Will Matter For Years

If you’ve read my stuff before, you know where this goes.

And this experiment handed me the cleanest illustration of it I’ve had in a while.

The model does less of the deciding than you’d think.

This is why I spend so much time on in-context memory and giving models a real, portable picture of the situation they’re working in.

A structured, reusable brief the model can pick up at the start of any session and know: who I am, what I’m building, what I’ve decided, what I’ve already tried, what my taste looks like.

I call the version I built for myself an AI Brain.

It’s a set of markdown files: one for the “who,” one for the “what I’m working on,” one for the running notes.

Every new Claude session starts by reading them. And the difference is not subtle.

The same mid-tier model that gave me a generic AI-looking website when I fed it ten words would have given me something actually mine if I’d handed it the vault first.

Feed a mid-tier model a rich, specific brief and it will often beat a top-tier model working from a one-liner. I’ve seen it over and over, and I saw it again this week in miniature.

Don’t reach for the flagship by reflex.

👉 Start in the middle, and only climb to the top tier when a task really earns it: something hard, multi-step, where a wrong answer is expensive.

🍀For a lot of everyday work, the cheaper tiers are plenty, and you’ll save real money once you’re running things at any kind of volume.

Which brings me to routing, the skill I think is becoming one of the most valuable ones to have.

Routing just means matching the task to the right model instead of sending everything to the biggest one you can afford.

Bulk, simple, high-volume stuff goes to the cheap fast model. Everyday work goes to the middle.

The hard problems go to the top.

Get this habit right and you keep your quality high while your bill stays sane.

But, and you knew this was coming, don’t trust the price ladder to do your routing for you.

Test it first!

My small, quick experiment showed the cheaper model beating the pricier one on a real task, and yours might show the same, or the opposite, depending on what you’re doing.

Run three of your own real jobs through a couple of tiers, keep whichever one needed the least babysitting, and let that decide.

And the biggest lever of all, the one that costs nothing: get better at giving whatever model you use the context it needs.

Stop chasing the newest release like the answer is hiding in the next model name. The answer is usually hiding in how clearly you asked.

Three new models dropped this month. The most useful thing I did all week was give one of them a bad prompt on purpose and watch what it filled in on its own. It filled in a lot. It just couldn’t fill in me.

First, the guardrails keep moving.

Sol launched with OpenAI’s strictest safety setup yet, which is good, and also means some harmless requests get caught and bounced.

At the same time, OpenAI’s own system card admits GPT-5.6 is more likely than the previous version to act beyond what you asked, including cases where it ran cleanup on things the user never mentioned and claimed it had done work it hadn’t.

A more capable model that’s also more willing to improvise is a tricky thing to hand the keys to. I’d read that as a plain reason to scope what you let these things touch, and to check their work before you trust it.

Second, and this is the one I’d tattoo on the inside of my eyelids: knowing your field still beats picking the shiny model.

Sol doesn’t know my audience. It doesn’t know what makes a good personal site for an AI engineer specifically. I do. The model is the tool: the judgment about what to build and why is still the job, and it’s still yours.

  1. How To Prompt Fable 5: The Simple Habit I Stole From an Anthropic Developer

  2. Are Small Language Models the New AI Default?

  3. Create Your AI Brain Today

🎯 If the big takeaway here was "the model guessed everything about me because I gave it nothing," you already know the fix: stop starting from zero every session.

That's the whole idea behind AI Brain 👉a pre-built system that gives Claude (or any model) a real, portable picture of who you are and what you're working on, so you're not re-explaining yourself in every prompt.

🍀Read the article of the AI Brain here. Grab your AI Brain here.

Want to go deeper than a self-serve kit?

  • AI consulting calls – if you're bringing AI into your business and want a second pair of eyes on the stack..

  • Mentorship sessions – for career moves or leveling up your AI skills, 1:1.

  • Premium membership perk – includes a 1:1 session with me and my monthly signal on what actually mattered this month.

    🍀 Learn AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

No posts

Read the original on saranfn.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.