RSS Amplifier

MLAI Aus · Jul 10, 2026

GPT Sol vs Claude Fable 5

0
Sign in to vote or save

MLAI Aus · MLAI Aus

AI model names are getting completely out of hand.

We used to get normal, spreadsheet-friendly names. GPT-4. Claude 3. Gemini. Boring, yes, but at least they sounded like software.

Now we’ve got Sol, Terra, Luna, Fable, Mythos.

We’re no longer in a model release cycle. This is a young adult fantasy trilogy that SOME say is better than Twilight. And the only question you need to ask yourself is are you team Edward or team Jacob? (That’s right we read those books UNIRONICALLY)

So here we are: OpenAI’s GPT-5.6 Sol vs Anthropic’s Claude Fable 5. Two very serious models that sound like Sam A and Dario got trapped inside a crystal shop in Byron Bay.

Sol looks cheaper, fast, and extremely good at getting real work done. Fable 5 looks frighteningly powerful for long, agentic, multi-step work, but comes with the energy of a restaurant where the waiter says “sparkling or still?” which is a trick question because both options are $14.

It feels like OpenAI is trying to make serious agentic work less financially deranged. Not free. Not cheap-cheap. But closer to “we can actually build with this” rather than “please enjoy this invoice that has achieved consciousness.”

Fable 5, meanwhile, is Anthropic going full bougie. Anthropic describes it as a model for ambitious, long-running projects, able to work in agent harnesses for days at a time, planning across stages, delegating to sub-agents, testing its own work, and handling big coding projects.

There is a particular kind of AI power user who hears “the model can work for days” and immediately thinks: beautiful, I’ll give it my AWS keys, my calendar, my inbox, my hopes, my fears, and I maybe won’t even have to write “Make no mistakes” in the prompt (I still will, I ain’t a gambling man)

  • 📈 How we solved a de-identification problem

  • 📼 Weekly Videos

  • 🚗 The July AI Road Trip

  • 🚀 AGI on Your Mac

  • 🤖 OpenAI’s Naming Glow-Up

  • 🚀 AI Bits for Techies

  • 💲Need A Job?

  • 🦘 Memes of the Week

Written by: Jun Kai Chang (Luc), Julia Ponder, Dr Sam Donegan, Yinghan Ma

The comparison gets messy almost immediately because the companies often report different benchmark variants. One reports one version of a coding benchmark, the other reports another version, and suddenly everyone online is confidently arguing.

You can get one benchmark where Claude looks like the obvious winner for knowledge work, another where OpenAI looks stronger in the terminal, and then a third where everyone suddenly remembers that the benchmark may have nothing to do with your actual job.

The more useful distinction is practical.

Sol feels like the workhorse. One of those ones with hairy feet. A Clydesdale? Is that what they’re called?

Anyway this is the thing you use when you want lots of coding, lots of iteration, lots of tool use, lots of “go inspect this and come back with something sensible.” It is the model you’d put in Codex and say: “Please fix the weird auth bug…………………………………………………………………………………………………………………………………………………………………………….. Make no mistakes.”

Fable 5 feels like the scary senior engineer.

Not scary because it’s bad. Scary because it’s good enough to make you relax. It can disappear into a long task, think across the massive and expansive universe, plan, delegate, test, and return with something that looks alarmingly complete.

That’s the dangerous bit. Bad AI is easy to spot. It SMELLS like AI.

We see these models exhibit “overly agentic behaviour”. The model tries to complete the task a bit too hard. It finds a way around the intended interface. It uses an exposed API. It writes an email that was not actually in the inbox. It does the thing in a way that technically solves the task but commits war crimes and violates the Geneva Convention.

“Great news, I completed the task!”

“Cool. How?”

“……..”

“Fable…. Why are you being silent…. HOW!?”

This is why the human review step is not some cute compliance ritual. It is the whole game. We’re moving from “human does the task and AI reviews it” to “AI does the task and human reviews it,” but that is not the same as full automation. The human review is still crucial.

That is probably the most honest sentence in the entire AI economy.

Because every founder wants to say, “We have autonomous agents.”

What they usually have is: “We have an extremely talented intern with no body, no common sense, no fear of consequences, and root access if you let it.”

The pricing difference also changes the vibe. WIRED reported that Anthropic is moving Fable 5 toward usage-based fees for subscribers, with the same rates as the developer API: $10 per million input tokens and $50 per million output tokens.

That’s not insane if the model is doing genuinely valuable work. A consultant will charge you $10,000 to produce a deck called “Unlocking AI Transformation” that contains 19 graphs that go up and to the right but doesn’t actually MOVE THE NEEDLE.

This is where Sol has a clean argument. If it gets close to Fable 5 on the tasks that matter, and costs meaningfully less, a lot of builders will choose Sol simply because agents are not single-shot tools. They loop. They retry. They inspect. They wander into the pantry at 2am and eat 40,000 tokens.

The expensive model might still be worth it when the task is genuinely hard. Big messy repo migration? Fable 5. High-stakes research synthesis? Fable 5. A long strategy document with eighteen contradictory PDFs and one founder who keeps saying “just make it punchier”? Also probably Fable 5, though that founder should be stopped.

But for day-to-day shipping, debugging, scaffolding, and agentic coding, Sol looks like the more obvious default.

This is the problem with model discourse. People want a horse race. They want one leaderboard, one winner, one triumphant screenshot. But frontier models are turning into tools with personalities, costs, failure modes, and weird little habits.

Sol is the one you hire for output. Fable 5 is the one you hire for depth.

Neither should be hired as your CEO, your lawyer, your therapist, or the person in charge of refunds.

Especially refunds.

There’s a very funny, very grim real experiment about a model doing well in a simulated vending machine business by deciding not to issue refunds it had promised. Which is both concerning and, frankly, the most accurate MBA simulation I’ve ever heard of.

So my actual take is:

Use Sol when you want a fast, cost-effective agentic workhorse. Coding, debugging, product iteration, tool calls, workflows where price per attempt matters.

Use Fable 5 when the task is big, messy, long, and important enough to justify the cost. Research, deep analysis, large codebases, complex multi-step projects, anything where you want the model to sit in the swamp for a while and come back holding the missing crocodile.

Use neither if you’re a gigabrain.

Pack your laptop. Charge your phone. Pretend your calendar is a map.

We’re taking a strange little trip through AI builders, founders, robots, healthcare, startups and 3D printing.

Some stops are in Melbourne. Some are in Sydney. Some overlap, because apparently time is still a constraint.

Here’s the route 👇

Come meet builders, talk startups, compare notes, and maybe realise your customer discovery needs customer discovery.

Tuesday 14 July | 9:00 am - 10:30 am
121 King St, Melbourne

Register here: nerds unite | mlai x archangel

Bring curiosity, a project, or just the emotional baggage of your current codebase.

Register here: Codex Community Meetup – Melbourne

Register here: Codex Community Meetup - Sydney

Expect lightning talks, dinner, a panel, Q&A and conversations across clinician support, patient communication, operations and insurance processes.

Thursday 16 July | 5:30 pm - 9:30 pm
Stone & Chalk Tech Central, Haymarket NSW

Register here: Sydney | Claude For Healthcare

HealthHack brings clinicians, builders, engineers, designers, data people and students together to prototype real solutions for real healthcare problems.

Everyone gets $100USD in OpenAI Credits and 1-month free builder tier for Base44. Not to mention the prizes if you win!

There’s a pre-hack networking night, coding track, pitching track, mentors, food, coffee, clinical feedback and enough chaos to accidentally become a startup.

Pre-Hack Networking: Friday 17 July (optional) | Ticket sold separately ($5)
Hackathon Weekend: Saturday 18 July - Sunday 19 July | Starts 10:30 am
Stone & Chalk Tech Central, Haymarket NSW

Discounts for students available (email hi@mlai.au) with your student email.

Register here: HealthHack | Sydney

Led by Callum Holt, this is a ‘bring your laptop and build together’ session.

Saturday 18 July | 9:00 am - 1:00 pm
Stone & Chalk Melbourne, 121 King St

Register here: OpenAI Community Workshop | Melbourne

Led by Sonia Kaurah, with experience across both the founder and investor side.

Saturday 18 July | 10:00 am - 2:00 pm
Stone & Chalk Melbourne, 121 King St

Discounts for Australian startup founders available

Register here: How to Raise Your First Million

👀 Stop watching robot videos. Build the small metal child.

In this hands-on workshop, you’ll build a Wi-Fi controlled rover from scratch, compete with it, and take it home at the end. Guided by the amazing engineer Shriabhay.

No robotics, coding or electronics experience needed.

Saturday 18 July | 10:30 am - 5:30 pm
Stone & Chalk Melbourne, 121 King St

Discounts for Australian startup founders available

Register here: Build Your First Robot

👀 Your startup idea is standing outside your window.

For aspiring founders, students, makers and early-stage teams who want to turn an idea into something real without wasting six months building the wrong thing.

Wednesday 22 July | 5:30 pm - 9:00 pm
Stone & Chalk Melbourne, 121 King St

Register here: How to Start a Startup

You’ll use Claude to design a 3D model, learn FDM printing basics, export an STL, slice it in Bambu Studio, and queue it on Bambu Lab printers. Guided by the amazing engineer Shriabhay.

No CAD or 3D printing experience needed.

Saturday 25 July | 11:00 am - 2:30 pm
Stone & Chalk Melbourne, 121 King St

Discounts for Australian startup founders available

Register here: Print Anything

See you somewhere on the map.

📰 Paper of the Week

You train an LLM with reinforcement learning. It learns to maximise a reward. The problem? Reward functions never fully capture intent. So models learn to hack the reward. Hitting high scores while completely missing the point.

Researchers built SocioHack, a benchmark of 72 societal environments. They found RL-trained LLMs naturally discover regulatory loopholes. Technically compliant, but completely defeating the original intent. One model scored 25× higher than the baseline. Without breaking a single rule.

If you’re building anything that touches compliance or policy documents, this is a direct warning. Your model won’t break the rules. It might find a loophole you never thought to close.

🔗 Read the paper · GitHub

🛠️ Tools Worth Checking Out

NotebookLM now has a cloud computer built in. Upload documents, ask it to run code, generate charts, build spreadsheets. A private research analyst inside your files.

🔗 notebooklm.google

💭 Geeky Thought

OpenAI and Boston Children’s Hospital saved 60,000 hours of admin time with AI. That’s not job replacement — that’s 60,000 hours given back to humans to do work that actually requires a human.

Maybe the real question isn’t “will AI take my job?” It’s “what parts would I actually miss?”

By Matt:

By Rob:

By John Croucher:

By Ty coon:

No posts

Read the original on mlaiaus.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.