The house style of frontier LLMs still seems rather obnoxious at times, which I particularly noticed with GPT-5.5 in Codex. Anthropics new Fable 5 does best on the Creative Writing EQ-Benchmark (2229.6 score vs 2028.5 with GPT-5.5), but Ethan Mollick calls it “overwrought language” that cannot be prompted away:
Ethan Mollick@emollick
The overwrought language of Fable is an ongoing problem in software and design projects where the AI is otherwise excellent. It just creeps into little bits of text and toasts and menus, even if you try to instruct otherwise. And once it gets in it is hard to purge completely.
5:24 PM · Jul 7, 2026 · 50.3K Views
50 Replies · 12 Reposts · 520 Likes
Peter Gostev confirms limited steerability for GPT-5.6 too: “Sol feels quite difficult to align to what I want to say or explain things to me simply”:
Peter Gostev@petergostev
My view of: Fable 5 vs GPT-5.6-Sol. They are not easy models to compare, these are my vibes - take them as you will. My overall feel is that Fable is a 'wise owl' who is very thoughtful and very well spoken, GPT-5.6-Sol is like a rottweiler who will grab the problem by the
6:06 PM · Jul 8, 2026 · 2.14K Views
2 Replies · 5 Reposts · 75 Likes
And he confirms GPT Pro writing capabilities - which seems wasteful, but worth-it to me: “Though I think the 'Pro' model writes clearer [than Fable]”.
Yet, both Ethan and Peter may not have been persistent enough.
To tackle the slop, Louis-Francois Bouchard shared his “How to use AI without sounding AI” skill and Shreya Shankar did likewise.
Louis-Francois starts before generation. His guide is a prompt scaffold: task, audience, source material, structure, style, banned phrases, and a review loop with a second model. This is useful when the model has not written yet. It is mostly about reducing degrees of freedom.
Shreya starts after generation, or at least inside the writer’s head while reading the draft. Her skill is stricter and more local: simple words, complete sentences, no fake emphasis, no vague claims, no ornamental structure. In a second pass, every clause has to justify why it is still there.
So Louis is closer to prompt engineering. Shreya is closer to editing. And Peter shares my recommendation of using GPT Pro.
Update 2026-07-26:
Anthropic’s Opus 5 has surpassed Kimi K3 on an internal writing benchmark at Towards AI:
Louis-François Bouchard 🎥🤖@Whats_AI
Big news from our internal writing benchmark (early results): Claude Opus 5 by @AnthropicAI is now #1 for writing in our editorial voice, at 2817 Elo, surpassing Claude Fable 5 and Kimi. Already! That is a jump from #15 to #1 over its predecessor (if we take all thinking

Claude @claudeai
Introducing Claude Opus 5. It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price.
12:49 AM · Jul 25, 2026 · 22.3K Views
13 Replies · 25 Reposts · 299 Likes
Kimi K3, when loosely instructed to use the Skills linked above and given samples of prior writing, already beats the Pangram feature on Substack - its result coming out as “Fully Human Written” after light manual edits.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.