RSS Amplifier

Fairly AI · Jul 12, 2026

Claude Fable, GPT 5.6, and a Glimpse of Superhuman

0
Sign in to vote or save

Wei Chen · Fairly AI

Disclaimer: This article reflects my personal opinion based on one single experiment and does not constitute legal advice.

The coders seem to have seen the future. Andrej Karpathy said in March that he had not typed a line of code since December. Andrew Ng said that coding agents have exceeded his expectations. My husband, an AI researcher who is not easy to impress, said coding agents have achieved superhuman level.

Whatever they have seen in coding, superhuman or AGI, I have been looking for in legal tasks. This weekend, I might have gotten a glimpse of it in Claude Fable 5.

I first reported my findings in February, when the world of software experienced what some analysts have called the “Claude Crash,” wiping out approximately $830 billion in market value in less than a week.

Using Claude’s agent, Claude Cowork and its legal plug-in, I tested a legal task often called “contract review” or “contract due diligence”. I pointed Cowork to a local folder with 33 license agreements from the CUAD dataset, an open-source dataset curated by The Atticus Project, and asked Cowork to review these contracts and create a due diligence summary sheet.

The summary, at best, matched the level of an untrained law student, with all the mistakes a rookie makes:

  • IP Ownership: It struggled to differentiate between standard boilerplate that allows each party to own their own IP and specific IP transfer and joint ownership provisions.

  • Termination for Convenience: It mistakenly included clauses allowing termination for material breach following a cure period.

  • Change of Control: It included clauses that allowed assignment despite instructions to the contrary.

  • Cap On Liability: It highlighted waivers of consequential damages by mistake.

  • Non-Verbatim Text: Sections were summarized or truncated, making human validation nearly impossible.

  • Missing Labels: It missed a number of clauses that should have been included.

Five months later, Claude Fable 5 returned after nineteen days under a government order, and I was eager to try it. I gave the model the same 33 agreements and contract review task as in February.

When I got the result back, I read it with a heavy dose of skepticism.

“Is this really a change of control provision? I think Claude is wrong. Oh, never mind. Claude is right.”

“Are these really termination for convenience? Wow, they really are.”

“These summaries are way too short, must be missing nuance. Well, I guess not.”

By the fifth agreement, I was reliving my first ride in a Waymo. At first, I stared at the eerily turning wheel and wondered whether I had been too rash to risk my life. Five minutes in, when Waymo dodged an unexpected pedestrian and made a perfect turn, I sat back in wonder. Ten minutes in, I forgot that I was in a car without a human driver.

That is how I felt. If this summary had been created by a human lawyer, I would have stopped checking and trusted the results whole-heartedly. But to make sure I was not prematurely hyping Claude, I read all of the entries. I even compared them against the ChatGPT 5.6 summary, looking for mistakes. I found only 2 (highlighted in red), out of thousands of entries.

I was utterly floored. I may not be the most skilled in contract review, but after 25 years of reviewing tens of thousands of agreements, I think I am a high bar. Claude Fable 5 has outdone me in quality, and way faster. This would make Claude Fable 5 superhuman in contract review.

You can check out the results here.

In March, when GPT 5.4 came out, I published the results of a redline face-off of the work products of five experienced attorneys and three frontier models: ChatGPT 5.4, Claude Opus 4.6, and Gemini 3 Pro.

I asked each model to mark up three heavily negotiated clauses in a SaaS license agreement (i.e., warranty disclaimer, limitation of liability, and indemnification). This task is often called “contract redline.” The goal is to secure a reasonable market position from the seller’s perspective with minimal, precise edits.

Back in March, two attorneys took the top places. ChatGPT 5.4 came third, ahead of three of the five humans. Claude Opus 4.6 came last, for what I called a “bloody redline,” rewriting clauses and adding changes unnecessarily. In July, the result from Claude Fable 5 improved, but still lags behind the two top human attorneys. GPT 5.6 ranked last. Unlike the contract review task, there is no sign of superhuman here.

One interesting note is that, unlike in March, both redlines labeled the edits under my name. I was jolted a bit to see edits made by Wei Chen instead of Claude or ChatGPT. That said, practically speaking, this may be necessary; sending a draft to opposing counsel with “Claude” or “ChatGPT” listed as the author might not be a wise move.

You can check out the redlines here.

Compared to Claude Fable 5, the results from GPT 5.6 on both tasks are disappointing:

  • Contract Review: The summary sheet only included yes or no answers, without referencing the section numbers or the relevant clauses. Even with just the yes or no answers, GPT 5.6 made more than 10 errors.

  • Redline: GPT 5.6 scored significantly lower than the GPT 5.4 result from March. The redline added a service guarantee, a termination right and a right to refund. Ouch! Needless to say, it ranked last.

  • ChatGPT 5.4: which beat three attorneys in March, could no longer replicate its result. Four months later, it struggled to follow the basic instructions required to generate a redline.

Here is my takeaways from the weekend experiment, beyond the wow.

  • Use frontier models for contract review. If a model can produce a due diligence summary with 2 errors in thousands of entries, the scarce human skill shifts from doing the first pass to validating it. Ask the model to provide verbatim quotes and section numbers in the output.

  • Keep humans in the loop for contract redline. The lack of publicly available high-quality redlining dataset continues to impact model performance. For now, use AI output as a reference point and not trusted work product.

  • Retest frontier models every few months. Whatever conclusion you drew about AI from last year’s pilot, or even last quarter’s, is already stale. Keep a small, fixed set of your own documents and rerun it on each major release.

  • Treat “AI for contracts” as different skills. Claude Fable 5 performed at superhuman level for contract review, but not at contract redline. Each type of legal task is different. Test before trust.

  • Do not assume a newer model is better. GPT 5.6 regressed on contract redlining compared to GPT 5.4 in March, and GPT 5.4 in July could no longer reproduce its own March result. Your own evaluation set matters more than any leaderboard.

We now have access to models that can review contracts like experts, but we are still waiting for the ones that can negotiate them. When that day comes, I suspect it will feel like sitting in a Waymo: a few minutes of anxiety and distrust, and then, almost as quickly, forgetting that I was in a Waymo.

For more practical tips on AI governance and innovation, check out GenAI for the Legal Profession: Power User Edition, AI Strategy for Legal Leaders, Atticus AI Habits Workshop and my Fairly AI blogs.

No posts

Read the original on weichen221.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.