Anthropic released Claude Opus 4.7 yesterday. I spent the morning testing it on legal tasks. Here’s what stood out.
The headline improvement for legal work isn’t the benchmark scores, it’s how the model handles instructions. Ask it to do something specific and it now tends to do that thing, rather than taking creative liberties. For legal work, where precision matters more than flair, that’s a welcome improvement.
Earlier Claude models had a tendency to interpret instructions more loosely. For example, if you asked to flag deviations from a playbook, it would sometimes rewrite clauses, or quietly skip parts of the instruction. Opus 4.7 is noticeably more disciplined. In my testing it tends to stay on task and followed multi-step instructions without losing the thread halfway through.
The flip side is also worth flagging. If your existing prompts are a bit loose, this model might expose that. It does what you tell it, which means you need to tell it exactly what you want. Teams running prompt templates built for 4.6 should retest them before switching over.
Another thing to note, if you’ve written system prompts that lean hard on steering the model, Opus 4.7 can overcorrect. In my own testing it was oversensitive to flagging potential prompt injection, to the point where it flagged some of my own skills as potentially suspicious. My guess is that this behaviour sits in the system prompt for the Claude app itself. Either way, worth keeping in mind when you evaluate the model (especially in a product context versus via raw API).
Harvey ran their BigLaw Bench on the new model and it’s the highest-scoring Claude model in Harvey to date. They noted it now correctly distinguishes assignment provisions from change-of-control provisions, something that has historically tripped up frontier models. That matches what I saw on similar tasks.
It isn’t a clean sweep, though. Harvey’s evaluators flagged some polish gaps compared to the previous version (Opus 4.6): occasional tone mismatch, a tendency to overshoot on detail, and non-standard citations. I’ve seen the same things. These are exactly the kind of issues that create rework for associates and that benchmarks don’t always capture.
An underrated improvement: Opus 4.7 processes images at roughly three times the resolution of previous Claude models. For legal work, that means better handling of scanned contracts, dense annexes, tables inside PDFs, and documents with small print.
One caveat. If you’re using a legal AI platform, this is mostly relevant for pure image inputs. Most legal AI providers don’t rely on the built-in vision capabilities for document processing; they use something like Reducto, which already outperforms the foundation models’ native document handling. If you’re using Claude directly, though, you’ll notice the difference.
Pricing stays the same as Opus 4.6 ($5 per million input tokens and $25 per million output tokens). There’s a catch, though. The new tokenizer can use up to 35% more tokens on the same input depending on content type, and the model tends to think longer at higher effort levels. For teams processing high volumes of contracts or documents, that adds up. Worth factoring in when you budget for production use, especially if you’re paying per token (for example when integrating the API directly).
What I’m seeing lines up with the benchmarks. The model is more disciplined, follows instructions more precisely, and is better at working through complex multi-step tasks without losing the thread. For contract review, summarisation, and document analysis, it feels like a meaningful step forward.
It isn’t without friction. The stricter instruction following means your team needs tighter prompts. The occasional polish gaps (tone, citation style, over-detailing) mean human review remains non-negotiable. None of that is a reason to hold back, but it is a reason to test on your own workflows before rolling it out widely.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.