This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.
At AI week, Gian Segato from Anthropic said something offhand that I have not been able to put down. He mentioned that a lot of people inside Anthropic keep their own personal eval for Claude. Not the big public benchmarks. A small, private test, tuned to something they personally care about, that they trust more than any leaderboard to tell them whether a new model is actually better. That one…
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.