This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.
Throughout recent years, LLM capabilities have outpaced evaluation benchmarks. This is not a new development. What is new is that the set of standard LLM evals has further narrowed and there are questions regarding the reliability of even this small set of benchmarks.
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.