This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.
How good are LLMs at trivia? I used the Jeopardy! dataset from Kaggle to benchmark ChatGPT and the new Llama 3 models. Here are the results: There you go. You’ve already gotten 90% of what you’re going to get out of this article. Some guy on the internet ran a half-baked benchmark on a handful of LLM models, and the results were largely in line with popular benchmarks and received…
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.