This page cannot be shown here. You can still read it on the original site — the toolbar below keeps your place in the directory.
We argue that the sum total of everything LLMs do is a form of mutation or fuzzing, and that exactly one hard benchmark will suffice: number of tokens expended to converge to the mathematically correct output. We certainly don't have heated arguments about tar , ls , or any other computer program that takes some input and produces some output. The difference is that the program in question…
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.