Read a blog recently where the author said “LLMs use fewer tokens for input than output so they’re cheap if you don’t use them to generate code” which does kind of ignore that “reasoning” involves models generating text to talk to themselves ad nauseum









