This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.
High end NVIDIA cards, and the server and power needed to run them, cost a lot of money, especially if you plan to reach enough VRAM to run massive models. The alternative, so far, has been Apple hardware, or the DGX Spark that, even if severely limited because of memory bandwidth, still allows to run LLMs prompt processing (prefill) fast enough. The Mac Studio provided up to 512GB unified memory,…
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.