Sasank Chilamkurthy · Dec 30, 2023
Hardware Design for LLM Inference: Von Neumann Bottleneck
0Sign in to vote or save
This page cannot be shown here. You can still read it on the original site — the toolbar below keeps your place in the directory.
I was speaking to Prof. Veeresh Deshpande from IIT Bombay about optimal hardware system design for LLM inference. I explained to him how my ‘perfect’ hardware should have equal number of floating point operations per second (FLOPS) and memory band...
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.