This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.
Sub-INT8 quantization-aware training and inference kernels (W83); Local-first LLM serving stacks on consumer hardware (W82); Sleep-like memory consolidation for long-context LLMs (W79)
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.