RSS Amplifier

Machine Learning · Aug 22, 2026

I developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R]

0
Sign in to vote or save

This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.

I trained a 250M parameter model from scratch on 30B tokens of fineweb. It’s quantized to under 2 bits so the whole deployment is 60 MB and it needs about 80 MB of RAM to run. Runs around 400 tok/s on a normal laptop CPU, no GPU needed. How the long context works: the most recent 2048 tokens stay in fp16 like a normal KV cache. Everything older gets compressed to 1 bit and written to disk, about…

Read on reddit.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.