A minimal Qwen3-0.6B in pure PyTorch, running on Apple Silicon, with a live UI that makes the memory and latency trade-offs of KV cache physically visible. Toggle the cache off and watch attention go quadratic in real time.
A deep dive into making Plymouth boot animations work with NVIDIA's proprietary driver on Debian 13, involving kernel module timing, DRM framebuffer initialization, and initramfs configuration
Hi there! Thanks for checking out my blog. Update: follow Coding with Intelligence (CoWI) for curated updates about LLMs and Machine Learning: codingwithintelligence.substack.com My name is Rick Lamers . I’m an AI Researcher and I live in Amsterdam. Previously, I ran a startup called Orchest. We built a workflow orchestration tool not unlike Apache Airflow. We had an open source project and…