Vol. IV · 18
Quantization From the Bits Up: How 4-Bit Models Actually Work
Quantization · Part 1
Every series I've written so far has bumped into quantization from the side — QLoRA compressing frozen bases, the optimization guide's memory math, GGUF files on Hugging Face with cryptic suffixes.
Read entry