Q8 KV cache: a quarter of the bytes, seventeen percent slower decode on Metal
In July this blog argued a Q8 KV cache is close to free on an RDNA4 card. On August 17 we turned it off by default for Muse Glimmer on Apple Silicon, because it was costing 15 to 22 percent of decode. Both results are correct. Q8_0 stores a KV element in 1.06 bytes instead of 4, so the quantized path reads 3.8 times fewer bytes and still loses, which means the flash-attention kernel at ordinary context lengths is not bandwidth-bound at all. KV quantization is a capacity decision that a decode kernel charges you for in time.