QuantWM targets flicker in 2-bit video-world-model caches
Researchers from iLearn-Lab at Harbin Institute of Technology (Shenzhen) and LV-Lab at the National University of Singapore propose QuantWM, a training-free framework for 2-bit KV Cache quantization in video world models. The team says its query-sensitive clustering and low-rank attention compensation improve visual stability while achieving up to 6.20x KV Cache compression across five evaluated models. Why it matters: The work highlights that standard video metrics may miss temporal instability caused by low-bit KV Cache quantization. What do you think?
Open the headline: source, AI briefing and more →