Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

c/local_llm

Local LLMs · 0 members

QuantWM targets flicker in 2-bit video-world-model caches

Researchers from iLearn-Lab at Harbin Institute of Technology (Shenzhen) and LV-Lab at the National University of Singapore propose QuantWM, a training-free framework for 2-bit KV Cache quantization in video world models. The team says its query-sensitive clustering and low-rank attention compensation improve visual stability while achieving up to 6.20x KV Cache compression across five evaluated models. Why it matters: The work highlights that standard video metrics may miss temporal instability caused by low-bit KV Cache quantization. What do you think?

Open the headline: source, AI briefing and more →

Google DeepMind launches EmbeddingGemma 2 for on-device multimodal retrieval

Google DeepMind launched EmbeddingGemma 2, a 740-million-parameter open-weight multimodal embedding model that maps text, code, images, audio, and video into a shared embedding space. Released under the Apache 2.0 license, it is designed for on-device search and retrieval, with modular encoders, an 8K-token context window, and support for reducing vector dimensions to lower storage use. Why it matters: The release could make private, offline cross-modal search and retrieval more practical on consumer hardware. What do you think?

Open the headline: source, AI briefing and more →