创建

登录 ReadmeX

登录后可以加入社区、发帖、投票和聊天。

还没有账号?

资讯

Transformers v5.19.0支持EmbeddingGemma 2

AI 总结

Hugging Face Transformers v5.19.0 新增对 EmbeddingGemma 2 的支持。这款 Google 多模态嵌入模型可将文本、图像、音频和视频编码到共享的 768 维向量空间;本次版本还更新了 MoE 路由器输出、专家并行、连续批处理注意力机制和量化缓存处理。

为什么重要:该版本让新的多模态嵌入模型进入常用开源框架,并同步更新分布式推理能力。

Hugging Face TransformersHugging FaceGoogle

5
来源原文Transformers Releases · 约 5 分钟读完

Release v5.19.0

New Model additions

EmbeddingGemma2

image

EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture. It encodes text, images, audio, and video, individually or combined in one input, into a shared 768-dimensional vector space for cross-modal retrieval, semantic similarity, clustering, and classification. It uses Matryoshka Representation Learning, so embeddings can be truncated to 512, 256, or 128 dimensions. It also offers configurable visual and video token budgets, and unused vision or audio towers can be disabled at load time to save memory.

Links: Documentation

Breaking changes

All MoE models whose routers compute logits now return them when output_router_logits=True, following the Qwen3-MoE pattern (a router_logits recorder on the base model, MoeModelOutputWithPast from the backbone, and a MoE causal LM output from the head), so code that relied on the previous outputs or their absence should read the router logits from these output classes.

  • 🚨 Return router logits from every MoE model that computes them (#48920) by @qgallouedec

Owlv2ForObjectDetection.embed_image_query now selects the query box with the highest objectness score, as in the original OWLv2 notebook, instead of the OWL-ViT heuristic, so image-guided query embeddings and detections may differ from earlier releases.

The "paged|" prefix for SDPA and flash attention implementations is deprecated, so users should set the regular attention implementation (e.g. sdpa or flash_attention_2) for continuous batching instead of paged|sdpa or paged|flash_attention_2.

  • 🚨 Attention 🚨 Deprecate "paged|" prefix for SDPA and flash (#49112) by @remi-or

The regular flash and SDPA attention functions (flash_attention.py, sdpa_attention.py) now support continuous batching directly, and "paged|..." implementations for these are redirected to them, while eager still requires the "paged|eager" prefix.

  • 🚨 Attention 🚨 Make regular attention support CB (#49101) by @remi-or

In continuous batching, the cache update for the index-based and block-table paths is now fused into a single call, which slightly changes the cache update function's behavior and affects any custom code that calls the separate update paths.

  • 🚨 [CB] 🚨 Fuse update for index and block table path (#49088) by @remi-or

Continuous batching internals changed in preparation for removing "paged": `max

  • 🚨 [CB] 🚨 Little fixes before removing "paged" (#49069) by @remi-or

Parallelization

Expert parallelism gains a token-dispatch implementation, selected via the new ep_dispatch_experts plan rule and now the default for Qwen3 MoE and Mellum, which removes the requirement that EP size equal TP size. The Trainer was also adapted to work with expert parallelism, and the docs now note that PEFT adapters support tensor parallelism. A CI-related fix for pipeline-parallel chart2table inference was also included.

Cache

This release fixes quantized cache handling: generate no longer mutates the user's cache_config, and QuantizedLayer.reorder_cache is repaired. It also adds per-layer cache configuration, so DynamicCache and StaticCache initialize each layer from its own config (sliding window, attention chunk size, conv states, and attention head counts) to better support heterogeneous models. Separately, the deprecation cycle on mask and cache

Bugfixes and improvements

Significant community contributions

The following contributors have made significant changes to the library over the last release:

  • @vasqu
  • @remi-or
    • [Fix] Remove old test file with two ancient tests (#49280)
    • [CB] Make CB more device agnostic and add support for XPU (#49156)
    • 🚨 Attention 🚨 Deprecate "paged|" prefix for SDPA and flash (#49112)
    • 🚨 Attention 🚨 Make regular attention support CB (#49101)
    • 🚨 [CB] 🚨 Fuse update for index and block table path (#49088)
    • [Refactor] Make some flash-attention utils more readable (#49071)
    • 🚨 [CB] 🚨 Little fixes before removing "paged" (#49069)
阅读原始来源 →

前因后果

  1. AI硬件正在重定义什么算录制The Verge AI · Google
  2. OpenTPU开源AI加速器可在FPGA运行模型Hacker News · AI(100+ 分) · Gemma 4
  3. 保险商警惕AI智能体责任索赔The Decoder · Hugging Face
  4. 谷歌发布 Nano Banana 2.1:出图价减半Google AI Studio · Google
  5. 开放权重模型的网络风险争论已经失焦Interconnects · Hugging Face
  6. Gamma 5 发布:让 AI 演示文稿摆脱“AI 味”SiliconANGLE AI · Google

评论

我用过:分享经验 我怎么看:发表观点
你觉得这条新闻有多重要?还没有评分

还没有评论,来说说你的看法。