Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

AI News

215

Transformers v5.19.0 adds EmbeddingGemma 2 support

Hugging Face Transformers v5.19.0 adds support for EmbeddingGemma 2, Google’s multimodal embedding model for text, images, audio, and video in a shared 768-dimensional vector space. The release also changes MoE router-logit outputs, expands expert-parallel token dispatch, updates continuous-batching attention behavior, and fixes quantized and per-layer cache handling.

Transformers Releases·
225

OpenTPU open-sources an AI accelerator built by AI

FeSens has published OpenTPU, an open-source AI accelerator whose hardware, instruction set, simulator, compiler, and host software are contained in one monorepo. The project reports running models including LFM2.5-230M, Qwen3, Qwen3.5, Gemma 4, SmolLM3, and Phi-4-mini on an FPGA card with bit-for-bit agreement between the simulator and hardware.

Hacker News · AI(100+ 分)·
234

vLLM v0.31.0 expands serving, speculation, and model support

vLLM v0.31.0 adds broad serving and model-support updates, including DeepSeek-V4.1-Flash optimizations, Model Runner V2 speculative decoding, larger-scale expert parallelism, and improved multimodal support. It also introduces the `vllm preload` weight-cache daemon for faster engine restarts and experimental CRIU-based initialized-engine snapshots, alongside several security and breaking changes.

vLLM Releases·
246

llm-mistral 0.16 adds support for reasoning models

Simon Willison released llm-mistral 0.16, which adds support for reasoning models to the LLM command-line tool. The post cites the newly released Mistral Large 4 as an example. No further details about the update or the model are given in the source.

Simon Willison's Weblog·
256

Cline pushes an open-source coding harness as usage grows 20x

Speaking at CoreWeave's Fully Connected 2026 event, Cline Bot CEO Renee Huang said the company's open-source harness is expanding beyond coding into more general knowledge work, letting users choose faster, cheaper inference when a task doesn't need a frontier model. Huang said Cline's token consumption has grown 20x over the last four months, with a turning point in May when several strong open-weight models came out, and that much of that inference runs on CoreWeave's serverless platform. Cline recently open-sourced its evaluations for open-weight agents and launched Cline Desktop, an open-source app for open-weight models.

SiliconANGLE AI·
263

Strata reportedly runs 125B Qwen model on 12GB GPUs

Developer Niko1221 has open-sourced the Strata engine, which reportedly runs a quantized Qwen3.8-Flash-Next model with 125 billion parameters on consumer GPUs with at least 12GB of VRAM. The engine keeps the MoE model in system RAM, loads frequently used experts into VRAM, and uses a lightweight model for speculative decoding; reported tests reached 94 tokens per second on an RTX 5070 with a 2-bit quantization.

IT之家 AI·
273

Termexo v0.10.10 adds 19 local MCP tools and agent auto-connect

MIT-licensed Windows AI coding workspace Termexo released v0.10.10, exposing its terminal, tasks, and some settings as 19 local MCP tools. These connections are automatically added to newly launched Claude Code, Codex, OpenCode, Grok Build, and Antigravity terminals in Termexo.

开源中国·
287

RemoveMacAI removes Apple Intelligence features from macOS 27

RemoveMacAI is an open-source command-line tool that removes selected or all Apple Intelligence features from macOS 27 using an approved configuration profile and Apple’s asset service. Its developer says the approach leaves System Integrity Protection enabled and avoids directly modifying /System; users can revert the changes and restore the models when features are re-enabled.

Ars Technica AI·
290

GitHub launches ReviewBench benchmark for AI code review

GitHub introduced ReviewBench, an open benchmark for evaluating AI code review agents. The benchmark covers 219 pull requests from 187 public repositories across 19 languages, uses a multi-source golden set, and supports precision, recall, severity, and category-based analysis. GitHub says its offline results have shown alignment with production experiments for GitHub Copilot code review.

GitHub Blog · AI & ML·
300

Only 11 of 26 popular Chinese 'de-AI-flavor' rules survive testing

A hospital worker who writes science explainers in his spare time reviewed 13 Chinese open-source 'de-AI-flavor' writing projects and found most of their rule lists had never been validated: only lieflat-less-ai-tone had tested anything, having five models write 60 pieces each (300 AI articles) and comparing them with 329 human-written public articles (~2.83 million characters), scoring 26 circulating criteria — just 11 held up, while 15 were refuted, including sentence-length uniformity. The author then built a Claude Code skill that tiers rules by evidence and adds hard constraints such as information conservation (every fact in the final draft must exist in the source) and no changing the strength of claims, plus a scanning script that uses only the Python standard library. On the Chinese HC3 dataset, 3.4% of 12,000-plus human answers were flagged as AI-like, versus 18.9% for GPT-3.5 answers and 85% for Claude answers and 83% for 12 science drafts.

少数派 AI·
310

SmartCall v1.0.7 adds model and voice management

SmartCall v1.0.7 is described as an update to an AI customer-service call-center system built on AI foundation models and Asterisk. The release adds model management and voice management, alongside capabilities including AI voice bots, IVR orchestration, real-time speech recognition, speech synthesis, and intent recognition.

开源中国·

That's everything.

← Previous