Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

AI News

All dates
016

llm-openai-decisions 0.1a0 plugin wraps OpenAI's Jev-style Decisions API

Simon Willison released llm-openai-decisions 0.1a0, an LLM plugin for OpenAI's new Decisions API announced at last week's DevDay; he had GPT-6 Astra read the new API docs and modeled the plugin on his existing llm-typesafe plugin for Jev. The new decision model, gpt-6-luna, accepts image input as well as text. According to the post, both models charge for input only and not output — OpenAI at 10 cents per million input tokens versus Jev's 4.2 cents — and both support yes/no, choice and score question types.

Simon Willison's Weblog·
025

Ollama v0.40.0 defaults to MLX on Apple Silicon

Ollama v0.40.0 makes MLX the default runtime for supported model architectures on Apple Silicon. The release also adds MLX support for decision models and the EmbeddingGemma 2 embedding model, with additional model families listed in the release notes.

Ollama Releases·
035

Transformers v5.19.0 adds EmbeddingGemma 2 support

Hugging Face Transformers v5.19.0 adds support for EmbeddingGemma 2, Google’s multimodal embedding model for text, images, audio, and video in a shared 768-dimensional vector space. The release also changes MoE router-logit outputs, expands expert-parallel token dispatch, updates continuous-batching attention behavior, and fixes quantized and per-layer cache handling.

Transformers Releases·
045

OpenTPU open-sources an AI accelerator built by AI

FeSens has published OpenTPU, an open-source AI accelerator whose hardware, instruction set, simulator, compiler, and host software are contained in one monorepo. The project reports running models including LFM2.5-230M, Qwen3, Qwen3.5, Gemma 4, SmolLM3, and Phi-4-mini on an FPGA card with bit-for-bit agreement between the simulator and hardware.

Hacker News · AI(100+ 分)·
054

vLLM v0.31.0 expands serving, speculation, and model support

vLLM v0.31.0 adds broad serving and model-support updates, including DeepSeek-V4.1-Flash optimizations, Model Runner V2 speculative decoding, larger-scale expert parallelism, and improved multimodal support. It also introduces the `vllm preload` weight-cache daemon for faster engine restarts and experimental CRIU-based initialized-engine snapshots, alongside several security and breaking changes.

vLLM Releases·
066

llm-mistral 0.16 adds support for reasoning models

Simon Willison released llm-mistral 0.16, which adds support for reasoning models to the LLM command-line tool. The post cites the newly released Mistral Large 4 as an example. No further details about the update or the model are given in the source.

Simon Willison's Weblog·
076

Cline pushes an open-source coding harness as usage grows 20x

Speaking at CoreWeave's Fully Connected 2026 event, Cline Bot CEO Renee Huang said the company's open-source harness is expanding beyond coding into more general knowledge work, letting users choose faster, cheaper inference when a task doesn't need a frontier model. Huang said Cline's token consumption has grown 20x over the last four months, with a turning point in May when several strong open-weight models came out, and that much of that inference runs on CoreWeave's serverless platform. Cline recently open-sourced its evaluations for open-weight agents and launched Cline Desktop, an open-source app for open-weight models.

SiliconANGLE AI·

That's everything.