Qwen3.5-9B
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 3
- Microsoft launches Decision-1, a Qwen3.5-based decision-scoring model
Microsoft released Microsoft-Decision-1, a purpose-built model for fast structured decision scoring, available now in Microsoft Foundry and live on OpenRouter. Post-trained from Qwen3.5-9B, it takes a fixed set of options and returns a calibrated probability for each via a structured API call, supporting yes/no, multiple-choice and rating formats for routing, classification, prioritization, verification and agent control. Microsoft says it ranked highest in accuracy across 36 benchmarks covering nearly 150,000 withheld questions, and that its P50 latency is 35 times faster than GPT-6 Sol, with pricing at $0.042 per million input tokens and free output tokens; Microsoft plans to rebase the system on other models, including Microsoft AI and OpenAI technology. All performance figures come from Microsoft's own testing and have not been independently verified.
TestingCatalog · 🔥 36 - JetBrains releases Mellum2.1 coding model
JetBrains launched Mellum2.1, keeping Mellum2's 12B mixture-of-experts architecture with 2.5B active parameters and the Apache 2.0 license, and focusing the upgrade on agentic coding. JetBrains says reinforcement learning was expanded from a short post-training step into the main training stage, letting the model explore codebases, edit files, identify root causes of failing tests and draft and verify fixes. The model is available on Hugging Face, with GGUF builds for llama.cpp, Ollama and LM Studio and a vLLM MTP speculative-decoding component promised later.
IT之家 AI · 🔥 0 - Rapid-MLX: open-source LLM server for Apple Silicon tops GitHub Trending
Rapid-MLX is an Apache-2.0 local LLM inference server built on MLX for Apple Silicon, exposing both OpenAI (/v1/chat/completions, /v1/responses) and Anthropic Messages (/v1/messages) APIs and focusing on reliable tool calling for coding agents. The project claims up to 4× faster decoding than Apple's mlx-lm on identical weights: a 1.50× median across 18 streaming tasks with Qwen3.5-9B 4-bit on an M4 Pro Mac mini, and up to 4.27× on a whole-file code edit (46.4 to 198.4 tok/s), plus 3.0× Ollama's aggregate decode at 8 concurrent streams on Qwen3.6-35B-A3B. All figures are the project's own measurements, and much of its comparison table is drawn from other projects' own documentation.
GitHub Trending · Python · 🔥 0
Experience and discussion from the community
Share my Qwen3.5-9B experienceAsk about Qwen3.5-9B
Nobody has shared their experience with Qwen3.5-9B yet.