MLX
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 2
- Rapid-MLX: open-source LLM server for Apple Silicon tops GitHub Trending
Rapid-MLX is an Apache-2.0 local LLM inference server built on MLX for Apple Silicon, exposing both OpenAI (/v1/chat/completions, /v1/responses) and Anthropic Messages (/v1/messages) APIs and focusing on reliable tool calling for coding agents. The project claims up to 4× faster decoding than Apple's mlx-lm on identical weights: a 1.50× median across 18 streaming tasks with Qwen3.5-9B 4-bit on an M4 Pro Mac mini, and up to 4.27× on a whole-file code edit (46.4 to 198.4 tok/s), plus 3.0× Ollama's aggregate decode at 8 concurrent streams on Qwen3.6-35B-A3B. All figures are the project's own measurements, and much of its comparison table is drawn from other projects' own documentation.
GitHub Trending · Python · 🔥 11 - Ollama v0.40.0 defaults to MLX on Apple Silicon
Ollama v0.40.0 makes MLX the default runtime for supported model architectures on Apple Silicon. The release also adds MLX support for decision models and the EmbeddingGemma 2 embedding model, with additional model families listed in the release notes.
Ollama Releases · 🔥 5
Experience and discussion from the community
Share my MLX experienceAsk about MLX
Nobody has shared their experience with MLX yet.