Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

Rapid-MLX

AI overview

Sign in and the AI will write an overview from our coverage.

Headlines · 1

  1. Rapid-MLX: open-source LLM server for Apple Silicon tops GitHub Trending

    Rapid-MLX is an Apache-2.0 local LLM inference server built on MLX for Apple Silicon, exposing both OpenAI (/v1/chat/completions, /v1/responses) and Anthropic Messages (/v1/messages) APIs and focusing on reliable tool calling for coding agents. The project claims up to 4× faster decoding than Apple's mlx-lm on identical weights: a 1.50× median across 18 streaming tasks with Qwen3.5-9B 4-bit on an M4 Pro Mac mini, and up to 4.27× on a whole-file code edit (46.4 to 198.4 tok/s), plus 3.0× Ollama's aggregate decode at 8 concurrent streams on Qwen3.6-35B-A3B. All figures are the project's own measurements, and much of its comparison table is drawn from other projects' own documentation.

    GitHub Trending · Python · 🔥 11

Experience and discussion from the community

Share my Rapid-MLX experienceAsk about Rapid-MLX

Nobody has shared their experience with Rapid-MLX yet.