Ollama v0.40.0 defaults to MLX on Apple Silicon
Ollama v0.40.0 makes MLX the default runtime for supported model architectures on Apple Silicon. The release also adds MLX support for decision models and the EmbeddingGemma 2 embedding model, with additional model families listed in the release notes.
Why it matters: Apple Silicon users may get improved compatibility and performance for supported models without changing their configuration.
v0.40.0
LatestChoose a tag to compare
Sorry, something went wrong.
FilterLoading
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
No results found
github-actions
released this
25 Sep 03:31
· 2 commits to main since this release
This commit was created on GitHub.com and signed with GitHub’s verified signature.
GPG key ID: B5690EEEBB952194
Verified
Learn about vigilant mode.What's Changed
Models run on MLX on Apple Silicon by default
In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.
ollama pull qwen3.8
ollama run qwen3.8
Additional models include gemma4, qwen3.6 and qwen3.5
Decision models are now available on MLX as well: Nimble tev1 clef clef-flash
MLX now has support for an embedding model: embeddinggemma-2
We will continue testing and enabling additional models.
Full Changelog: v0.35.1...v0.40.0
Assets 19
Loading
Uh oh!
There was an error while loading. Please reload this page.
38 people reacted
How we got here
- Transformers v5.19.0 adds EmbeddingGemma 2 supportTransformers Releases · EmbeddingGemma 2
- Google releases EmbeddingGemma 2, a 740M on-device multimodal embedding modelGoogle Developers Blog · EmbeddingGemma 2