
Ollama
Run open-source LLMs locally on your machine (Llama, Mistral, Gemma)
From the Ollama blog
Published by Ollama, not by us. Every card opens the original post.
Ollama now supports Jev-style decision models
Ollama now supports decision models, based on TypeSafe's Jev API for fast, typed decisions. Decision models can now be run at no cost with low latency. Based on text, decision models answer yes-or-no questions, choices and scores about it, with a probability for every option.
Ollama's transparent pricing
Ollama's Pro, Max, and Team plans now use industry-standard per-token pricing with usage included on every plan.
Claude Desktop support with Ollama
Claude Desktop can now be configured to work with Ollama as a third-party gateway provider, making it possible to use open models in Claude.
NVIDIA Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is now available on Ollama. It's a 30 billion parameter (3B active) open model built for agents that stay running, gathering context, calling tools, and working through multi-step tasks on your own hardware.
Muse Glimmer from Meta Superintelligence Labs is now available
Meta's Muse Glimmer, the first open model released by Meta Superintelligence Labs, is now available. Muse Glimmer is a 30B multimodal model released under the Apache 2.0 license, designed for local coding agents, and accelerated by Ollama's MLX engine with new native DFlash and…
Ollama: all aboard open models
Serving 8.9 million developers, Ollama has raised $88M from Benchmark, Theory Ventures, 8VC, Y Combinator, and many incredible angel investors.
Faster Gemma 4 on MLX with multi-token prediction
Gemma 4 is now significantly faster in Ollama 0.31 on Apple Silicon via multi-token prediction (MTP), powered by MLX. Performance is now up to 90% faster when used with coding agents, as measured using the Aider polyglot benchmark.
Ollama's highest performance on Apple Silicon yet with MLX
Ollama's MLX engine has been updated to deliver its highest performance on Apple Silicon yet. Models output higher quality responses, respond faster, and use less memory.
Improved performance and model support with GGUF
Ollama 0.30 is now available with improved performance and GGUF model compatibility through llama.cpp. This augments Ollama's MLX engine on Apple silicon, bringing support to more models on a wider range of hardware.
NVIDIA Nemotron 3 Ultra
NVIDIA Nemotron 3 Ultra is built for high-throughput reasoning and long-running agent workflows.
OpenJarvis: a local-first personal AI is now available to run with Ollama
OpenJarvis v1.0 is now available: an open-source framework for building personal AI agents that run on your own hardware, with Ollama support built-in.
Ollama is now powered by MLX on Apple Silicon in preview
Today, we're previewing the fastest way to run Ollama on Apple silicon, powered by MLX, Apple's machine learning framework.