Ollama runs open-weight language models such as Llama, Gemma and Qwen locally on Windows and serves them through a local API. Version 0.35.0, open source under MIT.
Rate Ollama
Average 0.0 · ratings: 0
Thank you! Your rating is saved; you can change it. Could not save the rating. Please try again later. Too many ratings from your network today. Please try again tomorrow. The average appears after 5 ratings.
About Ollama
Ollama 0.35.0 for Windows is a free, open-source tool for running large language models on your own computer instead of in someone else's cloud. You download a model once, and from then on prompts and answers stay on the PC. The Windows download on this page is the official installer from ollama.com, always pointing at the latest release; 0.35.0 came out on 28 September 2026.
What Ollama does
Think of Ollama as a model runner with two faces. On one side there is a command line: ollama run followed by a model name pulls the model and opens a chat right in the terminal. On the other there is a small desktop app for people who would rather click than type. Behind both sits a local HTTP server, and that server is the reason many developers install Ollama in the first place: other programs on the same machine can send it requests, including through an OpenAI-compatible endpoint, so tools written for cloud APIs can often be pointed at a local model with a change of address. Popular front ends such as Open WebUI and AnythingLLM connect to it this way.
The models themselves are not part of the installer. Ollama fetches open-weight models — Llama, Gemma, Qwen, DeepSeek and gpt-oss are typical examples — from its library when you ask for them, and each one comes with its own licence and its own appetite for disk space and memory.
System requirements
- Windows: Windows 10 22H2 or later; Ollama is also available for macOS and Linux.
- GPU (optional): NVIDIA cards through CUDA with driver 551.61 or newer, AMD cards through ROCm, and other hardware through Vulkan.
- No GPU: models run on the CPU, just more slowly.
- Disk: the installer, OllamaSetup.exe, is about 1.5 GB because it bundles the GPU libraries; every model you pull needs additional space.
Memory is the real limit. A model has to fit into RAM, or into video memory if you want the GPU to do the work, so the size of the model you can use depends on your machine rather than on Ollama.
Getting started
- Download OllamaSetup.exe with the button on this page and run it; no administrator setup wizard with dozens of options is involved.
- Open a terminal and start a model, for example with
ollama run gemma3. The first run downloads the weights. - Chat in the terminal, or open the desktop app and pick the model there.
- If you want other apps to use it, point them at the local server that Ollama starts in the background.
The Windows notes in the official documentation cover GPU setup, where models are stored and how to move that folder to another drive.
Is Ollama safe to download?
The installer is signed by Ollama Inc., and the source code is public in the ollama/ollama repository on GitHub, where every release is listed. Because prompts are processed locally, nothing you type is sent to a model provider. The usual caution still applies to models: download them through Ollama's own library rather than from random file shares.
Ollama or something else?
Ollama suits people who like the terminal or need a local API for their own scripts and tools. If you prefer a complete chat window with a model browser, LM Studio or Jan are closer to that. GPT4All focuses on chatting with your own documents, and AnythingLLM bundles documents, agents and several model back ends in one workspace — and can use Ollama as its engine.
Pros and cons
- Pro: open source under the MIT licence, free to use.
- Pro: one local API that many other apps already support.
- Pro: works on NVIDIA, AMD and Vulkan GPUs, and on the CPU alone.
- Con: the terminal is still the main way to manage models.
- Con: large download, and models add far more on top.