Get up and running with Kimi-K2.5, GLM-5, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models. https://ollama.com
Find a file
2026-08-13 15:19:02 -07:00
.github Release v0.32.7 (#17646) 2026-08-10 04:04:56 -07:00
agent agent: allow multiple edits per edit tool call (#17711) 2026-08-12 16:45:53 -07:00
anthropic anthropic: close text block before starting thinking block (#17225) 2026-07-17 10:24:38 -04:00
api openai: support web search in Responses API (#17686) 2026-08-12 11:51:54 -07:00
app launch: add DeepSeek Harness integration (#17733) 2026-08-13 15:19:02 -07:00
auth auth: fix problems with the ollama keypairs (#12373) 2025-09-22 23:20:20 -07:00
cmake mlx: enable CUDA backend in CUDA builds (#17688) 2026-08-12 07:31:00 -07:00
cmd launch: add DeepSeek Harness integration (#17733) 2026-08-13 15:19:02 -07:00
convert llama: default qwen2.5vl window attention metadata (#16868) 2026-06-24 10:35:29 -07:00
discover discover: fall back to standard CUDA when the JetPack runner is absent (#16949) 2026-07-02 08:34:35 -07:00
docs launch: add DeepSeek Harness integration (#17733) 2026-08-13 15:19:02 -07:00
envconfig runner: Remove CGO engines, use llama-server exclusively for GGML models (#16031) 2026-05-29 13:35:47 -07:00
format
fs create: harden GGUF create flows (#17062) 2026-07-06 16:20:20 -07:00
harmony runner: Remove CGO engines, use llama-server exclusively for GGML models (#16031) 2026-05-29 13:35:47 -07:00
integration Release v0.32.7 (#17646) 2026-08-10 04:04:56 -07:00
internal Merge pull request #17483 from ollama/drifkin/suggest-cloud 2026-07-31 14:28:04 -07:00
llama llama.cpp update (#17545) 2026-08-04 09:51:52 -07:00
llm Release v0.32.7 (#17646) 2026-08-10 04:04:56 -07:00
logutil runner: Remove CGO engines, use llama-server exclusively for GGML models (#16031) 2026-05-29 13:35:47 -07:00
manifest runner: Remove CGO engines, use llama-server exclusively for GGML models (#16031) 2026-05-29 13:35:47 -07:00
middleware openai: support web search in Responses API (#17686) 2026-08-12 11:51:54 -07:00
ml llama: enable FA on CUDA CC 6.x GPUs (#16994) 2026-07-02 17:11:39 -07:00
model model/renderers: match Muse Glimmer reasoning template (#17732) 2026-08-13 14:35:50 -07:00
openai launch: add Muse Code integration (#17594) 2026-08-13 13:10:29 -07:00
parser model: add Laguna v8 chat support and fix Metal inference (#17291) 2026-07-21 16:06:29 -07:00
progress progress: fix data races on ticker, states, spinner, and bar state (#17445) 2026-08-04 15:06:15 -07:00
readline runner: Remove CGO engines, use llama-server exclusively for GGML models (#16031) 2026-05-29 13:35:47 -07:00
runner imagegen: remove MLX image generation code (#16615) 2026-07-28 15:35:28 -07:00
scripts win: support CUDA on Windows ARM64 (#16931) 2026-07-21 10:53:30 -07:00
server mlx: avoid pulling MLX models when MLX is missing (#17710) 2026-08-12 14:42:17 -07:00
template runner: Remove CGO engines, use llama-server exclusively for GGML models (#16031) 2026-05-29 13:35:47 -07:00
thinking thinking: fix double emit when no opening tag 2025-08-21 21:03:12 -07:00
tokenizer tokenizer: fix multi-regex BPE offset handling (#15844) 2026-04-27 14:14:27 -07:00
tools tools: ignore braces inside JSON strings when detecting tool call end (#16937) 2026-06-27 12:00:55 -07:00
types manifests: remove OCI rootfs from the model config (#17619) 2026-08-08 19:44:57 -07:00
version
x nn: speed up prefill on double-scale nvfp4 models 2026-08-12 13:25:33 -07:00
.dockerignore
.gitattributes .gitattributes: add app/webview to linguist-vendored (#13274) 2025-11-29 23:46:10 -05:00
.gitignore create: Clean up experimental paths, fix create from existing safetensor model (#14679) 2026-04-07 08:12:57 -07:00
.golangci.yaml ci: restore previous linter rules (#13322) 2025-12-03 18:55:02 -08:00
AGENTS.md Add AGENTS.md and CLAUDE.md to root repository (#16604) 2026-06-07 10:57:59 -07:00
CLAUDE.md Add AGENTS.md and CLAUDE.md to root repository (#16604) 2026-06-07 10:57:59 -07:00
CMakeLists.txt runner: Remove CGO engines, use llama-server exclusively for GGML models (#16031) 2026-05-29 13:35:47 -07:00
CMakePresets.json runner: Remove CGO engines, use llama-server exclusively for GGML models (#16031) 2026-05-29 13:35:47 -07:00
CONTRIBUTING.md docs: fix typos in repository documentation (#10683) 2025-11-15 20:22:29 -08:00
Dockerfile imagegen: remove MLX image generation code (#16615) 2026-07-28 15:35:28 -07:00
go.mod agent/tui: stream thinking traces (#17611) 2026-08-07 13:11:45 -07:00
go.sum server: remove OLLAMA_EXPERIMENT=client2 (#16962) 2026-07-06 13:15:39 -07:00
LICENSE
LLAMA_CPP_VERSION llama.cpp bump (#17702) 2026-08-12 12:10:18 -07:00
main.go
MLX_C_VERSION Update MLX and MLX-C with threading fixes (#15845) 2026-05-03 10:03:14 -07:00
MLX_VERSION MLX update (#17704) 2026-08-12 12:09:50 -07:00
README.md launch: add DeepSeek Harness integration (#17733) 2026-08-13 15:19:02 -07:00
SECURITY.md docs: fix typos in repository documentation (#10683) 2025-11-15 20:22:29 -08:00

ollama

Ollama

Start building with open models.

Download

macOS

curl -fsSL https://ollama.com/install.sh | sh

or download manually

Windows

irm https://ollama.com/install.ps1 | iex

or download manually

Linux

curl -fsSL https://ollama.com/install.sh | sh

Manual install instructions

Docker

The official Ollama Docker image ollama/ollama is available on Docker Hub.

Libraries

Community

Get started

ollama

You'll be prompted to run a model or connect Ollama to your existing agents or applications such as Claude Code, OpenClaw, OpenCode , Codex, Copilot, and more.

Coding

To launch a specific integration:

ollama launch claude

Supported integrations include Claude Code, Codex, Copilot CLI, DeepSeek Harness, Droid, and OpenCode.

AI assistant

Use OpenClaw to turn Ollama into a personal AI assistant across WhatsApp, Telegram, Slack, Discord, and more:

ollama launch openclaw

Chat with a model

Run and chat with Gemma 4:

ollama run gemma4

See ollama.com/library for the full list.

See the quickstart guide for more details.

REST API

Ollama has a REST API for running and managing models.

curl http://localhost:11434/api/chat -d '{
  "model": "gemma4",
  "messages": [{
    "role": "user",
    "content": "Why is the sky blue?"
  }],
  "stream": false
}'

See the API documentation for all endpoints.

Python

pip install ollama
from ollama import chat

response = chat(model='gemma4', messages=[
  {
    'role': 'user',
    'content': 'Why is the sky blue?',
  },
])
print(response.message.content)

JavaScript

npm i ollama
import ollama from "ollama";

const response = await ollama.chat({
  model: "gemma4",
  messages: [{ role: "user", content: "Why is the sky blue?" }],
});
console.log(response.message.content);

Supported backends

  • llama.cpp project founded by Georgi Gerganov.

Documentation

Community Integrations

Want to add your project? Open a pull request.

Chat Interfaces

Web

Desktop

  • Dify.AI - LLM app development platform
  • AnythingLLM - All-in-one AI app for Mac, Windows, and Linux
  • Maid - Cross-platform mobile and desktop client
  • Witsy - AI desktop app for Mac, Windows, and Linux
  • Cherry Studio - Multi-provider desktop client
  • Ollama App - Multi-platform client for desktop and mobile
  • PyGPT - AI desktop assistant for Linux, Windows, and Mac
  • Alpaca - GTK4 client for Linux and macOS
  • SwiftChat - Cross-platform including iOS, Android, and Apple Vision Pro
  • Enchanted - Native macOS and iOS client
  • RWKV-Runner - Multi-model desktop runner
  • Ollama Grid Search - Evaluate and compare models
  • macai - macOS client for Ollama and ChatGPT
  • AI Studio - Multi-provider desktop IDE
  • Reins - Parameter tuning and reasoning model support
  • ConfiChat - Privacy-focused with optional encryption
  • LLocal.in - Electron desktop client
  • MindMac - AI chat client for Mac
  • Msty - Multi-model desktop client
  • BoltAI for Mac - AI chat client for Mac
  • IntelliBar - AI-powered assistant for macOS
  • Kerlig AI - AI writing assistant for macOS
  • Hillnote - Markdown-first AI workspace
  • Perfect Memory AI - Productivity AI personalized by screen and meeting history

Mobile

SwiftChat, Enchanted, Maid, Ollama App, Reins, and ConfiChat listed above also support mobile platforms.

Code Editors & Development

Libraries & SDKs

Frameworks & Agents

RAG & Knowledge Bases

  • RAGFlow - RAG engine based on deep document understanding
  • R2R - Open-source RAG engine
  • MaxKB - Ready-to-use RAG chatbot
  • Minima - On-premises or fully local RAG
  • Chipper - AI interface with Haystack RAG
  • ARGO - RAG and deep research on Mac/Windows/Linux
  • Archyve - RAG-enabling document library
  • Casibase - AI knowledge base with RAG and SSO
  • BrainSoup - Native client with RAG and multi-agent automation

Bots & Messaging

Terminal & CLI

Productivity & Apps

Observability & Monitoring

  • Opik - Debug, evaluate, and monitor LLM applications
  • OpenLIT - OpenTelemetry-native monitoring for Ollama and GPUs
  • Lunary - LLM observability with analytics and PII masking
  • Langfuse - Open source LLM observability
  • HoneyHive - AI observability and evaluation for agents
  • MLflow Tracing - Open source LLM observability

Database & Embeddings

Infrastructure & Deployment

Cloud

Package Managers