Ollama
by Ollama (plain text, no company note)
Open-source tool to run LLMs locally (and optionally on Ollama’s cloud) behind a simple CLI and REST API.
See https://ollama.com (the old ollama.ai domain in earlier notes now maps to ollama.com). Repo: https://github.com/ollama/ollama (MIT licence, 182,000+ stars).
Features (2026-09-30)
- Run open models locally, for example
ollama run gemma4; library covers Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and others. - CLI, REST API, Python and JavaScript libraries; Docker image
ollama/ollama. - OpenAI-compatible API:
http://localhost:11434/v1/supports chat completions (streaming, JSON mode, vision, tool calling), completions, embeddings and a non-stateful Responses endpoint. Not supported: logprobs, stateful responses, image URLs (base64 only). - Desktop app for macOS and Windows.
- Cloud models at
https://ollama.com/v1with US, Europe and Singapore servers; Ollama states prompts are not tracked or trained on. ollama launch claudewires Claude Code to Ollama models; integrations include VS Code and n8n. Model packaging through Modelfiles.- Install: installer from ollama.com,
irm https://ollama.com/install.ps1 | iex(Windows),curl -fsSL https://ollama.com/install.sh | sh(macOS/Linux).
Pricing
Local use is free. Cloud plans: Free, Pro (includes usage credits); higher Max tier exists (details not confirmed).
Environment variables
See set-ollama-env-vars-in-powershell for OLLAMA_HOST and OLLAMA_ORIGINS on Windows.
Related
windows-11-wsl-ubuntu-setup-guide, model-context-protocol, Ollama MCP Support
Sources
- https://ollama.com (accessed 2026-09-30)
- https://github.com/ollama/ollama (accessed 2026-09-30)
- https://docs.ollama.com/api/openai-compatibility (accessed 2026-09-30)