Ollama

by Ollama (plain text, no company note)

Open-source tool to run LLMs locally (and optionally on Ollama’s cloud) behind a simple CLI and REST API.

See https://ollama.com (the old ollama.ai domain in earlier notes now maps to ollama.com). Repo: https://github.com/ollama/ollama (MIT licence, 182,000+ stars).

Features (2026-09-30)

  • Run open models locally, for example ollama run gemma4; library covers Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and others.
  • CLI, REST API, Python and JavaScript libraries; Docker image ollama/ollama.
  • OpenAI-compatible API: http://localhost:11434/v1/ supports chat completions (streaming, JSON mode, vision, tool calling), completions, embeddings and a non-stateful Responses endpoint. Not supported: logprobs, stateful responses, image URLs (base64 only).
  • Desktop app for macOS and Windows.
  • Cloud models at https://ollama.com/v1 with US, Europe and Singapore servers; Ollama states prompts are not tracked or trained on.
  • ollama launch claude wires Claude Code to Ollama models; integrations include VS Code and n8n. Model packaging through Modelfiles.
  • Install: installer from ollama.com, irm https://ollama.com/install.ps1 | iex (Windows), curl -fsSL https://ollama.com/install.sh | sh (macOS/Linux).

Pricing

Local use is free. Cloud plans: Free, Pro (includes usage credits); higher Max tier exists (details not confirmed).

Environment variables

See set-ollama-env-vars-in-powershell for OLLAMA_HOST and OLLAMA_ORIGINS on Windows.

windows-11-wsl-ubuntu-setup-guide, model-context-protocol, Ollama MCP Support

Sources