Microsoft Foundry Local

by Microsoft

End-to-end local AI runtime and SDK that ships inside your application: model download, hardware acceleration and inference run on the user’s device, with no Azure subscription required.

Status (2026): generally available since 2026-04-09 (announced on the Microsoft Foundry devblog). Announced as a preview at Build 2025. The original 2025 note framed it as a CLI plus localhost REST server in preview; that is no longer the core product (see below).

What it is now

  • A native library (.dll/.so/.dylib) loaded in-process by your app through SDKs for C#, JavaScript, Python and Rust (pip install foundry-local-sdk, npm install foundry-local-sdk, dotnet add package Microsoft.AI.Foundry.Local, cargo add foundry-local-sdk). Runtime adds roughly 20 MB to an app package.
  • Inference by ONNX Runtime with automatic execution-provider selection: NVIDIA CUDA (Windows/Linux), WebGPU via Dawn (Windows/Linux/macOS incl. Apple silicon via Metal), AMD Vitis and Qualcomm NPUs and Intel OpenVINO (Windows), CPU fallback. On Windows it integrates with Windows ML.
  • A curated, hardware-optimised model catalog (chat: e.g. GPT OSS, Qwen, DeepSeek, Mistral, Phi; audio transcription: Whisper). Models download on first use, are cached locally and are versioned. Own models can be compiled to ONNX.
  • Optional OpenAI-compatible local web server (supports the Responses API format) for tools such as LangChain or Open WebUI; for embedded apps the SDK runs in-process without a server.
  • Platforms: Windows, macOS (Apple silicon), Linux. Licence: SDK MIT; the CLI (still labelled public preview in the repo README) is under Microsoft Software License Terms; models carry their own licences.
  • Prompts and outputs stay on the device; the network is used for model/execution-provider downloads and optional diagnostics. No per-token costs.

Limits

  • Designed for single-user, hardware-constrained devices; Microsoft’s own FAQ says to use vLLM or Triton for multi-user serving (vllm).
  • Catalog is intentionally curated, not “any model”.
  • A separate “Foundry Local on Azure Local” offering targets Kubernetes/Arc edge deployments (reported public preview April 2026; not verified in detail here).

Compare

ollama (general-purpose local runner, default port 11434) is the closest alternative; Foundry Local targets apps that bundle the runtime. Cloud sibling: microsoft-foundry-note.

Open items

None blocking. Not verified: the 2025 claims removed from this note (sub-50 ms latency, role-based access, audit logging, cryptographic model verification, foundry cache list command, ~/.foundry cache path, and the localhost:11434 endpoint, which is Ollama’s port and was wrong).

Sources