Open-Source AI

Open-source AI means an AI system whose parts can be used, studied, modified and shared by anyone. The term is contested: most popular “open” models are only open-weight (downloadable weights, often with a restrictive licence, without training data or code).

Definitions: open source vs open weights

  • The Open Source Initiative’s Open Source AI Definition (OSAID) 1.0 (announced at All Things Open, Oct 2024) requires the freedoms to use, study, modify and share, access to the “preferred form for making modifications”: training/inference code under an OSI-approved licence, weights under free terms, and enough information on training data for a skilled person to recreate a substantially equivalent system. The Software Freedom Conservancy criticised it for not requiring the training data itself.
  • Open-weight models publish weights only. Under OSAID most do not qualify.
  • Meta’s Llama licence is not OSI-approved (per the OSI forum discussion cited below: usage restrictions such as a licence request above 700M monthly users, naming/attribution terms, and a restriction on Llama 4’s multimodal use in the EU).
  • Truly “fully open” example: Ai2’s OLMo 3 (Nov 2025; 7B and 32B) releases weights, Dolma 3 training data, code and checkpoints under Apache 2.0.

Licences commonly seen

LicencePermissivePatent grantCommercial useCopyleft
Apache 2.0yesyesyesno
MITyesnoyesno
GPL v3noimpliedyesyes
OpenRAIL familymedium (use-based restrictions)yesyesno
Llama / Gemma-style custom licencesrestrictedvariesconditionalno
Always read each model card: licences differ per model and per version.

Landscape (2026)

  • Hugging Face Hub passed ~3 million public models in Aug 2026 (counter read 3,012,377 in late August, per a Hugging Face community blog post, 2026) up from ~2 million in Aug 2025. See hugging-face.
  • Google Gemma 4: see gemma-4 for its licence (press reports say Apache 2.0; Google’s docs mention a “Gemma 4 license”; check the model card).
  • Meta: frontier work moved to closed Muse models; Llama 4 (2025) is the newest Llama; see Meta Muse and Llama.
  • China-based families dominate open weights: qwen (Alibaba), deepseek, GLM/Z.ai, Kimi, MiniMax (see the notes in ai-models/open-source/).
  • OpenAI gpt-oss (2025, Apache 2.0) is the only OpenAI open-weight release confirmed in this vault.
  • Per UNU’s summary of the Stanford AI Index 2026 (c3.unu.edu), the closed-vs-open performance gap is roughly 3%.
  • Specialised: qwen3-tts (speech), whisper-and-asr-models (ASR); Stable Diffusion/FLUX-style image models have their own licences (check each).

Deployment

  • Local: ollama and similar runtimes, quantised models on consumer GPUs.
  • Self-hosted serving: vllm on cloud GPUs or Kubernetes.
  • Managed: Hugging Face inference providers and other hosts (see ai-api-platforms).
  • Customisation: quantisation, distillation, fine-tuning (low-rank-adaptation).

Advantages and risks

Advantages: no per-token licence cost, privacy and data control, vendor independence, inspectability, customisation, reproducible research.
Risks: no warranty or support contract, uneven quality, hardware and expertise needs, licence traps (non-commercial or use-restricted terms), and weaker or removable safety guardrails on open weights.

large-language-model, neural-networks, machine-learning.

Limits

The Hugging Face model count and the Stanford AI Index gap are attributed to secondary summaries; Llama licence terms are attributed to the OSI forum, not the licence text. Always check each model card.

Sources