Historical snapshot (verified 2026-10-01)

This note describes the Gemini 3 launch (2025-11-18). Since then: gemini-3-pro-preview was shut down 2026-03-09 and now points to gemini-3.1-pro-preview (preview since 2026-02-19; still the only Pro-tier model on the API). Gemini 3.5 Flash went GA 2026-05-19, 3.6 Flash 2026-07-21, 3.7 Flash 2026-08-13 and 3.8 Flash 2026-09-02 (current stable). Gemini 3.5 Pro never shipped (see gemini-3-dot-5-pro). Gemini 4 Argon was announced 2026-09-30 (restricted to trusted cyber defenders via the Fairwind Program; not generally available). The “Gemini CLI” mentioned in launch coverage was sunset on 2026-06-18 and replaced by Antigravity CLI (Google Developers Blog, via search result, 2026-10-01). See gemini-3-x-family and gemini-3-series.

Gemini 3.0 and Gemini 3 Pro (Nov 2025 launch)

by Google / DeepMind

Google’s foundational model with three core architectural upgrades: deeper reasoning, stronger multimodality, and 1M-token context. Foundation for Pro, Deep Think, Flash and Flash-Lite variants.

What changed in 2026

DateEvent (source: Gemini API changelog / Google blog)
2025-11-18Gemini 3 Pro (preview) launched; Gemini 3 Deep Think “coming soon” to Ultra subscribers
2025-12-03Gemini 3 Deep Think (preview, per Wikipedia)
2025-12-17Gemini 3 Flash preview (Gemini API changelog)
2026-02-19Gemini 3.1 Pro preview
2026-03-09gemini-3-pro-preview shut down, alias to 3.1 Pro preview
2026-05-19Gemini 3.5 Flash GA (Google I/O); 3.5 Pro never shipped
2026-07-21 / 08-13 / 09-02Gemini 3.6 / 3.7 / 3.8 Flash
2026-09-30Gemini 4 Argon announced, restricted

Overview

Gemini 3.0 is the foundational architecture powering Google’s Gemini 3 family. Rather than a single model, it comprises variants optimized for different use cases (Pro, Deep Think, Flash, Flash-Lite). The earlier note also listed a “Gemini 3 Ultra” tier: no Google page or API listing shows one (“Ultra” is a consumer subscription, Google AI Ultra), so it was removed.

Core Architecture

Three Key Upgrades

  1. Deeper Reasoning: System 2 thinking at speed—plan chains of complex tasks without losing track
  2. Stronger Multimodality: Native processing of text, images, audio, video, code in single transformer stack
  3. Expanded Context: 1 million-token context window with caching capabilities

Native Multimodality

  • Unified Architecture: Not separate encoders stitched together
  • Genuine Cross-Modal Reasoning: Interpret sketches and generate code, analyze videos and explain science
  • Real-Time Audio: Low-latency audio encoder and Live API for natural speech-to-speech interaction
  • Document Intelligence: Processes PDFs as visual and textual objects simultaneously
  • Video Understanding: Google’s launch post reports 87.6% Video-MMMU and 81% MMMU-Pro for Gemini 3 Pro (vendor figures)

Gemini 3.0 Model Family

Tier Breakdown

ModelRoleUse Case
Gemini 3 ProFlagship general modelMultimodal apps, agents, advanced chat
Pro Deep ThinkHigh-depth reasoning modeComplex scientific analysis, planning
Gemini 3 FlashCost-efficient, high-throughputLarge-volume consumer apps
Gemini 3 Flash-LiteLightweight variantOn-device features, efficiency
Gemini NanoOn-device lightweight (separate line; not part of the Gemini 3 launch post)Mobile, privacy-sensitive, offline; see gemini-nano-android

Strategic Positioning

  • Pro tier: Core general-purpose model
  • Deep Think mode: Configurable enhanced reasoning for complex problems
  • Flash / Flash-Lite: speed and cost efficiency

Capabilities confirmed in Google’s launch post (2025-11-18)

  • Native multimodal input (text, images, video, audio, code) with a 1M-token context window.
  • Launch figures (vendor): 1501 Elo on LMArena, 81% MMMU-Pro, 87.6% Video-MMMU.
  • Use cases named by Google: learning (papers, interactive visualisations), building (“vibe coding”, interactive web apps) and multi-step planning with consistent tool use.
  • Surfaces at launch: Gemini app, AI Mode in Search, Google AI Studio, Vertex AI and the new Google Antigravity platform; Deep Think mode to follow for Ultra subscribers.
  • Details in the earlier version of this note (zoom-and-inspect, image annotation, reduced sycophancy / prompt-injection resistance, Workspace integrations, document-intelligence claims, “Code Assist 3.0”) could not be confirmed in the launch post and were removed.

Context Window & Processing

  • 1 Million Token Window: Process entire codebases or lengthy reports
  • Caching Capabilities: Efficient handling of repeated contexts
  • Multi-Modal Processing: Text, audio, images, video, PDFs simultaneously
  • Extended Analysis: Vast datasets from diverse information sources

Competitive Position (2026-10-01)

Gemini 3.0 is no longer a frontier reference. Current rivals: Anthropic Fable 5.1 / Opus 5.5 (anthropic-claude) and OpenAI GPT-6 Astra / 6.1 Sol (gpt-5-and-gpt-6-family). Google’s own current line is Gemini 3.8 Flash on the API and, restricted, Gemini 4 Argon. For leaderboard positions use live boards, not this note.

Pricing model and availability (Gemini API pricing page, fetched 2026-10-01)

  • Gemini 3.1 Pro Preview: token-based pricing with a long-prompt surcharge above 200k tokens, unchanged from the Nov 2025 Gemini 3 Pro list price; API limits 1,048,576 input / 65,536 output tokens. Flash and Flash-Lite models are cheaper; Gemini 3.8 Flash has introductory pricing until 2026-12-31. Batch API gives 50% off. Current prices: see the vendor pricing page https://ai.google.dev/gemini-api/docs/pricing.
  • Free tier in Google AI Studio / the API for some models with rate limits; enterprise via Vertex AI.
  • Gemini 4 Argon (restricted): introductory pricing, then a higher standard rate (Google blog).
  • For production use gemini-3.1-pro-preview (preview models may change or be shut down with notice) or a GA Flash model; the 3.1 Pro model is still “preview” on the 2026-10-01 models page and this can change.

Caveats

  • Benchmark figures (LMArena 1501 Elo, MMMU-Pro 81%, Video-MMMU 87.6%) are vendor claims from Google’s launch post, re-checked 2026-10-02; a GPQA Diamond 91.9% figure was not re-confirmed and is left out. Early community “checkpoint” hints (e.g. X58/2HT) are obsolete and unverified.

Key Innovation

The shift from “a model” to “a family of models” optimized for different use cases within single architectural foundation—from Flash-Lite to Pro and Deep Think, all sharing native multimodality.

See Also

Open items

  • None blocking. The Gemini 3 Deep Think date (2025-12-03) is from Wikipedia only.

Sources (accessed 2026-10-01 and 2026-10-02)