Google Cloud Run

by Google Cloud

Fully managed serverless container platform: run stateless containers that scale to zero, for HTTP services, jobs and event-driven workloads.

Features

  • Deploy any container image; scale to zero and autoscale on requests.
  • Services (request-driven), jobs (run-to-completion) and event-driven invocation via Eventarc; Cloud Run functions are deployed on the same platform.
  • Per-revision CPU and memory settings, configurable concurrency, min/max instances.
  • GPUs (see below), IAM-based invocation, Secret Manager, Cloud Logging/Monitoring.
  • CI/CD through gcloud, Cloud Build, Artifact Registry, Terraform.

Limits verified against docs (2026-10-05)

  • Request timeout for services: default 5 minutes, maximum 60 minutes (over 15 minutes Google advises retries and reconnect tolerance).
  • Concurrency: default 80 requests per instance, configurable.
  • Memory per instance: up to 32 GiB; maximum 8 vCPU (8 vCPU needs 4-32 GiB); default memory 512 MiB.
  • GPUs: both generally available per the GPU docs - NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM) and NVIDIA L4 (24 GB); one GPU per instance; instances with GPUs can scale to zero but need instance-based billing; region lists in the docs are limited and change.

Billing model

Pay per use (CPU, memory, requests, networking), with request-based and instance-based billing modes and a free allowance. Pricing: https://cloud.google.com/run/pricing

Typical uses

Web APIs, ML inference endpoints (CPU or GPU), background workers and batch jobs, microservices needing a custom runtime. Compare GKE for complex, stateful or multi-container systems.

Deployment sketch

gcloud run deploy my-service --image=REGION-docker.pkg.dev/PROJECT/REPO/IMAGE:TAG \  
  --region=us-central1 --concurrency=80 --memory=4Gi --cpu=1 --max-instances=50  

Practices and pitfalls (author’s suggestions, opinion)

  • Handle SIGTERM for graceful shutdown; keep images small for cold starts; use min instances for latency-sensitive services.
  • Tune concurrency to the workload; single-threaded CPU-bound apps need low concurrency.
  • Keep secrets in Secret Manager, not in images.
  • Front with Cloud Load Balancing and optionally IAP for access control.

Sources

Open items

  • Free-tier allowance and AI-agent hosting features not checked; the Artifact Registry / gcr.io statement was removed as unsourced.