Google Cloud Run
by Google Cloud
Fully managed serverless container platform: run stateless containers that scale to zero, for HTTP services, jobs and event-driven workloads.
Features
- Deploy any container image; scale to zero and autoscale on requests.
- Services (request-driven), jobs (run-to-completion) and event-driven invocation via Eventarc; Cloud Run functions are deployed on the same platform.
- Per-revision CPU and memory settings, configurable concurrency, min/max instances.
- GPUs (see below), IAM-based invocation, Secret Manager, Cloud Logging/Monitoring.
- CI/CD through gcloud, Cloud Build, Artifact Registry, Terraform.
Limits verified against docs (2026-10-05)
- Request timeout for services: default 5 minutes, maximum 60 minutes (over 15 minutes Google advises retries and reconnect tolerance).
- Concurrency: default 80 requests per instance, configurable.
- Memory per instance: up to 32 GiB; maximum 8 vCPU (8 vCPU needs 4-32 GiB); default memory 512 MiB.
- GPUs: both generally available per the GPU docs - NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM) and NVIDIA L4 (24 GB); one GPU per instance; instances with GPUs can scale to zero but need instance-based billing; region lists in the docs are limited and change.
Billing model
Pay per use (CPU, memory, requests, networking), with request-based and instance-based billing modes and a free allowance. Pricing: https://cloud.google.com/run/pricing
Typical uses
Web APIs, ML inference endpoints (CPU or GPU), background workers and batch jobs, microservices needing a custom runtime. Compare GKE for complex, stateful or multi-container systems.
Deployment sketch
gcloud run deploy my-service --image=REGION-docker.pkg.dev/PROJECT/REPO/IMAGE:TAG \
--region=us-central1 --concurrency=80 --memory=4Gi --cpu=1 --max-instances=50 Practices and pitfalls (author’s suggestions, opinion)
- Handle SIGTERM for graceful shutdown; keep images small for cold starts; use min instances for latency-sensitive services.
- Tune concurrency to the workload; single-threaded CPU-bound apps need low concurrency.
- Keep secrets in Secret Manager, not in images.
- Front with Cloud Load Balancing and optionally IAP for access control.
Sources
- Request timeout: https://docs.cloud.google.com/run/docs/configuring/request-timeout (fetched 2026-10-05)
- Concurrency: https://docs.cloud.google.com/run/docs/configuring/concurrency (fetched 2026-10-05)
- Memory/CPU limits: https://docs.cloud.google.com/run/docs/configuring/services/memory-limits (fetched 2026-10-05)
- GPU support: https://docs.cloud.google.com/run/docs/configuring/services/gpu (fetched 2026-10-05)
Open items
- Free-tier allowance and AI-agent hosting features not checked; the Artifact Registry / gcr.io statement was removed as unsourced.