Google Cloud Dataproc (now “Managed Service for Apache Spark”)

by Google Cloud

Google Cloud’s managed Spark, Hadoop and related open-source data processing.

Renamed

Per Google’s docs and release notes, Dataproc and “Google Cloud Serverless for Apache Spark” were reportedly unified in April 2026 (exact date not verified) under the brand Managed Service for Apache Spark. The Dataproc API, client libraries, CLI and IAM names are unchanged. Videos and older material still say “Dataproc” and “Dataproc Serverless”.

Deployment options (as of 2026-10-05)

  1. Clusters on Compute Engine (formerly “Dataproc on Compute Engine”): from a few to hundreds of nodes.
  2. Serverless for Apache Spark (formerly Dataproc Serverless): run Spark workloads with no cluster management; also surfaced inside BigQuery.
  3. On GKE (GKE).

Components and features

  • Spark, Hadoop, Hive, Pig; optional components include Flink, Trino/Presto, Iceberg, Delta Lake, HBase, Jupyter, Ranger, Zeppelin, Zookeeper.
  • Fast cluster create/delete (docs claim cluster operations in about 90 seconds or less), autoscaling, spot/preemptible workers, initialization actions, custom images, workflow templates.
  • Integrates with Cloud Storage, BigQuery, Sub, Cloud Composer (Airflow), Vertex AI, Cloud Monitoring/Logging; compare Dataflow (Beam-based, fully managed streaming/batch).
  • Security: IAM, VPC and private clusters, encryption, Kerberos.

When to use (opinion)

Migrating on-premises Hadoop/Spark with few code changes, ephemeral ETL (data in Cloud Storage, cluster created per job), Spark-based ML preprocessing. Use Dataflow for Beam pipelines and BigQuery SQL when Spark is not required.

Billing model

Compute Engine resources plus a Dataproc management fee for clusters; serverless is billed by workload usage. See https://cloud.google.com/dataproc/pricing and the Spark serverless pricing page; no amounts recorded here.

Practices (opinion)

Keep data in Cloud Storage and clusters ephemeral; use spot workers only for stateless work; automate teardown of idle clusters; use workflow templates or Composer for repeatable pipelines.

Sources

Open items

  • Exact rename date taken from a search listing and the release-notes page title, not read in full.
  • Status of Dataproc Metastore and Dataproc Gemini features not checked.