Apache Iceberg
What it is
An open table format for large analytic tables on object storage. It adds ACID transactions, schema and partition evolution, time travel and snapshot isolation on top of Parquet (also Avro/ORC) files, so that several engines can read and write the same table safely. It is the de facto interchange format of the 2025-26 “open lakehouse”.
Maker, ownership and history
- Apache Software Foundation top-level project; licence Apache 2.0. Origin and incubation history is not re-verified here (see Open items).
- Databricks acquired Tabular (company founded by Iceberg’s creators) in 2024, reported at over $1 billion (Wikipedia, secondary).
- Latest release seen: 1.12.0, 2026-09-30 (GitHub release): Variant type support, Flink equality-delete conversion, Spark 4.2 support, REST catalog improvements.
Editions/deployment
Not a product but a spec plus libraries (Java, Python, Rust, Go, C++). Used through engines (Spark, Flink, Trino, DuckDB, Snowflake, Databricks, BigQuery, Athena…). No pricing; cost sits in the engine and storage.
Core architecture
Metadata tree: table metadata JSON → manifest lists → manifests → data files; catalog holds the pointer to the current metadata file. Spec versions (iceberg.apache.org/spec, checked 2026-10-07):
- v1: analytic tables on immutable files.
- v2: row-level deletes (position/equality delete files).
- v3: nanosecond timestamps,
variant,geometry,geography,unknowntypes; column default values; binary deletion vectors; table encryption; multi-argument transforms. Row lineage is also in v3 (per AWS S3 Tables docs). - v4: “under active development”, not adopted.
REST catalog protocol: an HTTP API for namespaces, tables, commits and credential vending; any engine that speaks it can use any compliant catalog. See Open lakehouse catalogs.
Role in an enterprise AI rollout
One governed copy of data readable by SQL engines, Spark/ML training jobs and agents; Variant helps with semi-structured/JSON-like agent logs and events. Avoids copying data between warehouse and ML platforms.
Vendor adoption (as of 2026-10-07, from vendor docs)
- Snowflake: Iceberg tables with Snowflake-managed or customer-managed storage; Snowflake or external catalogs (Glue, Unity Catalog, REST); supports spec v1, v2 and v3 (some limits on equality deletes). See Snowflake AI Data Cloud.
- Databricks: managed Iceberg tables in Unity Catalog and foreign tables; v1-v3; needs serverless for managed tables, DBR 16.4 LTS+. See Unity Catalog.
- AWS: S3 Tables store Iceberg tables in table buckets, v3 supported, integrated with Glue Data Catalog, Athena, Redshift, EMR. See AWS data and AI platform.
- Google: BigQuery “managed tables for Apache Iceberg” keep data in customer Cloud Storage; exports Iceberg V2 snapshots for external engines. See BigQuery.
- Microsoft: OneLake virtualizes Delta tables as Iceberg (V2 metadata) and Iceberg as Delta; Iceberg V3 only partially. See OneLake and Fabric.
AI features
Format-level: Variant type, v3 deletion vectors. No AI features of its own; AI sits in engines and catalogs. DuckDB, Trino and Dremio read it (Trino and Starburst, Dremio, DuckDB).
Integrations
Catalogs: Polaris, Unity, Lakekeeper, Gravitino, Glue. Interop with Delta via UniForm and XTable-style virtualization: see Delta Lake; other formats in Hudi and Paimon.
Strengths and weaknesses (opinion)
- Strongest multi-vendor support of any lakehouse format; engine-neutral spec.
- Many small metadata files; compaction and maintenance needed (managed by S3 Tables, Snowflake, Databricks predictive optimization).
- Feature uptake differs by engine: v3 coverage is uneven (see Fabric limits).
Self-learning
- Docs and spec: https://iceberg.apache.org/ (free)
- Spec: https://iceberg.apache.org/spec/ (free)
- Snowflake Iceberg docs: https://docs.snowflake.com/en/user-guide/tables-iceberg (free)
- Databricks Iceberg docs: https://docs.databricks.com/aws/en/iceberg/ (free)
- AWS S3 Tables: https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tables.html (free)
- OneLake Iceberg: https://learn.microsoft.com/en-us/fabric/onelake/onelake-iceberg-tables (free)
Sources (fetched 2026-10-07)
- https://iceberg.apache.org/spec/ and https://raw.githubusercontent.com/apache/iceberg/main/format/spec.md
- https://api.github.com/repos/apache/iceberg/releases
- Vendor docs listed under Self-learning
- https://docs.cloud.google.com/bigquery/docs/iceberg-tables
- https://en.wikipedia.org/wiki/Databricks (secondary, Tabular)
Open items
- Incubation/graduation dates and original Netflix origin not verified this session.
- Google BigQuery v3 support and Azure Iceberg write support not confirmed.