Apache Hudi and Apache Paimon
Short comparison of the two other Apache lake table formats next to Iceberg and Delta Lake.
| Hudi | Paimon | |
|---|---|---|
| Description (project site) | “open data lakehouse platform” on an open table format with database functionality | Apache project; per its docs a data lake platform unifying batch, streaming and multimodal AI workloads, built on LSM-tree technology (time travel, schema evolution, vector search) |
| Origin | Developed at Uber (hudi.apache.org) | Origin (Flink Table Store) not verified here |
| Strengths | Record-level upserts/deletes, incremental processing, Copy-on-Write and Merge-on-Read (per project docs, not individually checked) | LSM-based streaming updates plus analytics (per Paimon docs) |
| Latest release seen | 1.2.1 (2026-09-24, GitHub API); 0.14.2 (2026-06-08) | 2.0.0 (2026-08-07, GitHub), with PyPaimon 2.0.0 |
Notes from release pages: Hudi 1.2.1 release notes (GitHub, checked 2026-10-07): Trino-Hudi connector moved into the Hudi repo (org.apache.hudi:hudi-trino, built for Trino 483), Flink 2.1 backports and a variant-type adapter for Flink, hoodie.client.heartbeat.tolerable.misses default changed from 2 to 10.
Role for AI
In the author’s view (opinion, not sourced): the formats suit streaming/CDC upsert pipelines feeding feature or RAG tables, and may matter less than Iceberg or Delta unless the stack is Flink- or Hudi-centred.
Self-learning
- https://hudi.apache.org/ (free)
- https://paimon.apache.org/ (free)
Sources (fetched 2026-10-07)
https://hudi.apache.org/ ; https://api.github.com/repos/apache/hudi/releases ; https://api.github.com/repos/apache/paimon/releases ; https://paimon.apache.org/docs/master/ ; https://github.com/apache/hudi/releases/tag/release-1.2.1 ; https://github.com/apache/paimon/releases
Open items
- Paimon origin (Flink Table Store), incubation/graduation dates and Hudi graduation date not sourced (docs page gives no history; Wikipedia pages were unavailable); vendor adoption not sourced. agy search timed out 2026-10-07.