Skip to main content

Iceberg Compatibility

Paimon can publish Iceberg metadata that points to its existing data files. This lets applications query a Paimon table through an Iceberg connector while Paimon continues to manage writes, compaction, and data retention.

Paimon commits publish Iceberg metadata; both readers access the same data files.

Start Here​

TaskGuide
Read your first Paimon table through Flink or Spark's Iceberg connectorAppend tables
Read updates and deletes from a primary key tablePrimary key tables
Choose Hadoop, Hive, REST, or path-based accessCatalogs and metadata layout
Query a named historical snapshotTags
Connect Trino, Athena, or DuckDBQuery engines
Check column types and format requirementsData types
Look up table optionsConfiguration reference

How Publication Works​

  1. A writer commits a Paimon snapshot.
  2. Paimon generates Iceberg manifests and snapshot metadata for the files eligible for Iceberg reads.
  3. With Hive or REST storage, Paimon also publishes the metadata to the external catalog.
  4. An Iceberg reader loads the published metadata and reads the referenced data files directly.

Enable publication with the Paimon table option metadata.iceberg.storage; its default is disabled. For a first example, use hadoop-catalog. The Iceberg warehouse is then <paimon-warehouse>/iceberg, using the default metadata layout.

'metadata.iceberg.storage' = 'hadoop-catalog'

Metadata publication does not copy the table's data. Iceberg readers therefore need access to both the metadata location and the original Paimon data files, including the required filesystem configuration and credentials.

What Iceberg Readers See​

Paimon tableFiles eligible for incremental publicationWhen changes become visible
Append tableData files in the committed snapshotAfter metadata publication and reader refresh
Primary key table without Iceberg deletion vectorsFiles at the highest LSM levelAfter full compaction, metadata publication, and reader refresh
Primary key table with Iceberg v3 deletion vectorsFiles above L0, together with deletion vectorsAfter changes reach those files and metadata is published and refreshed

Initial publication and metadata rebuilds use snapshot splits that can be read directly without Paimon's merge logic. The table above describes subsequent incremental publication. Use compaction to establish a predictable visibility boundary for primary key tables.

The primary key guide explains both modes and their configuration. Disabling an Iceberg catalog's cache can help with interactive verification, but cannot make uncompacted or unpublished changes visible.

Manage the table through Paimon

Use the Iceberg representation for reads. Perform writes, schema changes, compaction, snapshot expiration, and file cleanup through Paimon. Both representations refer to shared data files; independent Iceberg mutations or cleanup can invalidate Paimon's view of the table.

Supported Types​

Compatibility depends on the column types, data file format, Iceberg format version, and reader. See supported data types and precision limits before enabling publication on an existing table. Primary key deletion vectors and geospatial columns require Iceberg format v3.