Append Tables
This walkthrough creates an append table through Paimon and reads the same data through an Iceberg Hadoop catalog. It also provides the catalog setup used by the primary key examples.
Before You Begin
Install the Paimon connector for your engine and an Iceberg runtime compatible with that engine. See the Paimon Flink quick start, Paimon Spark quick start, and Iceberg's Flink setup or Spark setup.
Replace the path and version placeholders below. Both connectors must be able to access the same warehouse. In a cluster, use shared storage accessible to every worker.
Prepare Catalogs
- Flink SQL
- Spark SQL
Install the Paimon and Iceberg Flink runtime JARs before starting the cluster and SQL client. Use batch mode for the bounded examples and synchronous DML so each insert finishes before its query.
SET 'execution.runtime-mode' = 'batch';
SET 'table.dml-sync' = 'true';
CREATE CATALOG paimon_catalog WITH (
'type' = 'paimon',
'warehouse' = '<path-to-warehouse>'
);
CREATE DATABASE IF NOT EXISTS paimon_catalog.`default`;
CREATE CATALOG iceberg_catalog WITH (
'type' = 'iceberg',
'catalog-type' = 'hadoop',
'warehouse' = '<path-to-warehouse>/iceberg',
'cache-enabled' = 'false'
);
Start spark-sql with both catalogs. The Iceberg artifact name contains the Spark major/minor and
Scala binary versions; its Maven version follows the colon. Select a combination published by
Iceberg and supported by your Paimon Spark connector.
spark-sql \
--jars <path-to-paimon-spark-jar> \
--packages org.apache.iceberg:iceberg-spark-runtime-<spark-major.minor>_<scala-binary-version>:<iceberg-version> \
--conf spark.sql.catalog.paimon_catalog=org.apache.paimon.spark.SparkCatalog \
--conf spark.sql.catalog.paimon_catalog.warehouse=<path-to-warehouse> \
--conf spark.sql.catalog.iceberg_catalog=org.apache.iceberg.spark.SparkCatalog \
--conf spark.sql.catalog.iceberg_catalog.type=hadoop \
--conf spark.sql.catalog.iceberg_catalog.warehouse=<path-to-warehouse>/iceberg \
--conf spark.sql.catalog.iceberg_catalog.cache-enabled=false \
--conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions,org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions
CREATE DATABASE IF NOT EXISTS paimon_catalog.`default`;
The examples disable Iceberg catalog caching to make newly published metadata easier to observe. For Hive or REST access, use the corresponding catalog guide.
Create the Paimon Table
- Flink SQL
- Spark SQL
CREATE TABLE paimon_catalog.`default`.cities (
country STRING,
name STRING
) WITH (
'metadata.iceberg.storage' = 'hadoop-catalog'
);
CREATE TABLE paimon_catalog.`default`.cities (
country STRING,
name STRING
) TBLPROPERTIES (
'metadata.iceberg.storage' = 'hadoop-catalog'
);
Write with Paimon, Read with Iceberg
Run the following SQL in either engine. The catalog name determines which connector handles each statement.
INSERT INTO paimon_catalog.`default`.cities VALUES
('usa', 'new york'),
('germany', 'berlin'),
('usa', 'chicago'),
('germany', 'hamburg');
SELECT country, name
FROM iceberg_catalog.`default`.cities
WHERE country = 'germany'
ORDER BY name;
Expected rows:
country name
germany berlin
germany hamburg
Insert another row through Paimon, then read the refreshed Iceberg representation:
INSERT INTO paimon_catalog.`default`.cities VALUES ('germany', 'munich');
SELECT country, name
FROM iceberg_catalog.`default`.cities
WHERE country = 'germany'
ORDER BY name;
country name
germany berlin
germany hamburg
germany munich
If the table or new rows are missing, first check that the insert completed, the Iceberg warehouse
ends in /iceberg, and the reader can access the published metadata and original data files.
For updates and deletes, continue with primary key tables.