Skip to main content

Hive Catalog

Use hive-catalog to publish an Iceberg table in a Hive metastore. Paimon writes Iceberg metadata and updates the metastore entry to point to it. Iceberg readers then discover the table through their own Hive catalog connector.

Before You Begin​

Prepare the Paimon and Iceberg engine connectors as described in the append-table walkthrough. The Paimon writer also needs the Paimon Hive catalog module and its Hive dependencies; see Hive catalog setup. Both the writer and reader need access to the metastore and the files referenced by the table.

Publish a Table​

The following Flink SQL uses a filesystem Paimon catalog and registers the Iceberg representation in Hive. Replace the warehouse and metastore placeholders.

SET 'execution.runtime-mode' = 'batch';
SET 'table.dml-sync' = 'true';

CREATE CATALOG paimon_catalog WITH (
'type' = 'paimon',
'warehouse' = '<path-to-warehouse>'
);

CREATE DATABASE IF NOT EXISTS paimon_catalog.`default`;

CREATE TABLE paimon_catalog.`default`.animals (
kind STRING,
name STRING
) WITH (
'metadata.iceberg.storage' = 'hive-catalog',
'metadata.iceberg.uri' = 'thrift://<metastore-host>:9083'
);

INSERT INTO paimon_catalog.`default`.animals VALUES
('mammal', 'cat'), ('mammal', 'dog'),
('reptile', 'snake'), ('reptile', 'lizard');

For Spark, use the same table options in TBLPROPERTIES; the append-table walkthrough shows the corresponding syntax.

Read through Iceberg​

In Flink, connect an Iceberg Hive catalog to the same metastore:

CREATE CATALOG iceberg_hive WITH (
'type' = 'iceberg',
'catalog-type' = 'hive',
'uri' = 'thrift://<metastore-host>:9083',
'cache-enabled' = 'false'
);

SELECT kind, name FROM iceberg_hive.`default`.animals
WHERE kind = 'mammal' ORDER BY name;
kind name
mammal cat
mammal dog

See Trino for a reader that uses the same Hive registration. For primary key tables, the compaction or deletion-vector requirements also apply.

Table Names and Metadata Location​

By default, the Iceberg database and table names match the Paimon names. When the Paimon catalog also uses the same Hive metastore, give the Iceberg representation a distinct name to avoid reusing the Paimon table's metastore entry:

'metadata.iceberg.database' = 'iceberg_analytics',
'metadata.iceberg.table' = 'animals'

Readers would then query iceberg_hive.iceberg_analytics.animals. For Hive publication, metadata.iceberg.database also accepts multiple databases separated by semicolons. Aliases change catalog registration names, not the physical metadata path.

Metadata is stored under <warehouse>/iceberg/<database>/<table>/metadata by default. Use metadata.iceberg.storage-location = table-location for metadata alongside the Paimon table, including databases with custom locations. See metadata layout.

Connection Options​

Paimon table optionWhen to use it
metadata.iceberg.uriSet the Hive metastore Thrift URI explicitly
metadata.iceberg.hive-conf-dirLoad Hive configuration such as hive-site.xml
metadata.iceberg.hadoop-conf-dirLoad Hadoop configuration needed by the Hive client
metadata.iceberg.hive-client-classUse a custom Hive metastore client
metadata.iceberg.hive-skip-update-statsSkip updating Hive statistics

If the URI is not set as a table option, provide hive.metastore.uris in the loaded configuration. See configuration reference for defaults and other publication options.

AWS Glue Catalog​

To publish through a Hive-compatible AWS Glue client, set:

'metadata.iceberg.hive-client-class' = 'com.amazonaws.glue.catalog.metastore.AWSCatalogMetastoreClient'

Install a Glue Hive client compatible with your Hive dependencies on the writer classpath, and configure its AWS region and credentials. The AWS Glue Data Catalog client provides build and configuration instructions. metadata.iceberg.glue.skip-archive controls whether publication skips archiving Glue table versions.

Reader support has separate constraints; see Athena.