Skip to main content

Hive Table Migration

Migrate existing Hive data files into Paimon append tables using a Paimon Hive catalog. The migrator moves files and registers them in Paimon metadata. It supports ORC, Parquet, and Avro source files, for both partitioned and non-partitioned tables.

By default, the original Hive table is replaced by a Paimon table with the same name. Use Clone to Paimon when you need a separate copy and a readable source table.

Migration is not atomic

Back up the source data and Hive metadata before migrating, and stop writes to the source tables. An interruption can leave files or metadata partially migrated. Do not rely on delete_origin => false as a backup: it keeps the Hive metadata, but files are still moved out of the source location.

Before You Start​

  • Configure a Paimon catalog with metastore = hive and access to the source Hive metastore. See Catalog, Flink Procedures, or Spark Catalogs for setup.
  • Give the migration runtime access to source and target storage, including the ability to move files and update Hive metadata. This is a file-move workflow, not a cross-storage copy.
  • Use a target append table with no primary key and bucket = -1. For an existing target, check that the source data schema is compatible and partition names and types match.
  • Record source row counts, partition values, and representative query results for validation. The examples below use an ORC source table named default.hivetable; adjust identifiers, paths, and file.format for your source.

Choose the Target Behavior​

Procedure argumentsTargetSource after success
Omit target_tableReplaces the original table under its original nameOriginal Hive table is removed
Set a different target_table; omit delete_origin or set it to trueCreates or imports files into the named Paimon tableOriginal Hive table is removed
Set a different target_table and delete_origin => falseCreates or imports files into the named Paimon tableHive metadata remains, but its files have moved

For an existing target, migration adds imported files; it does not use clone's partition overwrite behavior. Check the existing target data before importing to avoid duplicates. target_table and delete_origin are SQL procedure options; the Flink migration action shown below uses the default replacement behavior.

Migrate Hive Table​

Choose one of the following entry points. The SQL examples select the Paimon Hive catalog before calling the procedure.

CREATE CATALOG paimon_hive WITH (
'type' = 'paimon',
'metastore' = 'hive',
'uri' = 'thrift://localhost:9083',
'warehouse' = 'hdfs:///warehouse'
);

USE CATALOG paimon_hive;

-- Flink 1.19 and later: replace default.hivetable with a Paimon table.
CALL sys.migrate_table(
connector => 'hive',
source_table => 'default.hivetable',
options => 'file.format=orc'
);

For Flink 1.18, use positional arguments instead of the call above:

CALL sys.migrate_table('hive', 'default.hivetable', 'file.format=orc');

To import into a different target with Flink 1.19 or later, use this call instead of the replacement call. The original Hive table is removed by default.

CALL sys.migrate_table(
connector => 'hive',
source_table => 'default.hivetable',
target_table => 'default.paimon_target',
options => 'file.format=orc'
);

After success, read and write the target through Paimon. Existing jobs that expect the original Hive storage format must be updated. Hive queries need the Paimon Hive connector.

Migrate Hive Database​

Database migration attempts each table in the source database and replaces successfully migrated tables under their original names. It is not a transaction across all tables: a failure can leave a mixture of Hive and Paimon tables. Review every source table's format and requirements first. Use individual table migration for a selected subset.

Using the Paimon Hive catalog configured above:

USE CATALOG paimon_hive;

-- Flink 1.19 and later
CALL sys.migrate_database(
connector => 'hive',
source_database => 'default',
options => 'file.format=orc'
);

For Flink 1.18, use this call instead:

CALL sys.migrate_database('hive', 'default', 'file.format=orc');

Verify the Migration​

  1. Check the procedure output and logs. For a database, inspect both success and failure counts; completion of the call does not mean every table migrated.
  2. Inspect the target schema, partition keys, and table options. Confirm that it is a Paimon append table with bucket = -1.
  3. Query the target through Paimon and compare row counts and aggregates against the values recorded before migration. Include null partition values and representative partitions.
  4. Update reader and writer configurations, then verify them against the migrated table.

For example, in the selected Paimon catalog:

SHOW CREATE TABLE default.hivetable;
SELECT COUNT(*) FROM default.hivetable;

If migration fails, inspect the source and target metadata and file locations before retrying. For database migration, identify the failed tables and handle them individually after resolving the cause; do not assume another database-wide call will safely resume the previous run.