Migration
Choose a migration path based on whether you need to replace a Hive table, keep a separate copy, or expose historical Paimon data to existing Hive queries.
Choose a Guide
| Goal | Guide | Effect on the source | Result |
|---|---|---|---|
| Convert a Hive table or database to Paimon | Migrate from Hive | Moves data files; removes the original Hive table by default | Paimon append tables in a Hive catalog |
| Copy tables while keeping the source available | Clone to Paimon | Keeps source tables and data | Separate Paimon tables; Hive sources become append tables |
| Query daily views of an updating Paimon table using Hive partition filters | Expose Tags as Hive Partitions | Keeps the Paimon table and its write path | Hive partition values that select tags or preview snapshots |
Migration and clone import existing data. Tag-to-partition changes how Hive reads a Paimon table; it does not migrate Hive files or physically repartition the Paimon table.
Before Moving Data
- Choose the target table model. Hive migration and Hive clone create append tables. If the target needs primary-key updates or a different schema or partition layout, plan a data rewrite into a table with that model.
- Prepare the runtime and catalogs. Configure access to the Hive metastore and storage. Use the engine setup and connector requirements linked from each guide.
- Define the validation scope. Record source schemas, partitions, row counts, and key aggregates before the operation. For a consistent comparison, keep source data stable while it is being moved or copied.
- Plan the switch. In-place migration is not atomic and requires a backup. Clone lets you validate a separate target before switching readers and writers.
Related Guides
- Flink Procedures and Spark Migration Procedures: engine-specific arguments.
- COPY INTO: import files with Spark SQL.
- Manage Tags: create and retain historical views.
- Hive: install the connector used to query Paimon from Hive.