PostgreSQL CDC
Synchronize one or more PostgreSQL tables into one Paimon table. Tables may come from multiple schemas in the same source database.
See CDC Action Configuration for command syntax, computed columns, and job settings, and Schema Evolution and Type Mapping for shared behavior.
Prerequisites
Place these dependencies in <FLINK_HOME>/lib/ alongside the Paimon Flink bundled jar.
The dependency baseline for this Paimon version is Flink CDC 3.5.0.
All source tables selected by the table action must have primary keys, including when you supply
--primary_keys for the target. Schemas must be compatible when merging multiple source tables.
The partitioned examples below assume each source table has id and a non-null create_time.
Set REPLICA IDENTITY FULL on each selected source table so update and delete events include
the old create_time needed to locate the previous partition. For example:
ALTER TABLE public.source_table1 REPLICA IDENTITY FULL;
ALTER TABLE public.source_table2 REPLICA IDENTITY FULL;
Use a replication slot appropriate for the source job; the examples set slot.name=paimon_cdc.
This page covers postgres_sync_table; Paimon does not provide a PostgreSQL database action.
Synchronizing Tables
By using PostgresSyncTableAction in a Flink DataStream job or directly through flink run, users can synchronize one or multiple tables from PostgreSQL into one Paimon table.
If the Paimon table you specify does not exist, this action will automatically create the table. Its schema will be derived from all specified PostgreSQL tables. If the Paimon table already exists, its schema will be compared against the schema of all specified PostgreSQL tables.
Example 1: synchronize tables into one Paimon table
<FLINK_HOME>/bin/flink run \
/path/to/paimon-flink-action-2.2-SNAPSHOT.jar \
postgres_sync_table \
--warehouse hdfs:///path/to/warehouse \
--database test_db \
--table test_table \
--partition_keys pt \
--primary_keys id,pt \
--computed_column 'pt=date_format(create_time,yyyy-MM-dd)' \
--postgres_conf hostname=127.0.0.1 \
--postgres_conf username=root \
--postgres_conf password=123456 \
--postgres_conf database-name='source_db' \
--postgres_conf schema-name='public' \
--postgres_conf table-name='source_table1|source_table2' \
--postgres_conf slot.name='paimon_cdc' \
--catalog_conf metastore=hive \
--catalog_conf uri=thrift://hive-metastore:9083 \
--table_conf bucket=4 \
--table_conf changelog-producer=input \
--table_conf sink.parallelism=4
As example shows, the postgres_conf's table-name supports regular expressions to monitor multiple tables that satisfy the regular expressions. The schemas of all the tables will be merged into one Paimon table schema.
Example 2: synchronize shards into one Paimon table
You can also set 'schema-name' with a regular expression to capture multiple schemas. A typical scenario is that a table 'source_table' is split into schema 'source_schema1', 'source_schema2' ..., then you can synchronize data of all the 'source_table's into one Paimon table.
<FLINK_HOME>/bin/flink run \
/path/to/paimon-flink-action-2.2-SNAPSHOT.jar \
postgres_sync_table \
--warehouse hdfs:///path/to/warehouse \
--database test_db \
--table test_table \
--partition_keys pt \
--primary_keys id,pt \
--computed_column 'pt=date_format(create_time,yyyy-MM-dd)' \
--postgres_conf hostname=127.0.0.1 \
--postgres_conf username=root \
--postgres_conf password=123456 \
--postgres_conf database-name='source_db' \
--postgres_conf schema-name='source_schema.+' \
--postgres_conf table-name='source_table' \
--postgres_conf slot.name='paimon_cdc' \
--catalog_conf metastore=hive \
--catalog_conf uri=thrift://hive-metastore:9083 \
--table_conf bucket=4 \
--table_conf changelog-producer=input \
--table_conf sink.parallelism=4
Table Action Reference
Command syntax (square brackets indicate optional arguments):
<FLINK_HOME>/bin/flink run \
/path/to/paimon-flink-action-2.2-SNAPSHOT.jar \
postgres_sync_table \
--warehouse <warehouse_path> \
--database <database_name> \
--table <table_name> \
[--partition_keys <partition_keys>] \
[--primary_keys <primary_keys>] \
[--type_mapping <option1,option2...>] \
[--computed_column <'column-name=expr-name(args[, ...])'> [--computed_column ...]] \
[--metadata_column <metadata_column>] \
[--postgres_conf <postgres_cdc_source_conf> [--postgres_conf <postgres_cdc_source_conf> ...]] \
[--catalog_conf <paimon_catalog_conf> [--catalog_conf <paimon_catalog_conf> ...]] \
[--table_conf <paimon_table_sink_conf> [--table_conf <paimon_table_sink_conf> ...]]
| Configuration | Description |
|---|---|
--warehouse |
The path to Paimon warehouse. |
--database |
The database name in Paimon catalog. |
--table |
The Paimon table name. |
--partition_keys |
The partition keys for Paimon table. If there are multiple partition keys, connect them with comma, for example "dt,hh,mm". |
--primary_keys |
The primary keys for Paimon table. If there are multiple primary keys, connect them with comma, for example "buyer_id,seller_id". |
--type_mapping |
It is used to specify how to map PostgreSQL data type to Paimon type. Supported options:
|
--sync_primary_keys_from_source_schema |
This is used to specify if primary keys from source should be used in paimon schema if primary keys using --primary_keys are not specified. The default is true. |
--computed_column |
The definitions of computed columns. The argument field is from PostgreSQL table field name. See here for a complete list of configurations. |
--metadata_column |
--metadata_column is used to specify which metadata columns to include in the output schema of the connector. Metadata columns provide additional information related to the source data, for example: --metadata_column table_name,database_name,schema_name,op_ts. See its document for a complete list of available metadata. |
--postgres_conf |
The configuration for Flink CDC Postgres sources. Each configuration should be specified in the format "key=value". hostname, username, password, database-name, schema-name, table-name and slot.name are required configurations, others are optional. See its document for a complete list of configurations. |
--catalog_conf |
The configuration for Paimon catalog. Each configuration should be specified in the format "key=value". See here for a complete list of catalog configurations. |
--table_conf |
The configuration for Paimon table sink. Each configuration should be specified in the format "key=value". See here for a complete list of table configurations. |