Tables
The REST Catalog can expose three table categories. Choose one based on how the underlying data is organized and the capabilities implemented by your catalog service.
| Category | Data model | Guide |
|---|---|---|
| Paimon Table | Snapshot-managed data, with or without a primary key. | Primary-key table, Append table |
| Format Table | Files in a common format, including Hive-compatible tables. | Format Table |
| Object Table | Metadata for unstructured objects in a storage directory. | Object Table |
Paimon Table
Primary Key Table
A primary-key table merges records with the same primary key according to its configured merge engine. It supports streaming updates and changelog reads. See Primary Key Table for bucket modes, merge engines, and changelog configuration.
The following examples create a table with eight buckets:
- Flink SQL
- Spark SQL
CREATE TABLE my_table (
a INT PRIMARY KEY NOT ENFORCED,
b STRING
) WITH (
'bucket' = '8'
);
CREATE TABLE my_table (
a INT,
b STRING
) TBLPROPERTIES (
'primary-key' = 'a',
'bucket' = '8'
);
Append Table
A table without a primary key is an append table. Streaming writes append records rather than
applying primary-key upserts. Batch DELETE, UPDATE, and MERGE INTO operations are available
where supported by the compute engine. See Append Table for write and
read behavior.
CREATE TABLE my_table (
a INT,
b STRING
);
Format Table
A Format Table describes files in a common format, including compatible Hive tables exposed through the catalog. It lets engines read existing files and add new files without using Paimon's snapshot-managed table layout.
Common formats include CSV, Parquet, ORC, JSON, and text. Mosaic is also available with the corresponding format module. Format availability depends on the modules and table implementation used by the engine.
By default, partition discovery follows the directory layout. REST internal Format Tables can
instead use metastore.partitioned-table = true to read registered partitions from the catalog;
unregistered directories are then excluded from scans. This mode uses the Paimon implementation
and cannot be combined with format-table.implementation = engine. See
partition API compatibility for partition options and
custom locations.
The following examples assume that a REST Catalog and database are already selected:
- Flink-CSV
- Spark-CSV
- Flink-Parquet
- Spark-Parquet
- Flink-JSON
- Spark-JSON
CREATE TABLE my_csv_table (
a INT,
b STRING
) WITH (
'type' = 'format-table',
'file.format' = 'csv',
'csv.field-delimiter' = ','
);
CREATE TABLE my_csv_table (
a INT,
b STRING
) USING csv OPTIONS ('csv.field-delimiter' ',');
CREATE TABLE my_parquet_table (
a INT,
b STRING
) WITH (
'type' = 'format-table',
'file.format' = 'parquet'
);
CREATE TABLE my_parquet_table (
a INT,
b STRING
) USING parquet;
CREATE TABLE my_json_table (
a INT,
b STRING
) WITH (
'type' = 'format-table',
'file.format' = 'json'
);
CREATE TABLE my_json_table (
a INT,
b STRING
) USING json;
Object Table
An Object Table exposes metadata for unstructured files in a storage directory. Use SQL to query the file list and PVFS to access file contents through a REST Catalog.
REST, Hive, and JDBC catalogs support Object Tables. The available permissions and storage credentials depend on the catalog service. Create an Object Table without declaring user columns:
- Flink-SQL
- Spark-SQL
CREATE TABLE `my_object_table` WITH (
'type' = 'object-table'
);
CREATE TABLE `my_object_table` TBLPROPERTIES (
'type' = 'object-table'
);