Overview
An append table has no primary key. Each inserted row is stored as a new record, including rows whose values duplicate an existing record. Inserts do not perform key-based deduplication or upserts. Use a primary key table when incoming changelog records should update existing rows by key.
Append tables support batch and streaming workloads, with snapshots, time travel, schema evolution, and file-level query optimizations. You can also change stored rows with explicit row-level operations.
Create an Append Table
The examples below assume that you have configured and selected a Paimon catalog. See the Flink quick start or Spark quick start for catalog setup.
- Flink
- Spark
CREATE TABLE my_table (
product_id BIGINT,
price DOUBLE,
sales BIGINT,
dt STRING
) PARTITIONED BY (dt) WITH (
'bucket' = '-1'
);
INSERT INTO my_table VALUES (1, 10.0, 2, '2026-09-10');
SELECT * FROM my_table WHERE dt = '2026-09-10';
CREATE TABLE my_table (
product_id BIGINT,
price DOUBLE,
sales BIGINT,
dt STRING
) USING paimon
PARTITIONED BY (dt)
TBLPROPERTIES (
'bucket' = '-1'
);
INSERT INTO my_table VALUES (1, 10.0, 2, '2026-09-10');
SELECT * FROM my_table WHERE dt = '2026-09-10';
bucket = -1 is the default for append tables; it is shown explicitly here to identify the layout. Partitioning is
optional. File size, format, and compression can be configured with target-file-size, file.format, and
file.compression; see Configurations.
Choose a Layout
| Layout | Configuration | Use it when | Considerations |
|---|---|---|---|
| Unaware-bucket append | bucket = -1 (default) | You want flexible ingestion without distributing rows by a bucket key. | No streaming row-order guarantee. Supports row tracking. |
| Bucketed append | Positive bucket and a bucket-key | Queries filter on bucket keys, joins can reuse the distribution, or streaming consumers need ordering within a bucket. | Ordering is scoped to one partition and bucket, and requires bucket-append-ordered = true. |
Incremental clustering is a data-layout optimization available for both layouts. On bucketed append tables it requires giving up the ordered append guarantee. Bucketing, clustering, and row tracking solve different problems; choose them according to your query and ingestion requirements.
Read and Maintain the Table
Examples named my_table use the schema above. Individual guides introduce separate tables when they need a different
layout or schema.
Streaming
Streaming covers Flink ingestion, small-file compaction, scan startup modes, and watermarks. For ordering requirements, see Bucketed streaming.
Query Performance
Aggregate Pushdown
Some aggregate queries can use metadata instead of reading every data row. See Aggregate pushdown for examples and limits.
Clustering and File Statistics
Sorting can narrow the value ranges in each file and improve pruning. Start with File statistics and clustering, then configure Incremental clustering for recurring optimization.
File Indexes
File indexes provide additional filtering for supported predicates, including Bloom filter, bitmap, and range bitmap indexes.
Row-Level Operations
Row-level operations explains Spark SQL DELETE, UPDATE, and MERGE INTO, including the
choice between rewriting data files and using deletion vectors. Row tracking adds hidden row IDs
and row versions for tracking changes across these operations.