Skip to main content

Overview

An append table has no primary key. Each inserted row is stored as a new record, including rows whose values duplicate an existing record. Inserts do not perform key-based deduplication or upserts. Use a primary key table when incoming changelog records should update existing rows by key.

Append tables support batch and streaming workloads, with snapshots, time travel, schema evolution, and file-level query optimizations. You can also change stored rows with explicit row-level operations.

Create an Append Table​

The examples below assume that you have configured and selected a Paimon catalog. See the Flink quick start or Spark quick start for catalog setup.

CREATE TABLE my_table (
product_id BIGINT,
price DOUBLE,
sales BIGINT,
dt STRING
) PARTITIONED BY (dt) WITH (
'bucket' = '-1'
);

INSERT INTO my_table VALUES (1, 10.0, 2, '2026-09-10');
SELECT * FROM my_table WHERE dt = '2026-09-10';

bucket = -1 is the default for append tables; it is shown explicitly here to identify the layout. Partitioning is optional. File size, format, and compression can be configured with target-file-size, file.format, and file.compression; see Configurations.

Choose a Layout​

LayoutConfigurationUse it whenConsiderations
Unaware-bucket appendbucket = -1 (default)You want flexible ingestion without distributing rows by a bucket key.No streaming row-order guarantee. Supports row tracking.
Bucketed appendPositive bucket and a bucket-keyQueries filter on bucket keys, joins can reuse the distribution, or streaming consumers need ordering within a bucket.Ordering is scoped to one partition and bucket, and requires bucket-append-ordered = true.

Incremental clustering is a data-layout optimization available for both layouts. On bucketed append tables it requires giving up the ordered append guarantee. Bucketing, clustering, and row tracking solve different problems; choose them according to your query and ingestion requirements.

Read and Maintain the Table​

Examples named my_table use the schema above. Individual guides introduce separate tables when they need a different layout or schema.

Streaming​

Streaming covers Flink ingestion, small-file compaction, scan startup modes, and watermarks. For ordering requirements, see Bucketed streaming.

Query Performance​

Aggregate Pushdown​

Some aggregate queries can use metadata instead of reading every data row. See Aggregate pushdown for examples and limits.

Clustering and File Statistics​

Sorting can narrow the value ranges in each file and improve pruning. Start with File statistics and clustering, then configure Incremental clustering for recurring optimization.

File Indexes​

File indexes provide additional filtering for supported predicates, including Bloom filter, bitmap, and range bitmap indexes.

Row-Level Operations​

Row-level operations explains Spark SQL DELETE, UPDATE, and MERGE INTO, including the choice between rewriting data files and using deletion vectors. Row tracking adds hidden row IDs and row versions for tracking changes across these operations.