First Row
Set merge-engine = first-row to keep the first row for each primary key and ignore subsequent
rows for that key. This is useful for event or log deduplication. Unlike deduplicate, later
values do not replace the retained row.
Create a Table
CREATE TABLE events (
event_id BIGINT,
payload STRING,
PRIMARY KEY (event_id) NOT ENFORCED
) WITH (
'merge-engine' = 'first-row',
'changelog-producer' = 'lookup'
);
Inputs (1, 'first') and (1, 'later') leave (1, 'first') in the table.
Streaming Reads
The engine supports none and lookup changelog producers. Use
lookup for streaming reads that must emit a key only once:
lookup compaction checks existing keys and produces an insert-only changelog of newly retained
rows. none does not provide the same cross-commit streaming deduplication.
Managed BLOB storage requires none, so it cannot
use this lookup-changelog streaming pattern.
Ordering and Deletes
- User-defined sequence fields are not supported.
DELETEandUPDATE_BEFORErecords are rejected by default. Setignore-delete = trueto discard them when the source may emit retractions.- The retained row follows ingestion ordering, not the smallest value of a business timestamp.
Compaction and Visibility
By default, batch reads expose Level-0 data only after lookup compaction. Writers wait for lookup compaction by default; asynchronous compaction can delay visibility. See Lookup Compaction.
Do not enable deletion vectors for an ordinary first-row table. The
PK Clustering Override layout is a special case with its own requirements.