Skip to main content

Overview

Apache Paimon is a table format for data lakes that supports batch and streaming workloads. Compute engines read and write Paimon tables, while table data and file metadata live in a filesystem or object store. A catalog provides names and metadata operations for those tables.

Compute engines use a catalog to discover Paimon tables and read or write their snapshots and data in shared storage.

Start Here​

Read the core concepts in this order:

  1. Basic Concepts: choose a table type, understand partitions and buckets, and follow a read from snapshot to data files.
  2. Concurrency Control: understand how snapshots become visible and what happens when writers compete.
  3. Catalog: choose how engines discover and manage tables.

To run your first table, use the Flink quick start, Spark quick start, or Python API.

Unified Storage​

A Paimon table can serve several access patterns. The available operations depend on the table type, engine, and connector.

Access patternWhat the reader or writer doesLearn more
Batch readsQuery the latest table state or a retained historical snapshot.Flink queries, Spark queries
Streaming readsDiscover new snapshots and consume appended records or changes, according to the table's changelog configuration.Append streaming, Changelog producers
Streaming writesIngest new events or apply database changes to a primary-key table.CDC ingestion
Batch writesInsert records, overwrite data, or use the row-level operations supported by the engine and table.Flink writes, Spark writes

Historical snapshots and changelogs are subject to retention. Configure snapshot expiration and, when needed, tags to match your recovery and time-travel requirements.

Explore by Topic​

TopicPages
Table designAppend tables, Primary-key tables, Multimodal tables
Catalog and metadataCatalog, Data types, Views, Functions
Inspect a tableSystem tables for snapshots, files, partitions, indexes, and configuration
REST CatalogArchitecture and setup, authentication, table types, and API references
Storage specificationOn-disk layout, schemas, snapshots, manifests, data files, and indexes

The storage specification is a reference for readers implementing or inspecting the format. Start with Basic Concepts for the relationships between its parts.