Overview
Apache Paimon is a table format for data lakes that supports batch and streaming workloads. Compute engines read and write Paimon tables, while table data and file metadata live in a filesystem or object store. A catalog provides names and metadata operations for those tables.
Start Here
Read the core concepts in this order:
- Basic Concepts: choose a table type, understand partitions and buckets, and follow a read from snapshot to data files.
- Concurrency Control: understand how snapshots become visible and what happens when writers compete.
- Catalog: choose how engines discover and manage tables.
To run your first table, use the Flink quick start, Spark quick start, or Python API.
Unified Storage
A Paimon table can serve several access patterns. The available operations depend on the table type, engine, and connector.
| Access pattern | What the reader or writer does | Learn more |
|---|---|---|
| Batch reads | Query the latest table state or a retained historical snapshot. | Flink queries, Spark queries |
| Streaming reads | Discover new snapshots and consume appended records or changes, according to the table's changelog configuration. | Append streaming, Changelog producers |
| Streaming writes | Ingest new events or apply database changes to a primary-key table. | CDC ingestion |
| Batch writes | Insert records, overwrite data, or use the row-level operations supported by the engine and table. | Flink writes, Spark writes |
Historical snapshots and changelogs are subject to retention. Configure snapshot expiration and, when needed, tags to match your recovery and time-travel requirements.
Explore by Topic
| Topic | Pages |
|---|---|
| Table design | Append tables, Primary-key tables, Multimodal tables |
| Catalog and metadata | Catalog, Data types, Views, Functions |
| Inspect a table | System tables for snapshots, files, partitions, indexes, and configuration |
| REST Catalog | Architecture and setup, authentication, table types, and API references |
| Storage specification | On-disk layout, schemas, snapshots, manifests, data files, and indexes |
The storage specification is a reference for readers implementing or inspecting the format. Start with Basic Concepts for the relationships between its parts.