Skip to main content

Python API

Use the Python API to manage catalogs and control table reads, writes, and commits. Start with the quick start for a runnable example. For a compact interface to text, vector, and BLOB data, see the Multimodal API.

TaskGuide
Connect and define schemasCatalogs and tables
Write and commit dataBatch writes
Filter, project, and plan readsBatch reads
Follow new snapshotsStreaming and consumers
Choose Arrow types or use VARIANTData types
Maintain dataset versionsBranches and rollback, tags

Catalogs and tables​

Connect to filesystem, JDBC, or REST catalogs. Create databases and tables, define partition and primary keys, and alter schemas.

Read the guide.

Batch writes​

Write Arrow or pandas batches, publish a snapshot, overwrite data, and register commit callbacks.

Read the guide.

Batch reads​

Push down predicates and projections, plan splits, and read Arrow, pandas, Python iterators, or DuckDB results. Includes incremental scans, sharding, and plan inspection.

Read the guide.

Streaming and consumers​

Poll for new snapshots, filter streams, assign buckets to workers, and manage saved consumer positions.

Read the guide.

Data types​

Map Arrow types to Paimon types, store semi-structured VARIANT data, and update VARIANT paths.

Read the guide.

Branches and rollback​

Create an empty branch or initialize one from a tag, manage branch histories, and roll back to a retained snapshot. Use tags to pin a dataset version.

Read the guide.

Supported Features​

PyPaimon supports filesystem, JDBC, and REST catalogs; local storage, HDFS, S3, OSS, and GCS; and batch reads and writes for append and primary-key tables. The guides cover projections, predicates, overwrite, incremental and streaming reads, BLOB data, row sharding, and table version management.

Primary-key reads support deduplicate, first-row, basic partial-update, and aggregation with supported built-in aggregators. Support depends on the engine configuration: for example, partial-update sequence groups and field aggregators are not supported, and some delete/retract options are rejected. Do not assume that every Java merge-engine option is available in Python.

For media and column updates, use data-evolution tables. For optional file formats and compute engines, see Installation.