Python API
Use the Python API to manage catalogs and control table reads, writes, and commits. Start with the quick start for a runnable example. For a compact interface to text, vector, and BLOB data, see the Multimodal API.
| Task | Guide |
|---|---|
| Connect and define schemas | Catalogs and tables |
| Write and commit data | Batch writes |
| Filter, project, and plan reads | Batch reads |
| Follow new snapshots | Streaming and consumers |
| Choose Arrow types or use VARIANT | Data types |
| Maintain dataset versions | Branches and rollback, tags |
Catalogs and tables
Connect to filesystem, JDBC, or REST catalogs. Create databases and tables, define partition and primary keys, and alter schemas.
Batch writes
Write Arrow or pandas batches, publish a snapshot, overwrite data, and register commit callbacks.
Batch reads
Push down predicates and projections, plan splits, and read Arrow, pandas, Python iterators, or DuckDB results. Includes incremental scans, sharding, and plan inspection.
Streaming and consumers
Poll for new snapshots, filter streams, assign buckets to workers, and manage saved consumer positions.
Data types
Map Arrow types to Paimon types, store semi-structured VARIANT data, and update VARIANT paths.
Branches and rollback
Create an empty branch or initialize one from a tag, manage branch histories, and roll back to a retained snapshot. Use tags to pin a dataset version.
Supported Features
PyPaimon supports filesystem, JDBC, and REST catalogs; local storage, HDFS, S3, OSS, and GCS; and batch reads and writes for append and primary-key tables. The guides cover projections, predicates, overwrite, incremental and streaming reads, BLOB data, row sharding, and table version management.
Primary-key reads support deduplicate, first-row, basic partial-update, and
aggregation with supported built-in aggregators. Support depends on the engine
configuration: for example, partial-update sequence groups and field aggregators
are not supported, and some delete/retract options are rejected. Do not assume
that every Java merge-engine option is available in Python.
For media and column updates, use data-evolution tables. For optional file formats and compute engines, see Installation.