Spark
Use Spark SQL, the DataFrame API, or Structured Streaming to work with Paimon tables. Start with a local example, configure your catalog, then choose the guide for the operation you need.
Start Here
- Quick Start: create a table, write an update, and read the result locally.
- Installation: match the Spark and Scala versions and install the connector.
- Catalogs: connect to a filesystem, Hive, JDBC, or REST catalog.
- Configuration: distinguish catalog, table, session, and operation options.
Choose a Guide
| Task | Read next |
|---|---|
| Create a table, view, or tag | SQL DDL |
| Change a schema or table property | Alter Tables |
| Work with catalog-managed Format Table partitions | Format Table Partitions |
| Query current state, time travel, or bounded changes | SQL Queries |
| Insert, overwrite, update, delete, or merge rows | SQL Writes |
| Import or export CSV, JSON, or Parquet | COPY INTO |
| Work from Scala | DataFrame API |
| Add columns during writes | Schema Evolution on Write |
| Build a streaming job | Structured Streaming |
| Recover a stream and retain unread data | Streaming Recovery |
| Compact, expire, migrate, or build indexes | Procedures |
Reference
- Data Types: type mappings and version-specific timestamp and geospatial behavior.
- Default Values: defaults for new writes and how to change them.
- SQL Functions: partition helpers, blob functions, and user-defined functions.
- Inspect and Maintain Tables: schemas, partitions, statistics, and metadata refresh.
Feature requirements are called out on each page. In particular, SQL time travel and streaming
reads require Spark 3.3+, INSERT ... BY NAME requires Spark 3.5+, and SQL-defined scalar
functions require Spark 4.0+.