Skip to main content

Spark

Use Spark SQL, the DataFrame API, or Structured Streaming to work with Paimon tables. Start with a local example, configure your catalog, then choose the guide for the operation you need.

Spark SQL, DataFrames, and Structured Streaming access Paimon through the connector and catalog, with table data stored in the warehouse.

Start Here​

  1. Quick Start: create a table, write an update, and read the result locally.
  2. Installation: match the Spark and Scala versions and install the connector.
  3. Catalogs: connect to a filesystem, Hive, JDBC, or REST catalog.
  4. Configuration: distinguish catalog, table, session, and operation options.

Choose a Guide​

TaskRead next
Create a table, view, or tagSQL DDL
Change a schema or table propertyAlter Tables
Work with catalog-managed Format Table partitionsFormat Table Partitions
Query current state, time travel, or bounded changesSQL Queries
Insert, overwrite, update, delete, or merge rowsSQL Writes
Import or export CSV, JSON, or ParquetCOPY INTO
Work from ScalaDataFrame API
Add columns during writesSchema Evolution on Write
Build a streaming jobStructured Streaming
Recover a stream and retain unread dataStreaming Recovery
Compact, expire, migrate, or build indexesProcedures

Reference​

Feature requirements are called out on each page. In particular, SQL time travel and streaming reads require Spark 3.3+, INSERT ... BY NAME requires Spark 3.5+, and SQL-defined scalar functions require Spark 4.0+.