Skip to main content

Quick Start

Run a local Spark SQL session, create a primary key table, and read an update back. The examples use Spark 3.5 with Scala 2.12. Choose the matching connector for another version in Installation.

Preparation​

Install Spark and download the matching Paimon JAR from Installation. The connector's Spark minor version and Scala binary version must match your Spark distribution.

Setup​

Start a local session with the connector, catalog, and SQL extensions configured together:

spark-sql --master 'local[2]' \
--jars /path/to/paimon-spark-3.5_2.12-2.2-SNAPSHOT.jar \
--conf spark.sql.catalog.paimon=org.apache.paimon.spark.SparkCatalog \
--conf spark.sql.catalog.paimon.warehouse=file:/tmp/paimon \
--conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions

Replace the JAR path with your downloaded file. The local warehouse is for this single-machine example; a cluster needs storage accessible to the driver and every executor. See Catalogs for Hive, JDBC, REST, and SparkGenericCatalog setup.

Select the catalog and database:

USE paimon.default;

Use a fully qualified name such as spark_catalog.default.other_table to access a table in Spark's built-in catalog while paimon is selected.

Create Table​

CREATE TABLE my_table (
k INT,
v STRING
) TBLPROPERTIES (
'primary-key' = 'k',
'bucket' = '1'
);

A primary key table merges records with the same key. Without the primary-key property, Paimon creates an append table. See Create Tables, Views, and Tags for partitioned and external tables.

Insert Table​

INSERT INTO my_table VALUES (1, 'Hi'), (2, 'Hello');
INSERT INTO my_table VALUES (1, 'Hello Paimon');

The second write updates key 1 under the default deduplicate merge engine.

Query Table​

SELECT * FROM my_table ORDER BY k;
-- 1 Hello Paimon
-- 2 Hello

Next Steps​

TaskGuide
Read and write from ScalaDataFrame API
Filter data or query an older snapshotSQL Queries
Overwrite partitions, update, or merge rowsSQL Writes
Consume or produce streaming dataStructured Streaming
Tune the connector or table optionsConfiguration

Spark Type Conversion​

The mapping table and timestamp, variant, and geospatial compatibility notes are in Data Types.