Quick Start
Run a local Spark SQL session, create a primary key table, and read an update back. The examples use Spark 3.5 with Scala 2.12. Choose the matching connector for another version in Installation.
Preparation
Install Spark and download the matching Paimon JAR from Installation. The connector's Spark minor version and Scala binary version must match your Spark distribution.
Setup
Start a local session with the connector, catalog, and SQL extensions configured together:
spark-sql --master 'local[2]' \
--jars /path/to/paimon-spark-3.5_2.12-2.2-SNAPSHOT.jar \
--conf spark.sql.catalog.paimon=org.apache.paimon.spark.SparkCatalog \
--conf spark.sql.catalog.paimon.warehouse=file:/tmp/paimon \
--conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions
Replace the JAR path with your downloaded file. The local warehouse is for this single-machine
example; a cluster needs storage accessible to the driver and every executor. See
Catalogs for Hive, JDBC, REST, and SparkGenericCatalog setup.
Select the catalog and database:
USE paimon.default;
Use a fully qualified name such as spark_catalog.default.other_table to access a table in
Spark's built-in catalog while paimon is selected.
Create Table
CREATE TABLE my_table (
k INT,
v STRING
) TBLPROPERTIES (
'primary-key' = 'k',
'bucket' = '1'
);
A primary key table merges records with the same key. Without the primary-key property,
Paimon creates an append table. See Create Tables, Views, and Tags for partitioned
and external tables.
Insert Table
INSERT INTO my_table VALUES (1, 'Hi'), (2, 'Hello');
INSERT INTO my_table VALUES (1, 'Hello Paimon');
The second write updates key 1 under the default deduplicate merge engine.
Query Table
SELECT * FROM my_table ORDER BY k;
-- 1 Hello Paimon
-- 2 Hello
Next Steps
| Task | Guide |
|---|---|
| Read and write from Scala | DataFrame API |
| Filter data or query an older snapshot | SQL Queries |
| Overwrite partitions, update, or merge rows | SQL Writes |
| Consume or produce streaming data | Structured Streaming |
| Tune the connector or table options | Configuration |
Spark Type Conversion
The mapping table and timestamp, variant, and geospatial compatibility notes are in Data Types.