Skip to main content

Connecting Engines

An engine needs to discover a Paimon table and read its files. Use the same catalog backend and table location as the writer, then configure storage access in the engine that runs the query. The SQL catalog name is local to that engine and can differ between engines.

An engine resolves a table through a catalog, then reads Paimon metadata and data files from shared storage. Catalog and storage access are configured separately.

Identify the Catalog and Storage​

Collect these settings from the job or engine that created the table:

SettingWhat to check
Catalog backendFilesystem, Hive metastore, REST, or another backend supported by the reader.
Catalog endpointThe metastore or service URI, with the authentication needed to connect. A filesystem catalog has no separate metastore service.
Warehouse or table locationThe shared storage URI. A filesystem catalog uses the warehouse root; a Hive external table points to one table directory. REST catalogs can use a warehouse identifier.
Storage accessFilesystem libraries, endpoint configuration, and credentials available to the engine processes that access the files.
Table featuresPrimary keys, bucket mode, file format, types, and features that the reader must understand.

See Catalog for Paimon's catalog backends and Filesystems for Paimon storage dependencies. External engines can use their own filesystem implementations and configuration names.

Configure the Reader​

  1. Install the connector or enable the engine's built-in Paimon integration.
  2. Configure a catalog using the engine's own property names.
  3. Configure access to the catalog service and the warehouse storage on the relevant nodes.
  4. Query an existing table by its fully qualified name before adding optional read settings.

Catalog properties are not interchangeable. For example, a Hive-backed Paimon catalog uses metastore = hive in Flink, while StarRocks and Doris use paimon.catalog.type = hive and paimon.catalog.type = hms, respectively. Follow the matching guide: Hive, Trino, StarRocks, or Doris.

For a local experiment, a file: warehouse is sufficient when all processes share that path. For a distributed cluster, use shared storage that every participating process can access.

Check the Read Semantics​

Query a small, known table first. Compare values as well as row counts, especially for primary-key tables with updates and deletes.

Read pathWhat the result represents
Ordinary batch queryA table snapshot, subject to the engine's metadata caching and supported read mode.
Time travelA selected historical snapshot or tag; syntax and support depend on the engine. The referenced state must still be retained.
Read-optimized queryCompacted data from a primary-key table; recent changes may be absent until full compaction completes.
Streaming queryA snapshot and/or subsequent changes, according to the engine and scan configuration.

See Table Mode, System Tables, and Snapshot Retention for the underlying behavior.

Troubleshooting​

SymptomCheck first
Catalog exists but the table is missingConfirm the backend, warehouse, database, and table registration match the writer.
Tables can be listed but a query cannot open filesCheck storage endpoints, credentials, and filesystem dependencies on the nodes performing the read.
Class-loading or method-not-found errorsCheck the engine/connector version pair and remove conflicting connector jars.
New data or schema is not visibleCheck whether the writer committed a snapshot, then inspect engine metadata caches and any time-travel or read-optimized settings.
Updates or deletes produce unexpected rowsConfirm support for the table's merge engine and deletion vectors; use a Paimon-aware reader rather than scanning the underlying Parquet or ORC files directly.

Continue with the engine guide for exact SQL and version-specific limitations.