Connecting Engines
An engine needs to discover a Paimon table and read its files. Use the same catalog backend and table location as the writer, then configure storage access in the engine that runs the query. The SQL catalog name is local to that engine and can differ between engines.
Identify the Catalog and Storage
Collect these settings from the job or engine that created the table:
| Setting | What to check |
|---|---|
| Catalog backend | Filesystem, Hive metastore, REST, or another backend supported by the reader. |
| Catalog endpoint | The metastore or service URI, with the authentication needed to connect. A filesystem catalog has no separate metastore service. |
| Warehouse or table location | The shared storage URI. A filesystem catalog uses the warehouse root; a Hive external table points to one table directory. REST catalogs can use a warehouse identifier. |
| Storage access | Filesystem libraries, endpoint configuration, and credentials available to the engine processes that access the files. |
| Table features | Primary keys, bucket mode, file format, types, and features that the reader must understand. |
See Catalog for Paimon's catalog backends and Filesystems for Paimon storage dependencies. External engines can use their own filesystem implementations and configuration names.
Configure the Reader
- Install the connector or enable the engine's built-in Paimon integration.
- Configure a catalog using the engine's own property names.
- Configure access to the catalog service and the warehouse storage on the relevant nodes.
- Query an existing table by its fully qualified name before adding optional read settings.
Catalog properties are not interchangeable. For example, a Hive-backed Paimon catalog uses
metastore = hive in Flink, while StarRocks and Doris use paimon.catalog.type = hive and
paimon.catalog.type = hms, respectively. Follow the matching guide:
Hive, Trino,
StarRocks, or Doris.
For a local experiment, a file: warehouse is sufficient when all processes share that path.
For a distributed cluster, use shared storage that every participating process can access.
Check the Read Semantics
Query a small, known table first. Compare values as well as row counts, especially for primary-key tables with updates and deletes.
| Read path | What the result represents |
|---|---|
| Ordinary batch query | A table snapshot, subject to the engine's metadata caching and supported read mode. |
| Time travel | A selected historical snapshot or tag; syntax and support depend on the engine. The referenced state must still be retained. |
| Read-optimized query | Compacted data from a primary-key table; recent changes may be absent until full compaction completes. |
| Streaming query | A snapshot and/or subsequent changes, according to the engine and scan configuration. |
See Table Mode, System Tables, and Snapshot Retention for the underlying behavior.
Troubleshooting
| Symptom | Check first |
|---|---|
| Catalog exists but the table is missing | Confirm the backend, warehouse, database, and table registration match the writer. |
| Tables can be listed but a query cannot open files | Check storage endpoints, credentials, and filesystem dependencies on the nodes performing the read. |
| Class-loading or method-not-found errors | Check the engine/connector version pair and remove conflicting connector jars. |
| New data or schema is not visible | Check whether the writer committed a snapshot, then inspect engine metadata caches and any time-travel or read-optimized settings. |
| Updates or deletes produce unexpected rows | Confirm support for the table's merge engine and deletion vectors; use a Paimon-aware reader rather than scanning the underlying Parquet or ORC files directly. |
Continue with the engine guide for exact SQL and version-specific limitations.