Scan#
Bucket pruning#
For fixed-bucket append and primary-key tables, an equality predicate on every bucket key lets the scan derive the target bucket using the table’s bucket function. Other buckets are excluded from the scan plan without requiring an explicit bucket ID from the caller. An explicit bucket filter takes precedence. Queries that do not constrain all bucket keys with equality, and bucket-unaware tables, keep the existing scan behavior. Both scan types use a shared selector that computes the bucket with each manifest entry’s total bucket count, so rescaled files are not filtered using the current table’s bucket count. Historical schemas remain eligible when their ordered bucket-key field IDs, types and bucket function match. Changes to unrelated columns do not disable pruning. Incompatible bucket schemas and nonpositive total bucket counts retain the existing filtering behavior.
This inference prunes data files at the manifest-entry level. Manifest min/max-bucket skipping requires an explicit bucket filter. When the snapshot live-manifest-entry cache is enabled, inferred scans cache candidates by bucket, current bucket count and current schema ID. Files with other bucket counts or schema IDs remain in the cached candidates and are filtered after lookup. This permits cache reuse without discarding files that require a different bucket calculation or schema fallback.
Decimal literals are rescaled to the bucket field’s type only when the conversion is exact. NaN literals and decimals that cannot be represented exactly disable inferred bucket pruning.
Interface#
-
class TableScan#
A scanner interface for reading table’s meta and create a plan.
Public Functions
-
virtual ~TableScan() = default#
-
virtual Result<std::shared_ptr<Plan>> CreatePlan() = 0#
Create a scan plan.
- Returns:
A Result containing a shared pointer to the created
Planor an error status.
-
virtual std::shared_ptr<Metrics> GetMetrics() const#
Retrieve metrics related to scan planning operations.
- Returns:
A point-in-time snapshot of scan metrics. Mutating the returned object does not affect metrics collected by this scan.
-
virtual Result<std::vector<std::map<std::string, std::string>>> ListPartitions() const#
Lists the partitions the scan can see, each as its partition values keyed by field name, in a stable order.
A table with no partition keys has none.
Only a format table answers it today: its partitions are the directories a scan descends, so listing them costs no more than planning. Every other table type returns
NotImplemented.- Returns:
A Result containing the partitions, or an error status.
Public Static Functions
-
static Result<std::unique_ptr<TableScan>> Create(std::unique_ptr<ScanContext> context)#
Create an instance of
TableScan.- Parameters:
context – A unique pointer to the
ScanContextused for scan operations.- Returns:
A Result containing a unique pointer to the
TableScaninstance.
-
virtual ~TableScan() = default#
-
class ScanContextBuilder#
ScanContextBuilderused to build aScanContext, has input validation.Public Functions
-
explicit ScanContextBuilder(const std::string &path)#
Constructs a
ScanContextBuilderwith required parameters.- Parameters:
path – The root path of the table.
Constructs a
ScanContextBuilderfor a format table that is already loaded: the only way to scan one whose schema lives in a metastore rather than under its location, such as a table a REST catalog serves.The table carries what such a location does not say, so
SetTableSchema()andWithFileSystem()are refused here rather than ignored.- Parameters:
table – The format table to scan, as
Catalog::GetFormatTable()hands it back.
-
~ScanContextBuilder()#
-
ScanContextBuilder &SetLimit(int32_t limit)#
If limit is not set, it defaults to unlimited.
-
ScanContextBuilder &SetBucketFilter(int32_t bucket_filter)#
Set a bucket filter to scan only specific bucket.
-
ScanContextBuilder &SetPartitionFilter(const std::vector<std::map<std::string, std::string>> &partition_filters)#
partition_filters in vector is supposed to be OR, filter in map is supposed to be AND, e.g., partition_filters is {{k1=1,k2=10}, {k1=2,k2=20}} => OR(AND(k1=1,k2=10), AND(k1=2,k2=20))
Set a predicate for filtering data.
Sets the result of a global index search (e.g., row ids (may with scores) from a distributed index lookup).
This is used to push down index-filtered row ids into the scan for efficient data retrieval.
Enables process-local union reads with the real-time stores owned by
realtime_context.
-
ScanContextBuilder &AddOption(const std::string &key, const std::string &value)#
The options added or set in
ScanContextBuilderhave high priority and will be merged with the options in table schema.
-
ScanContextBuilder &SetOptions(const std::map<std::string, std::string> &options)#
Set a configuration options map to set some option entries which are not defined in the table schema or whose values you want to overwrite.
Note
The options map will clear the options added by
AddOption()before.- Parameters:
options – The configuration options map.
- Returns:
Reference to this builder for method chaining.
-
ScanContextBuilder &WithStreamingMode(bool is_streaming_mode)#
Set whether the scan is in streaming mode.
Note
if not set, is_streaming_mode = false
Set custom memory pool for memory management.
Note
if not set, memory_pool is default pool
- Parameters:
memory_pool – The memory pool to use.
- Returns:
Reference to this builder for method chaining.
Set custom executor for task execution.
Note
if not set, executor is default executor
- Parameters:
executor – The executor to use.
- Returns:
Reference to this builder for method chaining.
Sets a custom file system instance to be used for all file operations in this scan context.
This bypasses the global file system registry and uses the provided implementation directly.
Note
If not set, use default file system (configured in
Options::FILE_SYSTEM)- Parameters:
file_system – The file system to use.
- Returns:
Reference to this builder for method chaining.
Plans a native table through its own file system - including the per-table temporary credentials a catalog that issues them hands out through
Catalog::GetTableFileSystem.This is a shorthand for
WithFileSystem(catalog->GetTableFileSystem(identifier)); an explicitWithFileSystem()takes precedence, so the catalog is not asked.- Parameters:
catalog – Non-null catalog, read when
Finish()builds the context.identifier – The native table to scan.
- Returns:
Reference to this builder for method chaining.
-
ScanContextBuilder &SetTableSchema(const std::string &table_schema)#
Set the table schema as a string to avoid schema loading I/O operations.
This optimization allows the scanner to use a pre-loaded schema instead of reading it from the table metadata, which can improve performance especially in scenarios with many small scan operations.
Note
The user must ensure that the schema string is valid and matches the table.
Note
If not set, the schema will be loaded from the table path.
- Parameters:
table_schema – String representation of the table schema.
- Returns:
Reference to this builder for method chaining.
Inject a cache for scan operations.
Passing nullptr disables cache.
- Returns:
Reference to this builder for method chaining.
-
Result<std::unique_ptr<ScanContext>> Finish()#
Build and return a
ScanContextinstance with input validation.- Returns:
Result containing the constructed
ScanContextor an error status.
-
explicit ScanContextBuilder(const std::string &path)#
-
class ScanContext#
ScanContextis some configuration for table scan operations.Please do not use this class directly, use
ScanContextBuilderto build aScanContextwhich has input validation.See also
Public Functions
-
~ScanContext()#
-
inline const std::string &GetPath() const#
-
inline bool IsStreamingMode() const#
-
inline std::optional<int32_t> GetLimit() const#
-
inline std::shared_ptr<ScanFilter> GetScanFilters() const#
-
inline const std::map<std::string, std::string> &GetOptions() const#
-
inline std::shared_ptr<MemoryPool> GetMemoryPool() const#
-
inline std::shared_ptr<GlobalIndexResult> GetGlobalIndexResult() const#
-
inline std::shared_ptr<RealtimeContext> GetRealtimeContext() const#
Returns the optional process-local context used to plan real-time memory reads.
-
inline std::shared_ptr<FileSystem> GetSpecificFileSystem() const#
-
inline const std::optional<std::string> &GetSpecificTableSchema() const#
-
inline std::shared_ptr<Cache> GetCache() const#
-
inline const std::shared_ptr<FormatTable> &GetFormatTable() const#
The format table this context was built from, or null when it names a table path and the schema under that path says what kind of table it is.
-
~ScanContext()#
-
class Split#
An input split for reading operation.
Needed by most batch computation engines. Support Serialize and Deserialize.
This split can be a
DataSplit(for direct data file reads), anIndexedSplit(for reads leveraging global indexes), or aFormatDataSplit(for a format table’s plain data files). Only the first two are serializable; aFormatDataSplitis in-memory only andSerialize()refuses it.Subclassed by paimon::DataSplit, paimon::IndexedSplit
Public Functions
-
virtual ~Split() = default#
Public Static Functions
Deserialize a
Splitfrom a binary buffer.Creates a
Splitinstance from its serialized binary representation. This is typically used in distributed computing scenarios where splits are transmitted between different nodes or processes.
Serialize a
Splitto a binary string.Converts a
Splitinstance to its binary representation for storage or transmission. The serialized data can later be deserialized using the Deserialize method.- Parameters:
split – The
Splitinstance to serialize.pool – Memory pool for allocating temporary objects during serialization.
- Returns:
Result containing the serialized binary data as a string or an error status.
-
virtual ~Split() = default#
-
class DataSplit : public paimon::Split#
Input data split for reading operation. Needed by most batch computation engines.
Public Functions
-
virtual int32_t Bucket() const = 0#
Get the bucket id of this data split.
-
virtual std::vector<SimpleDataFileMeta> GetFileList() const = 0#
Get the list of metadata for all data files in this split.
Note
This method will be removed in future versions and is only used for append tables.
-
struct SimpleDataFileMeta#
Metadata structure for simple data files.
Contains essential information about a data file including its location, size, row count, sequence numbers, schema information, and timestamps. This structure is used to track file metadata without loading the actual file content.
Public Functions
-
inline SimpleDataFileMeta(const std::string &_file_path, int64_t _file_size, int64_t _row_count, int64_t _min_sequence_number, int64_t _max_sequence_number, int64_t _schema_id, int32_t _level, const Timestamp &_creation_time, const std::optional<int64_t> &_delete_row_count)#
-
bool operator==(const SimpleDataFileMeta &other) const#
-
std::string ToString() const#
Public Members
-
std::string file_path#
Absolute path of the data file.
If external path is enabled,
file_pathindicates the actual location in the external storage system.
-
int64_t file_size#
-
int64_t row_count#
-
int64_t min_sequence_number#
-
int64_t max_sequence_number#
-
int64_t schema_id#
-
int32_t level#
-
std::optional<int64_t> delete_row_count#
-
inline SimpleDataFileMeta(const std::string &_file_path, int64_t _file_size, int64_t _row_count, int64_t _min_sequence_number, int64_t _max_sequence_number, int64_t _schema_id, int32_t _level, const Timestamp &_creation_time, const std::optional<int64_t> &_delete_row_count)#
-
virtual int32_t Bucket() const = 0#
-
class ScanFilter#
Filter configuration for table scan operations.