File Format#
Interface#
-
class FileFormat#
FileFormatis used to createReaderBuilderandWriterBuilder.Public Functions
-
virtual ~FileFormat() = default#
-
virtual const std::string &Identifier() const = 0#
- Returns:
The corresponding identifier of file format, e.g., orc.
-
virtual Result<std::unique_ptr<ReaderBuilder>> CreateReaderBuilder(int32_t batch_size) const = 0#
- Returns:
A reader builder which will create reader with specific batch size.
-
virtual Result<std::unique_ptr<WriterBuilder>> CreateWriterBuilder(::ArrowSchema *schema, int32_t batch_size) const = 0#
- Returns:
A
WriterBuilderof the corresponding schema, or error status when schema is invalid.
-
virtual Result<std::unique_ptr<FormatStatsExtractor>> CreateStatsExtractor(::ArrowSchema *schema) const = 0#
- Returns:
A
FormatStatsExtractorof current file format.
-
virtual ~FileFormat() = default#
-
class FormatWriter#
File format writer, each writer corresponds to a data file.
Public Functions
-
virtual ~FormatWriter() = default#
-
virtual Status AddBatch(::ArrowArray *batch) = 0#
Add a batch of records to the format writer.
Note
The batch must conform to the schema expected by the writer.
Note
This method can be called multiple times to write data incrementally.
Note
After calling
Finish(), this method should not be called again.- Parameters:
batch – Pointer to an ArrowArray containing the batch data to write.
- Returns:
Status indicating success (OK) or failure with error information.
-
virtual Status Flush() = 0#
Flushes all intermediate buffered data to the format writer.
- Returns:
Error status returned if the encoder cannot be flushed, or if the output stream return an error.
-
virtual Status Finish() = 0#
Finishes the writing.
This must flush all internal buffer, finish encoding, and write footers.
Note
The writer is not expected to handle any more records via
AddBatch()after this method is called.Warning
This method MUST NOT close the stream that the writer writes to. Closing the stream is expected to happen through the invoker of this method afterwards.
- Returns:
Error status returned if the finalization fails.
-
virtual Result<bool> ReachTargetSize(bool suggested_check, int64_t target_size) const = 0#
Check if the writer has reached the
target_size.- Parameters:
suggested_check – Whether it needs to be checked, but subclasses can also decide whether to check it themselves.
target_size – The size of the target.
- Returns:
True if the target size was reached, otherwise false.
- Returns:
Error status returned if calculating the length fails.
-
virtual std::shared_ptr<Metrics> GetWriterMetrics() const = 0#
Get metrics of the writer.
- Returns:
The accumulated writer metrics to current state.
-
virtual ~FormatWriter() = default#
-
class WriterBuilder#
Create a file format writer based on the file output stream. Allows you to specify memory pool.
Subclassed by paimon::DirectWriterBuilder, paimon::SpecificFSWriterBuilder
Public Functions
-
virtual ~WriterBuilder() = default#
Set memory pool to use.
Build a file format writer based on the file output stream and file compression.
-
virtual ~WriterBuilder() = default#
-
class FileBatchReader : public paimon::BatchReader#
The batch reader for a single file supports returning the line number of the last batch read for deletion vector judgment.
Subclassed by paimon::PrefetchFileBatchReader
Public Functions
-
virtual Result<std::unique_ptr<::ArrowSchema>> GetFileSchema() const = 0#
- Returns:
The schema of the file.
Resets the read schema and predicate.
If
SetReadSchema()is not called,NextBatch()will return data with the file schema. After resetting the read schema,NextBatch()will read data starting from the first row.- Parameters:
read_schema – The schema to set for reading.
predicate – The predicate to apply for filtering data.
selection_bitmap – The bitmap to apply for filtering data.
- Returns:
The status of the operation.
-
virtual Result<uint64_t> GetPreviousBatchFileRowId(uint64_t batch_row_id) const = 0#
Get the file-level row ID for a given batch-relative row index in the previously read batch.
- Parameters:
batch_row_id – Zero-based index within the current batch.
- Returns:
The corresponding file-level row ID, or Status::Invalid if no batch has been read yet, the last batch was EOF, or batch_row_id is out of range.
-
virtual Result<uint64_t> GetNumberOfRows() const = 0#
Get the number of rows in the file.
-
virtual bool SupportPreciseBitmapSelection() const = 0#
Get whether or not support read precisely while bitmap pushed down.
-
inline virtual void Warmup()#
Starts whatever background work this reader would otherwise start on its first read, so a caller that knows this reader is next can pay that startup while still consuming the previous one.
This is an optional hint and never changes what the reader returns: ordering, filtering and metrics are the same with or without it. A reader with nothing to start, or one that is not yet ready to start it, does nothing. Calling it before the reader is configured (for example before
SetReadSchema()), or afterClose(), is such a no-op.It reports no error on purpose: the caller may warm up a reader it never ends up reading, and a hint about a file nobody reads must not fail the read in progress. An implementation that cannot start its work leaves it to be started by the first read, which reports the failure itself.
What an implementation starts is up to the implementation - a reader that prefetches on a background thread starts that thread, one that reads synchronously may do nothing at all.
Warning
Whatever it starts, the call itself is not thread-safe: make it from the same thread that reads this reader, and not concurrently with any other call on it.
-
virtual Result<ReadBatch> NextBatch() = 0#
Retrieves the next batch of data.
If EOF is reached, returns an OK status with a nullptr array. Returns an error status only for critical failures (e.g., IO errors). Once an error is returned, this method must not be retried, as it will repeatedly return the same error code.
Warning
A non-EOF ArrowArray and all its nested child arrays must have offset 0 to avoid potential issues during conversion through the Arrow C Data Interface.
Warning
A returned ArrowArray must retain every allocator and plugin resource needed by its release callback, so it remains releasable after this reader is destroyed.
Warning
Consumers must treat the returned ArrowArray and ArrowSchema as one complete Arrow C Data Interface ownership unit. Moving or retaining an individual child ArrowArray without its root array is unsupported because resource lifetimes are retained by the root array’s release chain.
- Returns:
A result containing a
ReadBatch, which consists of a unique pointer toArrowArrayand a unique pointer toArrowSchema. Returned array contains a_VALUE_KINDfield (the first field) to indicate the row kind of each row. Deleted or index-filtered rows are removed.
-
virtual Result<ReadBatchWithBitmap> NextBatchWithBitmap()#
Retrieves the next batch of data.
If EOF is reached, returns an OK status with a nullptr array. Returns an error status only for critical failures (e.g., IO errors). Once an error is returned, this method must not be retried, as it will repeatedly return the same error code.
Warning
A non-EOF ArrowArray and all its nested child arrays must have offset 0 to avoid potential issues during conversion through the Arrow C Data Interface.
Warning
A returned ArrowArray must retain every allocator and plugin resource needed by its release callback, so it remains releasable after this reader is destroyed.
Warning
Consumers must treat the returned ArrowArray and ArrowSchema as one complete Arrow C Data Interface ownership unit. Moving or retaining an individual child ArrowArray without its root array is unsupported because resource lifetimes are retained by the root array’s release chain.
- Returns:
A result containing a
ReadBatchand a valid bitmap.ReadBatchconsists of a unique pointer toArrowArrayand a unique pointer toArrowSchema. Returned array contains a _VALUE_KIND field (the first field) to indicate the row kind of each row. Deleted or index-filtered records maybe maintained inReadBatch, while bitmap indicates valid row id. If deletion vector or index are enabled, this function is more efficient thanNextBatch(). The default implementation callsNextBatch()and adds all rows to valid bitmap. Noted that the returned bitmap has at least one valid row id.
-
virtual Result<std::unique_ptr<::ArrowSchema>> GetFileSchema() const = 0#
-
class ReaderBuilder#
Create a file batch reader based on an input stream. Allows you to specify memory pool.
Public Functions
-
virtual ~ReaderBuilder() = default#
Set memory pool to use.
Inject a cache for reader-specific immutable metadata.
-
inline virtual ReaderBuilder *WithReadHints(const std::optional<ReadHints> &hints)#
Inject runtime read state from the framework layer, so the format can adapt its internal behavior accordingly.
When present, the hints describe the authoritative runtime state of this read; when absent, the format should fall back to its own options.
Build a file batch reader based on the created
InputStream.Every non-EOF ArrowArray returned by the reader must retain all allocator and plugin resources needed by its release callback. The array must remain releasable after the reader has been destroyed.
-
virtual ~ReaderBuilder() = default#
-
class FileFormatFactory : public paimon::Factory#
A factory for creating
FileFormatinstances.Public Functions
-
~FileFormatFactory() override#
-
virtual Result<std::unique_ptr<FileFormat>> Create(const std::map<std::string, std::string> &options) const = 0#
Create a
FileFormatwith the corresponding options.
Public Static Functions
-
static Result<std::unique_ptr<FileFormat>> Get(const std::string &identifier, const std::map<std::string, std::string> &options)#
Get
FileFormatcorresponding to identifier.- Pre:
Factory is already registered.
-
~FileFormatFactory() override#
-
class FormatStatsExtractor#
Extracts statistics directly from file.
Public Functions
-
virtual ~FormatStatsExtractor() = default#
Extracts statistics for each column of a data file based on the file path and file system.
Extracts statistics for each column and
FileInfoof a data file based on the file path and file system.
-
class FileInfo#
File info fetched from physical file, currently only include row count.
-
virtual ~FormatStatsExtractor() = default#
-
class WrittenFileStatsProvider#
Optionally implemented by a
FormatWriterthat can produce the column statistics of the file it has finished from the metadata it still holds in memory.The data file writers then take the statistics from it instead of reopening the file through
FormatStatsExtractorto read them back.It is implemented alongside
FormatWriterrather than added to it, soFormatWriterkeeps its ABI and format writers that do not implement it are unaffected.Public Functions
-
virtual ~WrittenFileStatsProvider() = default#
Extracts the statistics of each top-level column of the finished file.
They must equal what
FormatStatsExtractor::Extract()reads back from that file.- Parameters:
pool – Memory pool used to build the statistics.
- Returns:
The statistics, or an error status if the file has not been finished successfully.
-
virtual ~WrittenFileStatsProvider() = default#