Skip to main content

Local Cache

Paimon's local cache stores file blocks on disk or in memory to reduce repeated reads from storage such as S3, OSS, and HDFS. Configure it on the catalog that loads your tables.

ModeWhen it is selectedWhere blocks live
MemoryCache enabled without local-cache.dirProcess memory
DiskCache enabled with local-cache.dirThe configured local directory

Set an explicit size limit for your workload. The cache is disabled by default, and enabling it does not automatically cache data files: the default whitelist covers metadata and global indexes.

Cacheable reads check a local block cache. Hits return a cached block; misses fetch and cache a storage block. Other reads bypass the cache.

Cached File Types​

The cache classifies files by path. Set local-cache.whitelist to select the file types to cache.

File TypeConfig NameExamplesDefault Cached
METAmetasnapshot, schema, manifest, statistics, tagYes
GLOBAL_INDEXglobal-indexBTree and full-text index filesYes
BUCKET_INDEXbucket-indexHash, deletion vector index filesNo
DATAdataData files (ORC, Parquet, etc.)No
FILE_INDEXfile-indexData-file level bloom filter, bitmapNo

The default whitelist is meta,global-index, and * selects every file type. Files rewritten in place under a stable path always bypass the cache, whatever their type: the LATEST and EARLIEST hint files, tags, consumer and service files, _SUCCESS files, Iceberg-compatible metadata (v{N}.metadata.json, version-hint.text, retire-pending), and temporary files.

Enable Cache​

This is a catalog-level option. Configure it when creating the catalog:

import org.apache.paimon.catalog.Catalog;
import org.apache.paimon.catalog.CatalogContext;
import org.apache.paimon.catalog.CatalogFactory;
import org.apache.paimon.catalog.Identifier;
import org.apache.paimon.options.Options;
import org.apache.paimon.table.Table;

Options options = new Options();
options.set("warehouse", "s3://my-bucket/warehouse");
options.set("local-cache.enabled", "true");
// optional: use disk cache by specifying a directory
options.set("local-cache.dir", "/tmp/paimon-cache");
// Set a limit for cached blocks
options.set("local-cache.max-size", "2gb");
options.set("local-cache.block-size", "1mb");

CatalogContext context = CatalogContext.create(options);
try (Catalog catalog = CatalogFactory.createCatalog(context)) {
Table table = catalog.getTable(Identifier.create("my_db", "my_table"));
// Read table data here; eligible file reads use the cache.
}

The snippets assume the catalog and table already exist and storage credentials are configured. The Java snippet is a method-body fragment; declare or handle the catalog exceptions as in the Java setup.

Cache Options​

OptionTypeDefaultDescription
local-cache.enabledBooleanfalseWhether to enable local block cache for file reads.
local-cache.dirString(none)Directory for storing cached blocks on disk. If not configured, memory cache is used.
local-cache.max-sizeMemorySizeSee belowMaximum cached block bytes per cache manager. Least recently used blocks are evicted when the limit is exceeded.
local-cache.block-sizeMemorySize1 mbBlock size for caching. Files are logically divided into fixed-size blocks and cached independently.
local-cache.whitelistStringmeta,global-indexComma-separated list of file types to cache. Supported values: meta, global-index, bucket-index, data, file-index, or * for all of them.

When local-cache.max-size is omitted, Java has no configured size limit. PyPaimon uses a 256 MiB memory limit or a 10 GiB disk limit. Set the option explicitly when you want the same limit across clients. The limit accounts for cached block bytes, not all reader buffers or process memory.

How It Works​

  1. The reader requests bytes from an eligible file. The cache maps the byte range to fixed-size blocks, using 1 MiB blocks by default.
  2. A hit returns the cached block. A miss reads the block from storage and adds it to the cache.
  3. Least recently used blocks are evicted when a configured limit is exceeded.

A cache hit avoids fetching that block's contents again; it does not eliminate all remote metadata operations. In Java disk mode, opening a file still obtains its status. The disk cache identifies a file by path, length, and modification time, and separates entries by block size. Existing entries can be reused after restart when this identity matches. Memory caches do not survive process restarts.

Caching applies to reads. It does not change commit semantics or replace durable table storage.

Cache Lifecycle​

Java applications and distributed workers​

A catalog creates the cache used by the tables it loads. After a catalog-created FileIO is serialized and deserialized on a worker, it retains its cache configuration and lazily creates or reuses a cache manager in that worker JVM.

Deserialized wrappers with matching directory, maximum size, and block size share a JVM-local cache manager and its size limit. Memory entries remain isolated by the originating wrapper's namespace. The limit is local to that manager, not a cluster-wide budget. In disk mode the directory must be available and writable on each worker that uses it.

Close readers and other owned resources when finished. Closing a deserialized FileIO releases its reference to the shared manager; the last reference releases the shared entry. A manually constructed CachingFileIO without serialized cache configuration cannot recreate a cache after deserialization.

PyPaimon workers​

PyPaimon's cache is removed when CachingFileIO is pickled. An unpickled copy reads directly from storage. Create a catalog with cache options on each Python worker when you need caching there. The Java worker lifecycle above does not apply to Python serialization.