Local Cache
Paimon's local cache stores file blocks on disk or in memory to reduce repeated reads from storage such as S3, OSS, and HDFS. Configure it on the catalog that loads your tables.
| Mode | When it is selected | Where blocks live |
|---|---|---|
| Memory | Cache enabled without local-cache.dir | Process memory |
| Disk | Cache enabled with local-cache.dir | The configured local directory |
Set an explicit size limit for your workload. The cache is disabled by default, and enabling it does not automatically cache data files: the default whitelist covers metadata and global indexes.
Cached File Types
The cache classifies files by path. Set local-cache.whitelist to select the file types to cache.
| File Type | Config Name | Examples | Default Cached |
|---|---|---|---|
| META | meta | snapshot, schema, manifest, statistics, tag | Yes |
| GLOBAL_INDEX | global-index | BTree and full-text index files | Yes |
| BUCKET_INDEX | bucket-index | Hash, deletion vector index files | No |
| DATA | data | Data files (ORC, Parquet, etc.) | No |
| FILE_INDEX | file-index | Data-file level bloom filter, bitmap | No |
The default whitelist is meta,global-index, and * selects every file type. Files rewritten in
place under a stable path always bypass the cache, whatever their type: the LATEST and EARLIEST
hint files, tags, consumer and service files, _SUCCESS files, Iceberg-compatible metadata
(v{N}.metadata.json, version-hint.text, retire-pending), and temporary files.
Enable Cache
This is a catalog-level option. Configure it when creating the catalog:
- Java
- Python
import org.apache.paimon.catalog.Catalog;
import org.apache.paimon.catalog.CatalogContext;
import org.apache.paimon.catalog.CatalogFactory;
import org.apache.paimon.catalog.Identifier;
import org.apache.paimon.options.Options;
import org.apache.paimon.table.Table;
Options options = new Options();
options.set("warehouse", "s3://my-bucket/warehouse");
options.set("local-cache.enabled", "true");
// optional: use disk cache by specifying a directory
options.set("local-cache.dir", "/tmp/paimon-cache");
// Set a limit for cached blocks
options.set("local-cache.max-size", "2gb");
options.set("local-cache.block-size", "1mb");
CatalogContext context = CatalogContext.create(options);
try (Catalog catalog = CatalogFactory.createCatalog(context)) {
Table table = catalog.getTable(Identifier.create("my_db", "my_table"));
// Read table data here; eligible file reads use the cache.
}
import pypaimon
options = {
"warehouse": "s3://my-bucket/warehouse",
"local-cache.enabled": "true",
# optional: use disk cache by specifying a directory
"local-cache.dir": "/tmp/paimon-cache",
# Set a limit for cached blocks
"local-cache.max-size": "2gb",
"local-cache.block-size": "1mb",
}
catalog = pypaimon.create_catalog(options)
# Eligible file reads from tables loaded by this catalog use the cache
table = catalog.get_table("db.my_table")
The snippets assume the catalog and table already exist and storage credentials are configured. The Java snippet is a method-body fragment; declare or handle the catalog exceptions as in the Java setup.
Cache Options
| Option | Type | Default | Description |
|---|---|---|---|
local-cache.enabled | Boolean | false | Whether to enable local block cache for file reads. |
local-cache.dir | String | (none) | Directory for storing cached blocks on disk. If not configured, memory cache is used. |
local-cache.max-size | MemorySize | See below | Maximum cached block bytes per cache manager. Least recently used blocks are evicted when the limit is exceeded. |
local-cache.block-size | MemorySize | 1 mb | Block size for caching. Files are logically divided into fixed-size blocks and cached independently. |
local-cache.whitelist | String | meta,global-index | Comma-separated list of file types to cache. Supported values: meta, global-index, bucket-index, data, file-index, or * for all of them. |
When local-cache.max-size is omitted, Java has no configured size limit. PyPaimon uses a 256 MiB
memory limit or a 10 GiB disk limit. Set the option explicitly when you want the same limit across
clients. The limit accounts for cached block bytes, not all reader buffers or process memory.
How It Works
- The reader requests bytes from an eligible file. The cache maps the byte range to fixed-size blocks, using 1 MiB blocks by default.
- A hit returns the cached block. A miss reads the block from storage and adds it to the cache.
- Least recently used blocks are evicted when a configured limit is exceeded.
A cache hit avoids fetching that block's contents again; it does not eliminate all remote metadata operations. In Java disk mode, opening a file still obtains its status. The disk cache identifies a file by path, length, and modification time, and separates entries by block size. Existing entries can be reused after restart when this identity matches. Memory caches do not survive process restarts.
Caching applies to reads. It does not change commit semantics or replace durable table storage.
Cache Lifecycle
Java applications and distributed workers
A catalog creates the cache used by the tables it loads. After a catalog-created FileIO is
serialized and deserialized on a worker, it retains its cache configuration and lazily creates or
reuses a cache manager in that worker JVM.
Deserialized wrappers with matching directory, maximum size, and block size share a JVM-local cache manager and its size limit. Memory entries remain isolated by the originating wrapper's namespace. The limit is local to that manager, not a cluster-wide budget. In disk mode the directory must be available and writable on each worker that uses it.
Close readers and other owned resources when finished. Closing a deserialized FileIO releases its
reference to the shared manager; the last reference releases the shared entry. A manually constructed
CachingFileIO without serialized cache configuration cannot recreate a cache after deserialization.
PyPaimon workers
PyPaimon's cache is removed when CachingFileIO is pickled. An unpickled copy reads directly from
storage. Create a catalog with cache options on each Python worker when you need caching there.
The Java worker lifecycle above does not apply to Python serialization.