Maintenance
Use these guides to manage table history, clean up stored data, run compaction jobs, and tune and monitor workloads. Storage setup and configuration references are also collected here.
Choose a Task
| Task | Start here |
|---|---|
| Reduce storage usage | Choose a cleanup operation |
| Keep daily versions available for queries | Automatic tag creation and retention |
| Recover from an incorrect write | Choose a recovery operation |
| Validate changes on a separate branch | Manage Branches |
| Run compaction separately from writers | Dedicated Compaction |
| Investigate slow writes or memory pressure | Choose metrics, then see Write Performance |
| Connect to HDFS or an object store | Filesystems |
| Look up a table, catalog, or connector option | Configurations |
Clean Up Stored Data
Choose the operation based on what you want to remove:
- Expire partitions to remove old partitions from the latest table state. Physical file deletion depends on snapshot expiration.
- Expire snapshots to limit retained history and remove files that are no longer needed. Review the retention requirements of batch queries and streaming readers before changing the policy.
- Remove orphan files to clean up files that are no longer referenced. Follow the cleanup guide's age cutoff to account for files being added by active writers.
Tags preserve historical data independently of snapshot expiration. Review tag retention as part of
your storage policy. To remove empty directories left after file deletion, see the
snapshot.clean-empty-directories option in Expire Snapshots.
Recover Table Data
Choose a recovery mode, then follow the instructions for a snapshot ID or a tag. The linked guides include the supported engine commands and operation-specific limitations.
| Recovery mode | Effect on table history | Instructions |
|---|---|---|
| Rollback | Restore the target state and remove snapshots and tags after the target. | Snapshot or tag |
| Rollback as latest | Restore the target state as a new latest snapshot, preserving later snapshots and tags. | Snapshot or tag |
For a workflow that validates corrected data on a separate branch before updating the main branch, see Manage Branches and Fast Forward.
Browse by Topic
Data Lifecycle & Versioning
- Manage Snapshots: retain or expire table history, roll back, and remove orphan files.
- Manage Tags: preserve named versions, automate tag retention, and restore tagged data.
- Manage Branches: create isolated branches, validate changes, and fast-forward the main branch.
- Manage Partitions: expire partitions and mark partitions ready for downstream consumers.
Compaction & Data Layout
- Dedicated Compaction: run table or database compaction jobs and select compaction targets.
- Rescale Bucket: change bucket counts and reorganize existing data.
Performance & Monitoring
- Write Performance: tune parallelism, buffering, file formats, and memory usage.
- Metrics: access metrics through Flink and choose metrics for investigation.
Storage & Configuration
- Filesystems: install filesystem dependencies and configure storage access.
- Configurations: look up table, catalog, connector, and file-format options.