Skip to main content

Schema

A schema file defines the fields, keys, and options of a table at a particular schema ID. Readers use schema IDs in snapshots and data file metadata to interpret records written before or after a schema change.

Schema ID and Format Version​

These two numbers serve different purposes:

  • id identifies a table schema. IDs start at 0; an update creates a new schema ID.
  • version identifies the schema JSON format. The current format version is 3.

In the default layout, schema ID 0 is stored in schema/schema-0. A file named schema-3 does not imply format version 3.

JSON Fields​

FieldTypeMeaning
versionIntegerSchema JSON format version. See Compatibility.
idLongSchema ID, used in the file name and metadata references.
fieldsArray of DataFieldOrdered table fields, including stable field IDs.
highestFieldIdIntegerHighest field ID allocated, including nested fields; used when allocating IDs for new fields.
partitionKeysArray of stringsNames of the partition fields.
primaryKeysArray of stringsNames of the primary-key fields; empty for a table without a primary key.
optionsMap of strings to stringsTable options. Map entry order has no meaning.
commentString, optionalTable comment.
timeMillisLongSchema creation time in milliseconds since the Unix epoch.

Example​

This example is schema ID 0, serialized with format version 3:

{
"version" : 3,
"id" : 0,
"fields" : [ {
"id" : 0,
"name" : "order_id",
"type" : "BIGINT NOT NULL"
}, {
"id" : 1,
"name" : "order_name",
"type" : "STRING"
}, {
"id" : 2,
"name" : "order_user_id",
"type" : "BIGINT"
}, {
"id" : 3,
"name" : "order_shop_id",
"type" : "BIGINT"
} ],
"highestFieldId" : 3,
"partitionKeys" : [ ],
"primaryKeys" : [ "order_id" ],
"options" : {
"bucket" : "5"
},
"comment" : "",
"timeMillis" : 1720496663041
}

Compatibility​

Older schema files can omit options whose defaults have since changed. Readers preserve their original behavior when decoding those schemas:

Schema formatMissing field or optionReader behavior
No version fieldversionTreat as format version 1.
Version 1bucketSupply bucket = 1.
Versions 1 and 2file.formatSupply file.format = orc.
Older files without a timestamptimeMillisUse 0.

These compatibility defaults do not replace options explicitly stored in the schema. New tables use current defaults, including Parquet as the default data file format.

DataField​

A DataField describes one column, including fields nested inside a row type.

FieldTypeMeaning
idIntegerStable field identifier used for schema evolution.
nameStringField name.
typeString or JSON type objectLogical type, including nullability and nested type information.
descriptionString, optionalField comment.
defaultValueString, optionalStored default value definition. Engine support determines how it is used.

Primitive types commonly use strings such as BIGINT NOT NULL. Nested types need their structured type representation. See Data Types for the logical type reference.

Update Schema​

A schema update writes a new schema file rather than replacing the old definition:

my_table/
└── schema/
├── schema-0
├── schema-1
└── schema-2

The latest schema defines the current table, while existing snapshots and data files can still reference earlier schema IDs. Retain schema files that are needed to interpret existing data. Use the supported Flink schema changes or Spark schema changes instead of editing schema JSON directly.