Ch. 19 · System Design

System Design: Schema Evolution

Evolve stored and message schemas with additive changes, expand-contract migrations, and reader tolerance across services.

~2 min readadvancedupdated Oct 5, 2026

Data outlives code. A schema change deploys while old readers, queued messages and cached documents still exist, so the safe path is to change the schema in steps that are compatible with every version running at once.

Before you start

You should understand migrations, events and deployment. This article is conceptual with a short migration sketch.

Step-by-step walkthrough

Step 1: Add before you remove

Adding an optional field, a new column or a new event type is backward compatible: old readers ignore it, and new readers handle its absence with a default. Renames and type changes are breaking, because they remove or reinterpret what a reader expects.

Step 2: Use expand-contract for breaking changes

To rename or restructure, add the new shape (expand), backfill and migrate readers, then stop writing the old shape and remove it (contract). Each step is compatible with the fleet at that moment, so no deploy breaks a running version.

Step 3: Make readers tolerant

Readers should ignore unknown fields and default missing ones, so a producer can add data before consumers catch up. This tolerance is what lets producers and consumers deploy independently, which is the point of evolving a schema carefully.

Worked scenario

The rename happens in compatible steps.

1. expand:  add column full_name, keep name; write both
2. migrate: backfill full_name from name; readers read full_name
3. contract: stop writing name, remove it after all writers updated
Text

Walk through the example

At every stage, the old and new code agree on a readable column. Writing both during the expand phase keeps both readers working; migrating readers lets the old column be dropped safely. No single deploy requires all services to change at once, so a rollback at any step is safe.

Common mistake

Renaming a column or field in place and deploying, which breaks every reader that has not updated. Another is a consumer that rejects unknown fields, so adding a field on the producer side fails the consumer.

Verify the behavior

Run the old and new reader versions against the data during the expand phase and confirm both work. Send a message with an extra field and confirm the old consumer ignores it. Assert the contract step only happens after telemetry shows no readers use the old field.

Interview exercise

Why is adding a required field a breaking change while adding an optional one is not?

Answer and reasoning

An optional field can be absent, so old readers that never set it still produce valid data, and new readers default it. A required field must be present, so every producer must be updated before readers can rely on it, which couples the fleet. Keep new fields optional until all writers populate them, then treat them as required.

Continue learning

Compare contract handling in Contract evolution and Normalization trade-offs. Read the Avro schema evolution documentation and try the System design interview questions.

More in System Design

read ✓System Design · hard

System Design: Bloom Filters

Use a Bloom filter to skip lookups with a tiny memory footprint, and understand its false-positive-only guarantee.

~2 min readread →
read ✓System Design · hard

System Design: Idempotent APIs

Make retried requests safe with idempotency keys, store the result per key, and return the original response on a repeat.

~2 min readread →
esc