Partitioning divides data or work by a key. Distribution quality and common queries must be considered together.
Before you start
You should understand API requests, storage and basic capacity estimates. Begin with a concrete user action and its correctness requirement. Draw data flow and failure boundaries before selecting infrastructure; a technology name by itself does not explain why a design meets the requirement.
The practical goal is to reason through this situation: A celebrity account can overload one partition despite otherwise even user distribution. Read the walkthrough first, then try the interview exercise before opening its answer. The important part is explaining the decision and its consequences, rather than remembering a definition alone.
Step-by-step walkthrough
Step 1: Define partition key
The key controls where data and work concentrate.
Step 2: Inspect skew
Average distribution can hide a single extremely hot entity.
Step 3: Preserve required semantics
Splitting a hot key can complicate ordering and aggregation.
Worked scenario
A celebrity account can overload one partition despite otherwise even user distribution.
A celebrity account receives far more writes than ordinary accounts. Adding partitions does not automatically divide that one account’s key. Subpartitioning or caching may reduce pressure, but a globally ordered account activity stream then needs a deliberate merge and consistency policy.
Common mistake
More partitions do not automatically split one hot key.
Verify the behavior
Evaluate both average tenants and the largest hot entity under proposed changes.
Interview exercise
Mitigate a hot entity.
Answer and reasoning
Use workload-appropriate sharding or caching while preserving the entity’s required ordering and consistency rules.
Continue learning
Compare the scenario with the System Design interview questions and test your understanding with the System Design MCQs. For terminology and implementation details, consult the reference material.