A URL shortener maps a long URL to a short code and redirects the code back to the original. It is a classic design because it is read-heavy, needs a collision-free code space, and raises questions about caching, analytics and expiry.
Before you start
You should understand key-value storage, caching and hashing. This article uses a small code sketch.
Step-by-step walkthrough
Step 1: Generate a short code from a unique id
Assign each URL a unique numeric id and encode it in base 62, which yields short codes without collisions, because distinct ids map to distinct codes. Hashing the URL is an alternative but needs collision handling. Prefer the id approach for its simplicity.
Step 2: Serve reads from cache and a fast store
Reads dominate, so put a cache in front of the mapping and keep the store a simple key-value lookup by code. The redirect is a 301 for permanent links or 302 if you need to count each click; 301 lets browsers cache and reduces load but hides repeat visits.
Step 3: Plan for analytics, expiry and abuse
Click counting is high-volume and should be sent to a queue or analytics store, not written on the redirect’s critical path. Support optional expiry with a TTL, and scan submitted URLs to block malicious targets, since a shortener can otherwise launder a bad link.
Worked scenario
Base 62 turns a numeric id into a compact code.
const ALPHABET = '0123456789abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ';
function encode(n) {
let out = '';
do {
out = ALPHABET[n % 62] + out;
n = Math.floor(n / 62);
} while (n > 0);
return out;
}
console.log(encode(125)); // "21"Walk through the example
Each unique id produces one unique code, so there is no collision to resolve. Six characters cover more than 56 billion codes, which is ample. The redirect handler looks up the code, returns a redirect status, and asynchronously records the click so the read path stays fast.
Common mistake
Hashing and accepting collisions without a resolution strategy, or writing analytics synchronously on the redirect path so a slow analytics store slows every click. Another is using 301 while expecting accurate click counts.
Verify the behavior
Assert distinct ids produce distinct codes and that decoding returns the original id. Load-test the redirect with the cache disabled and enabled and compare latency. Confirm clicks are recorded asynchronously and that an expired code returns 404.
Interview exercise
Why prefer a generated id over hashing the long URL?
Answer and reasoning
A generated id guarantees uniqueness, so two different URLs never map to the same code and the same URL can even have multiple codes if desired. Hashing is deterministic but collisions require a resolution and retry path, which adds complexity. The id approach also makes codes shorter and predictable in length.
Continue learning
Compare key generation in ID generation and caching in Cache aside. Read the system design primer and try the System design interview questions.