System Design · cheat sheet

System Design

Frontend system design: the interview framework, rendering strategies, Core Web Vitals, caching, real-time data, pagination, state, APIs and five mini-designs.

A frontend-leaning system design round on one page: a framework to structure your answer, the trade-off tables interviewers expect, and outlines for the classic prompts.

The interview framework

  1. Requirements (about 5 min): functional (what users do), non-functional (devices, network, scale, SEO, offline, a11y, i18n, real-time?). Write them down and confirm scope.
  2. Architecture: boxes for client, CDN, BFF/API and third parties; the component tree; the rendering strategy and why.
  3. Data model & API: entities and their fields, endpoint or query shapes, pagination, error format, real-time channel.
  4. State & caching: what is local, global, server or URL state; cache layers and invalidation.
  5. Performance: Core Web Vitals targets, bundle strategy, images, lists, network.
  6. Accessibility: semantics, keyboard, focus, screen reader announcements, contrast.
  7. Edge cases: slow or offline network, errors and retries, races, empty and loading states, huge data, security, observability.
  • State trade-offs aloud (“SSR costs server time but helps SEO”) and check in before going deep.

Rendering strategies

Strategy HTML is built Strengths Costs Fits
CSR in the browser, by JS static hosting, app-like UX blank until JS runs; weaker SEO dashboards behind login
SSR on the server, per request fast first paint, SEO, fresh data server cost, slower TTFB, hydration cost personalized or fast-changing pages
SSG at build time fastest, CDN-cacheable stale until rebuilt; long builds docs, blogs, marketing
ISR statically, regenerated after a window or on demand SSG speed, fresher framework-specific; stale window large catalogs
Streaming SSR on the server, flushed in chunks slow sections don’t block the shell headers fixed after first flush pages with slow data
Islands static HTML; only interactive widgets hydrate very little JS sharing state across islands content sites
RSC server components render only on the server no client JS for them; fetch next to the data new mental model, framework-bound React apps
  • Hydration attaches handlers and state to server-rendered HTML; until then the page looks ready but may not respond. Mix strategies per route.

Core Web Vitals

Metric Measures Good Poor
LCP, Largest Contentful Paint loading: render time of the largest image or text block ≤ 2.5 s > 4 s
INP, Interaction to Next Paint responsiveness: latency of clicks, taps and key presses over the whole visit ≤ 200 ms > 500 ms
CLS, Cumulative Layout Shift visual stability: the largest burst of unexpected layout shifts ≤ 0.1 > 0.25
  • Assessed at the 75th percentile of page loads, mobile and desktop separately. INP replaced FID as a Core Web Vital in March 2024.
  • Diagnostics, not Core Web Vitals: TTFB (good ≤ 0.8 s), FCP (good ≤ 1.8 s), and TBT, the lab proxy for INP.
  • Improve LCP: split it into TTFB, resource load delay, load duration and render delay. CDN and caching; LCP image in the HTML with fetchpriority="high" or a preload; never lazy-load it; compress it; don’t wait for client JS.
  • Improve INP: break long tasks (> 50 ms) and yield to the main thread; ship less JS; paint feedback first, defer the rest; move heavy work to a Web Worker.
  • Improve CLS: dimensions or aspect-ratio on media; reserve space for ads and embeds; don’t insert content above the reader; animate transform. Shifts within 500 ms of a tap or key press don’t count.

Performance checklist

  • Network: CDN, HTTP/2 or HTTP/3, Brotli/gzip, preconnect to critical origins, long-lived caching of hashed assets, fewer redirects.
  • JavaScript: split by route with dynamic import(), tree-shake, audit dependencies, load third-party scripts async/defer or later, enforce a bundle budget in CI.
  • Rendering: server-render or pre-render first paint, inline critical CSS, font-display: swap, virtualize long lists, content-visibility: auto for long pages.
  • Images: AVIF/WebP with srcset/sizes, lazy-load below the fold, explicit dimensions, an image CDN for resizing.
  • Data: fetch in parallel (no waterfalls), prefetch on hover or when a link nears the viewport, debounce input (debounce vs throttle), paginate, cache server state.
  • Measure: lab tools (Lighthouse, DevTools) for debugging; field data (RUM, the Chrome UX Report) for truth.

Caching layers & headers

Layer Controlled by Good for Invalidate by
Browser HTTP cache Cache-Control, ETag, Last-Modified static assets, GET responses hashed file names; revalidation
CDN / edge s-maxage, CDN rules assets, SSG/ISR pages, public API responses purge API, versioned URLs
Service worker Cache Storage + fetch handler offline, app shell, custom strategies new cache name; delete old ones on activate
In-memory client cache TanStack Query, SWR, Apollo, RTK Query server state, dedup, instant back navigation invalidate or refetch after mutations
Browser storage localStorage, IndexedDB drafts, preferences, offline data schema version keys
Header Meaning
max-age=31536000, immutable fresh for a year, never revalidated: for hashed assets like app.3f9a1c.js
no-cache may be stored, but must be revalidated before every use: for HTML
no-store never stored anywhere: sensitive data
private / public browser only (personalized responses) / shared caches allowed
s-maxage=600 freshness for shared caches (CDN) only; overrides max-age there
stale-while-revalidate=60 serve stale while refreshing in the background
ETag + If-None-Match validator: unchanged means 304 Not Modified with no body
  • Deploy pattern: HTML no-cache, assets content-hashed and immutable. Keep old assets online for a while so open tabs don’t 404.

Real-time data & pagination

Pattern How Direction Pros Cons
Short polling request every N seconds client pulls trivial, cacheable wasted requests; up to N s stale
Long polling server holds the request until data or a timeout, client re-asks pull, near real-time works through any proxy a held request per client, reconnect overhead
SSE EventSource reading a text/event-stream response server → client auto-reconnect with Last-Event-ID, plain HTTP one-way, text only; over HTTP/1.1 limited to 6 connections per browser and domain
WebSocket HTTP upgrade to a full-duplex connection both ways low latency, binary frames stateful servers, sticky sessions or pub/sub to scale, manual reconnect and heartbeats
  • Pick the simplest that fits: polling for loose freshness, SSE for feeds, notifications and streamed LLM output, WebSockets for chat, games and collaborative editing.
Offset (?page=3&limit=20) Cursor (?after=c_91&limit=20)
Jump to page N, total count yes no (next/previous only)
Deep pages slow: the database skips OFFSET rows fast with an index on the sort key
Inserts while paging duplicates or skipped items stable
Fits admin tables, numbered results feeds, infinite scroll, real-time lists
  • A cursor encodes the last item’s sort key plus a unique tiebreaker (created_at, id); keep it opaque to clients.

State & optimistic updates

Layer Examples Typical home
Local UI input text, open menu component state, signals
Global client theme, session user, feature flags Context, Redux Toolkit, Zustand, NgRx
Server (cache) lists, entities TanStack Query, SWR, Apollo
URL filters, tab, page, search query router params: shareable, survives reload
Persistent drafts, preferences localStorage, IndexedDB
  • Keep state as local as possible, derive instead of copying, and treat server data as a cache with staleness, not as global state. Normalize entities by id so one update shows everywhere.
  • Optimistic update: snapshot, apply the change at once, send the request, roll back and tell the user on failure, then reconcile with the server’s answer. Use it for likely, reversible actions (like, rename, reorder), not payments.
async function toggleLike(postId) {
  const prev = cache.get(postId);
  cache.set(postId, { ...prev, liked: !prev.liked, likes: prev.likes + (prev.liked ? -1 : 1) });
  try {
    const saved = await api.post(`/posts/${postId}/like`, { liked: !prev.liked });
    cache.set(postId, saved);                 // the server's answer wins
  } catch {
    cache.set(postId, prev);                  // roll back, then tell the user
    toast('Could not save your like');
  }
}
JavaScript

API styles

REST GraphQL BFF
Shape resources + HTTP verbs one endpoint, typed schema, client picks fields a backend per frontend that aggregates services
Over/under-fetching common: extra fields or extra round trips avoided tailored to one UI
HTTP caching natural: GET URLs, ETags, CDN harder: usually POST; normalized client caches, persisted queries as you design it
Errors status codes often 200 with an errors array as you design it
Pitfalls chatty clients, versioning N+1 resolvers (batch with DataLoader), query cost limits one more service to own
  • A BFF can sit in front of REST or GraphQL services; it’s where auth tokens, aggregation and response shaping for one client live.

Security

Threat What happens Defenses
XSS attacker script runs in your origin (stored, reflected, DOM-based) framework auto-escaping; no innerHTML/dangerouslySetInnerHTML with untrusted data; sanitize HTML (DOMPurify); strict CSP; Trusted Types
CSRF another site makes the browser send a request with your cookies SameSite cookies, CSRF tokens, check Origin, no state changes on GET
Clickjacking your page framed invisibly under a decoy frame-ancestors 'none' (or 'self'); legacy X-Frame-Options: DENY
Token theft XSS reads tokens from localStorage HttpOnly cookie for the session or refresh token; short-lived access token in memory
Man-in-the-middle traffic read or altered HTTPS everywhere, Strict-Transport-Security, Secure cookies
  • Cookie flags: HttpOnly (no JS access), Secure (HTTPS only), SameSite=Strict (same-site only), Lax (also top-level GET navigations), None (cross-site; requires Secure). Max-Age beats Expires; neither means a session cookie. __Host- requires Secure, Path=/ and no Domain.
  • Strict CSP: script-src 'nonce-{random}' 'strict-dynamic'; object-src 'none'; base-uri 'none'. Roll it out with Content-Security-Policy-Report-Only and report-to. frame-ancestors and report-only don’t work in a <meta> tag.
  • Also: X-Content-Type-Options: nosniff, integrity (SRI) on third-party scripts, CORS allowlists (CORS controls who can read responses, not who can send requests).

Accessibility, i18n & observability

  • A11y at scale: an accessible design system, WCAG 2.2 AA as the bar, lint plus axe in CI (automation catches only some issues), manual keyboard and screen reader passes.
  • In SPAs, move focus to the new page’s heading (or announce it) after a route change, and announce async results with a live region.
  • i18n: externalize strings with ICU messages for plurals and gender, never concatenate fragments; format with Intl.NumberFormat, DateTimeFormat, PluralRules, RelativeTimeFormat; sort with Intl.Collator.
  • RTL via dir="rtl" and logical properties (margin-inline-start); room for longer text; store UTC, render in the user’s zone; lazy-load locale bundles.
  • RUM: collect LCP, INP and CLS from real users (a PerformanceObserver), segmented by route, device and country; alert on p75 regressions.
  • Errors: window error and unhandledrejection listeners, framework error boundaries, source maps uploaded privately to the error tracker, releases tagged.
  • Logs: structured, sampled, with a trace id passed to the backend (traceparent); send on visibilitychange with navigator.sendBeacon or fetch(url, { keepalive: true }); scrub PII.

Micro-frontends & offline

Micro-frontends: pros Cons
independent deploys and team autonomy duplicated dependencies, larger downloads
incremental migration from a legacy app inconsistent UX without a shared design system
smaller codebases, isolated failures cross-app state, routing and versioning complexity
  • Composition options: route-level split (simplest), runtime Module Federation, Web Components, server or edge composition, iframes (strong isolation, poor UX). Split by business domain and talk through the URL or events, not a shared store. Worth it for many teams, rarely for one.
  • Service worker: a script between the page and the network, HTTPS only (localhost excepted). Lifecycle: install (precache) → waiting → activate (clean old caches) → fetch. A new version waits for old tabs to close unless it calls skipWaiting().
  • Strategies: cache-first for hashed assets, network-first for HTML and fresh data, stale-while-revalidate for avatars and non-critical API calls.
  • A web app manifest (name, icons, start_url, display) makes the app installable. Store offline data in IndexedDB, queue writes and replay them when back online; resolve conflicts (last write wins, versions, or CRDTs for collaborative editing).

Back-of-envelope numbers

Number Value Why it matters
Frame budget at 60 Hz ~16.7 ms smooth scrolling and animation
Long task > 50 ms blocks input; hurts INP
Seconds per day 86,400 (~10⁵) 1M users × 10 requests/day ≈ 115 requests/s on average
Peak vs average traffic a multiple: assume 2-3× and say so size for peak
TCP + TLS 1.3 setup ~2 round trips before the request why preconnect and CDNs help
Cross-continent round trip ~100 ms or more put static content at the edge
First TCP flight ~14 KB (10 packets) keep critical HTML and CSS small
Cookie size ~4 KB each cookies ride on every request
localStorage ~5 MB per origin, synchronous not for large data; use IndexedDB
  • These are rules of thumb: state your assumptions and round aggressively.

Mini-designs

Autocomplete (full walkthrough)

  • Debounce input (~250 ms), minimum length, abort the previous request and ignore stale responses; LRU cache with a TTL per query.
  • ARIA combobox: focus stays in the input, arrows move aria-activedescendant, Enter selects, Esc closes. Plan loading, empty and error states.
let controller;
async function search(q) {
  controller?.abort();                        // cancel the previous request
  controller = new AbortController();
  try {
    const res = await fetch(`/api/search?q=${encodeURIComponent(q)}`, { signal: controller.signal });
    render(await res.json());
  } catch (err) {
    if (err.name !== 'AbortError') showError(err);
  }
}
JavaScript

News feed

  • Cursor pagination, normalized post cache, virtualized list, optimistic likes, reserved media sizes (CLS).
  • New posts: show a “5 new posts” pill (SSE or polling) instead of shifting the list; prefetch the next page near the bottom.

Chat

  • WebSocket with heartbeats and reconnect plus exponential backoff; after reconnecting, resync from the last seen message id.
  • Client-generated ids for dedup and idempotent retries; states sending → sent → delivered → read; order by server sequence.
  • Load history upward with a cursor, keep the scroll anchored at the bottom, queue unsent messages in IndexedDB, announce new messages politely.

Infinite scroll

  • IntersectionObserver on a sentinel with a rootMargin to prefetch; cursor pagination; virtualization to cap DOM nodes.
  • Restore scroll position on back navigation, dedupe, retry inline; a “Load more” button keeps the footer reachable.

Image upload

  • <input type="file" accept="image/*" multiple> plus a keyboard-reachable drop zone; validate type and size; preview with URL.createObjectURL (then revoke it).
  • Optionally resize or compress on the client in a worker; upload straight to object storage with a pre-signed URL; chunked, resumable uploads for big files.
  • Progress through XHR upload.onprogress (fetch has no upload progress events), cancel with AbortController. Server side: scan, strip EXIF (GPS), generate sizes, serve via CDN.

Quick answers

  • SSR vs CSR? SSR sends ready HTML (fast first paint, SEO) at a server cost; CSR builds pages in the browser (cheap hosting, slow first load).
  • What are the Core Web Vitals? LCP ≤ 2.5 s, INP ≤ 200 ms, CLS ≤ 0.1, judged at the 75th percentile.
  • no-cache vs no-store? no-cache stores but revalidates every time; no-store never stores.
  • How do you cache-bust? Content-hashed file names with immutable, and no-cache HTML that points at them.
  • SSE vs WebSocket? SSE is one-way over plain HTTP with built-in reconnect; WebSocket is two-way and needs more infrastructure.
  • Offset vs cursor pagination? Offset allows jumping to page N; cursor stays fast and stable on large, changing lists.
  • Where do auth tokens go? An HttpOnly, Secure, SameSite cookie (plus CSRF protection); not localStorage.
  • XSS vs CSRF? XSS runs the attacker’s code in your page; CSRF makes the victim’s browser send a request it didn’t intend.
  • How do you stop clickjacking? Content-Security-Policy: frame-ancestors 'none' (or X-Frame-Options: DENY).
  • When micro-frontends? When many teams need to deploy independently; otherwise a modular monolith is cheaper.
  • How do you handle search races? Abort earlier requests (or tag them) and render only the latest response.

Gotchas & traps

  • no-cache does not mean “don’t cache”; no-store does.
  • A personalized response without private can be cached by a CDN and shown to other users.
  • SameSite=None without Secure is rejected, and SameSite alone doesn’t stop attacks from other subdomains of the same site.
  • CORS is not CSRF protection: the request is still sent, the page just can’t read the response.
  • Lazy-loading or client-rendering the LCP element wrecks LCP. A Lighthouse page load can’t measure INP: use field data (TBT is only a lab proxy).
  • Offset pagination on a live feed duplicates items as new ones arrive.
  • Optimistic updates without rollback, or with out-of-order responses, leave the UI lying.
  • WebSockets behind a load balancer need sticky sessions or a pub/sub backplane, plus heartbeats to detect dead connections.
  • A buggy service worker can pin users to an old build: serve sw.js with no-cache, version caches, and plan an update prompt.
  • navigator.onLine === true only means there’s a network interface, not that your API is reachable.
esc