Microservices · cheat sheet

Microservices

Microservices interview prep: DDD decomposition, sync vs async calls, gateways and discovery, sagas and the outbox, resilience, observability and security.

The patterns, trade-offs and vocabulary interviewers expect for microservices, from splitting a domain to running it in production.

Architecture styles

Monolith Modular monolith Microservices
Deploy unit one one many, independent
Boundaries by convention, often eroded enforced modules, one codebase network plus separate pipelines
Data one database one database, tables owned per module database per service
Calls in-process in-process, through module APIs network: latency, partial failure
Scaling whole app whole app per service
Consistency ACID transactions ACID transactions sagas, eventual consistency
Ops cost low low high: CI/CD, observability, platform
Fits small team, new product most teams by default many teams needing independent deploys
  • A microservice is independently deployable, owns its data, and models one business capability. Size follows team ownership, not lines of code.
  • Buy the benefits (autonomy, independent scaling, fault isolation, tech choice) with distributed-systems costs. Start modular, extract when a boundary has proven itself.

Decomposition & DDD

  • Split by business capability (ordering, billing, shipping) or by DDD subdomain: core (your edge), supporting, generic (buy it).
  • Bounded context: a boundary inside which one model and its ubiquitous language stay consistent. “Customer” in Sales is not “Customer” in Support. A service maps to one bounded context, or part of one.
  • Aggregate: a cluster of entities with one root that forms a consistency boundary. One transaction changes one aggregate; other aggregates are referenced by ID.
  • Context mapping: anti-corruption layer (translate a legacy or external model), shared kernel, customer/supplier, conformist, open host service.
  • Good boundaries: high cohesion, few cross-service calls per use case, one owning team; things that change together live together.
  • Conway’s law: systems mirror the communication structure of the org, so shape teams around the architecture you want (the inverse Conway maneuver).
  • Strangler fig: put a facade in front of the monolith and route one slice at a time to new services until the old code can go.

Communication

Style Transport Strengths Costs
REST HTTP + JSON universal, cacheable, easy to debug chatty; contracts need OpenAPI
gRPC HTTP/2 + Protobuf fast, typed contracts, codegen, streaming browsers need gRPC-Web; binary payloads
GraphQL HTTP + JSON clients pick fields; good for BFFs caching, N+1 resolvers
Messaging Kafka, RabbitMQ, SQS temporal decoupling, buffering, fan-out eventual consistency, duplicates, harder tracing
  • Sync calls create temporal coupling: availabilities multiply, so five dependencies at 99.9% give about 99.5%.
  • Commands (“reserve stock”) have one receiver; events (“OrderPlaced”, past tense) have any number of subscribers.
  • A queue delivers each message to one of the competing consumers; a topic gives every subscriber its own copy.
  • Kafka: a partitioned log. Order is guaranteed only within a partition (choose the key, e.g. orderId); each partition goes to one consumer per group; retention allows replay.
  • Evolve contracts additively; readers ignore unknown fields (tolerant reader). Breaking changes get a new version (/v2, media type, new topic) and both run until consumers migrate.

Gateway, BFF & discovery

  • API gateway: the single entry point, handling routing, TLS termination, token validation, rate limiting, response aggregation and caching. Keep business logic out of it.
  • Backend for frontend (BFF): one gateway per client type (web, mobile), owned by that frontend team and shaped for its screens.
  • Gateway vs service mesh: north-south (client to cluster) vs east-west (service to service). A mesh (Istio, Linkerd) puts sidecar proxies next to each service for mTLS, retries, timeouts, traffic splitting and telemetry without code changes.
  • A service registry (Eureka, Consul, etcd) tracks healthy instances. Instances register themselves with heartbeats, or the platform registers them (Kubernetes endpoints).
Client-side discovery Server-side discovery
Who picks the instance the client queries the registry and load-balances a router or load balancer
Example Eureka + Spring Cloud LoadBalancer Kubernetes Service, cloud load balancer
Trade-off no extra hop, but logic in every client simple clients, but an extra hop to keep highly available
  • In Kubernetes, call http://orders (a Service DNS name); readiness probes decide which pods get traffic.

Data ownership & sagas

  • Database per service: private tables, schema or server; others go through the API or events. A shared database couples schemas and release schedules.
  • Cross-service queries: API composition (call and join in memory) or a CQRS read model fed by events.
  • Avoid 2PC/XA: it blocks while the coordinator is down, hurts availability, and many brokers and NoSQL stores don’t support it.
  • Saga: a sequence of local transactions. When a step fails, compensating transactions undo earlier steps semantically (refund, release stock), not by rollback.
Choreography Orchestration
Coordination services react to each other’s events an orchestrator sends commands and tracks state
Coupling services know events, not each other participants know only their commands
Visibility flow spread across services, hard to follow flow in one place, easy to monitor
Risk cyclic dependencies, hard to change orchestrator grows into a god service
Fits a few simple steps many steps, branching, timeouts
  • Orchestrators are often built on a workflow engine (Temporal, Camunda, AWS Step Functions).
  • Sagas are ACD without the I: other transactions can see intermediate state. Countermeasures: a semantic lock (PENDING status), commutative updates, rereading a value before overwriting it, and ordering steps so risky ones come last.
  • Step order: compensatable steps, then the pivot (the go/no-go point), then retriable steps that must eventually succeed.

Outbox, CDC, CQRS & ES

  • Dual-write problem: save() then publish() can half-fail, losing events or announcing changes that never committed.
BEGIN;
INSERT INTO orders (id, status) VALUES (42, 'PENDING');
INSERT INTO outbox (id, aggregate_id, type, payload)
VALUES (gen_random_uuid(), 42, 'OrderCreated', '{"orderId": 42}');
COMMIT;
-- a relay publishes new outbox rows to the broker, then marks or deletes them
sql
  • Transactional outbox: the event row commits atomically with the business change. The relay is a polling publisher or transaction-log tailing. Delivery is at-least-once, so consumers must be idempotent.
  • CDC (change data capture, e.g. Debezium) streams row changes from the database log (Postgres WAL, MySQL binlog) to Kafka without touching app code. Capture the outbox table, not internal tables, to avoid leaking your schema.
  • CQRS: separate the write model (commands, normalized) from read models (queries, denormalized, possibly another store such as Elasticsearch), updated from events. Reads are eventually consistent; use it when read and write needs differ sharply.
  • Event sourcing: store every state change as an append-only event and rebuild state by replaying (snapshots speed it up). It gives a full audit trail, time travel and events for free; it costs event schema evolution, projections for queries, and harder deletes.
  • Event sourcing is a persistence choice; event-driven architecture is a communication style. You can have either without the other.

Resilience patterns

Pattern Protects against Key idea
Timeout hung calls holding threads connect and read timeouts on every remote call, shorter than the caller’s
Retry transient faults idempotent calls only; few attempts; exponential backoff with jitter
Circuit breaker hammering a failing dependency fail fast while it’s down, probe to detect recovery
Bulkhead one slow dependency using all threads separate pools or semaphores per dependency
Fallback a feature failing completely cached, default or degraded response
Rate limiting overload, abuse token bucket or sliding window; reply 429 with Retry-After
Load shedding queues growing without bound bounded queues; reject early when saturated
  • Circuit breaker states: CLOSED (calls flow, failures counted), OPEN (after the failure rate crosses a threshold, calls fail immediately), HALF_OPEN (after a wait, a few trial calls go through; success closes, failure reopens).
long backoffMillis(int attempt) {                     // exponential backoff, full jitter
  long cap = 10_000, base = 100;
  long ceiling = Math.min(cap, base * (1L << attempt));
  return ThreadLocalRandom.current().nextLong(ceiling + 1);   // random in [0, ceiling]
}
java
  • Jitter spreads retries so clients don’t stampede in sync. Retry at one layer: 3 retries at each of 3 layers means up to 4 × 4 × 4 = 64 calls to the bottom service.
  • Propagate a deadline (remaining time budget) downstream instead of fixed per-hop timeouts.
  • Liveness probe: “restart me if I’m stuck”. Readiness probe: “don’t send me traffic yet”. Never check dependencies in liveness, or one database outage restarts everything.
  • Token bucket allows bursts up to the bucket size; a fixed window allows double bursts at window edges; a sliding window fixes that.

Idempotency & delivery

Guarantee How Risk
At-most-once no retries; ack or commit before processing lost messages
At-least-once retry until acked; commit after processing duplicates (the usual default)
Effectively exactly-once at-least-once plus deduplication or idempotent handlers extra state; Kafka transactions cover only Kafka-to-Kafka
  • HTTP: GET, HEAD, PUT, DELETE, OPTIONS are idempotent; POST and PATCH are not by definition.
  • Retry-safe POST: the client sends an Idempotency-Key header; the server stores key and response under a unique constraint and replays the stored response for repeats.
  • Idempotent consumer: record the message ID in the same transaction as the side effect.
INSERT INTO processed_messages (message_id) VALUES (:id)
ON CONFLICT (message_id) DO NOTHING;   -- 0 rows inserted: duplicate, skip the handler
sql
  • Prefer naturally idempotent writes (SET status = 'PAID') over increments (balance = balance - 10), or guard with a version: WHERE version = 7.
  • Ordering holds only per partition or key: carry a version or sequence number and drop stale events.
  • Poison messages: bounded retries, then a dead-letter queue, an alert and a replay tool.

Observability

  • Logs: structured JSON to stdout, with trace ID, service and level; no secrets or PII; centralized (ELK, Loki).
  • Metrics: RED for services (rate, errors, duration), USE for resources (utilization, saturation, errors), or the four golden signals (latency, traffic, errors, saturation). Alert on p95/p99 latency, not averages.
  • Traces: a trace is a tree of spans sharing one trace ID; each span records its parent, timing and attributes. Backends: Jaeger, Zipkin, Tempo.
  • Correlation ID: created at the edge if missing, forwarded in every HTTP and message header, and put in the logging MDC so one ID finds every log line.
  • W3C Trace Context header: traceparent: 00-<32-hex trace-id>-<16-hex parent-id>-01.
  • OpenTelemetry: vendor-neutral APIs, SDKs and auto-instrumentation (such as the Java agent) for traces, metrics and logs, sent over OTLP to the Collector (receivers, processors, exporters) and on to any backend.
  • Sampling: head-based decides at the first span (cheap); tail-based decides after the trace finishes, so it can keep every error and slow trace.
  • SLI is what you measure, SLO the target (“99.9% of requests under 300 ms over 30 days”), SLA the contract with penalties. Error budget = 1 − SLO.
  • Debug flow: dashboard shows the spike, a trace shows the slow span, logs filtered by trace ID show why.

Security

  • Authenticate at the edge, but every service still validates tokens (zero trust): signature via JWKS, iss, aud, exp, scopes.
  • OAuth 2 handles delegated authorization; OpenID Connect adds authentication with an ID token.
Flow Use for
Authorization code + PKCE users in web, single-page and mobile apps
Client credentials service-to-service calls with no user
Refresh token a new access token without logging in again
Token exchange swap a token for one scoped to a downstream service
Implicit, password legacy; current best practice drops them
  • JWT: header.payload.signature, Base64URL-encoded, signed but not encrypted. Prefer asymmetric keys (RS256, ES256) so services verify with the public key. Keep expiry short; early revocation needs a denylist, or opaque tokens with introspection.
  • Propagate identity by forwarding or exchanging the token; never trust a raw X-User-Id header from outside the gateway.
  • mTLS: both sides present certificates, giving encryption plus service identity. A mesh automates issuing and rotation.
  • Secrets live in Vault or a cloud secret manager. Kubernetes Secrets are only base64-encoded unless encryption at rest is on. Never bake secrets into images or git.
  • Also: a least-privilege database user per service, network policies, input validation and rate limits.

Deployment strategies

Strategy How Pros Cons
Recreate stop old, start new simple downtime
Rolling replace instances in batches (Kubernetes default) no extra capacity two versions live at once
Blue-green full parallel environment, switch traffic instant switch and rollback double capacity; DB must suit both
Canary small % of traffic first, watch metrics, ramp up small blast radius needs traffic splitting and good metrics
Shadow mirror live traffic, discard responses no user impact stub side effects; extra cost
Feature flags ship dark, release by toggle deploy separate from release; kill switch flag debt
  • Schema changes use expand/contract: add the new column, write both, backfill, switch reads, then drop the old one. Every release must work alongside the previous version.
  • One service, one pipeline, one immutable image promoted through environments with config injected at runtime.
  • A/B testing measures a business metric on user groups; canaries measure release health.

Testing

Test Scope Tools
Unit a class or function JUnit, Mockito
Integration service plus real DB or broker Testcontainers
Component one service, dependencies stubbed WireMock, in-memory broker
Contract the API agreement between consumer and provider Pact, Spring Cloud Contract
End-to-end the whole system a few critical journeys only (slow, flaky)
  • Consumer-driven contracts: consumer tests record the requests and responses they rely on as a contract; the provider’s build replays them against the real provider. A broker shares contracts and can block a deploy that would break a consumer.
  • Spring Cloud Contract starts from provider-side contracts and generates both provider tests and consumer stubs.
  • Test in production too: canaries, synthetic checks, chaos experiments (kill instances, inject latency).

CAP & PACELC

  • CAP: during a network partition, a replicated store must choose consistency (every read sees the latest write, i.e. linearizability) or availability (every live node answers). Partitions happen anyway, so the real choice is C vs A while partitioned.
  • CP systems refuse or fail some requests during a partition (ZooKeeper, etcd, HBase); AP systems answer with possibly stale data (Cassandra, CouchDB). Many stores let you tune this per query.
  • PACELC: if Partitioned, choose A or C; Else choose Latency or Consistency. Dynamo-style stores (Cassandra, Riak) are PA/EL; HBase and Spanner-style databases are PC/EC.
  • CAP’s C is not ACID’s C (which means invariants hold).
  • Quorums: with N replicas, reads of R and writes of W where R + W > N overlap, so reads see the latest acknowledged write.
  • BASE (basically available, soft state, eventually consistent) is the usual trade for availability.

The twelve factors

# Factor In one line
I Codebase one repo per app, deployed to many environments
II Dependencies declare every dependency explicitly; assume nothing on the host
III Config settings that vary per deploy come from the environment, not code
IV Backing services databases, queues and caches are attached resources swapped by URL
V Build, release, run build once, combine with config into an immutable release, run it
VI Processes stateless, share-nothing processes; state lives in backing services
VII Port binding self-contained app serves HTTP on its own port (embedded server)
VIII Concurrency scale out by running more processes
IX Disposability fast startup, graceful shutdown on SIGTERM
X Dev/prod parity keep dev, staging and prod as alike as possible
XI Logs write an event stream to stdout; the platform collects it
XII Admin processes run one-off tasks (migrations) as separate processes of the same release

Quick answers

  • When not to use microservices? A small team, an unclear domain or an early product: start with a modular monolith.
  • How do you split a monolith? Strangler fig along bounded contexts, starting where seams are clear and change is frequent.
  • Distributed transactions? Sagas with compensations plus the outbox; avoid 2PC.
  • Choreography or orchestration? Choreography for short, simple flows; orchestration once flows branch or grow.
  • How do services find each other? A registry, or platform DNS such as Kubernetes Services.
  • Gateway vs service mesh? Edge concerns for client traffic vs sidecars for service-to-service traffic.
  • Preventing cascading failures? Timeouts, retries with backoff and jitter, circuit breakers, bulkheads, fallbacks.
  • Duplicate messages? Assume at-least-once and make consumers idempotent with dedup keys.
  • Tracing a request across services? OpenTelemetry with W3C traceparent propagation, and the trace ID in every log line.
  • Securing service-to-service calls? mTLS for transport identity plus client-credentials tokens for authorization.
  • Can two services share a database? Avoid it: it couples schemas and releases. Share data through APIs or events.
  • How do you keep data consistent? Eventual consistency via events, with explicit pending states in the API and UI.

Anti-patterns

  • Distributed monolith: services that must deploy in lockstep, often through shared domain libraries or a shared database.
  • Nanoservices and entity services (CRUD per table): chatty calls whose overhead outweighs any benefit.
  • Long synchronous chains: latency adds up and availability multiplies down. Go async or merge services.
  • Dual writes without an outbox: lost or phantom events.
  • Retries without timeouts, backoff or idempotency: retry storms and duplicate side effects.
  • Smart gateway or ESB full of business logic: a new monolith in the middle.
  • Big-bang rewrite instead of strangling the old system piece by piece.
  • No correlation IDs or tracing: debugging means grepping twenty log streams.
  • Breaking API changes with no versioning or contract tests.

Interview tip

For any design question, name the trade-off out loud: “this buys independent deploys and costs eventual consistency, so here is how I handle it”.

esc