Microservices interview prep: DDD decomposition, sync vs async calls, gateways and discovery, sagas and the outbox, resilience, observability and security.
AMCompiled by Aditya Mishra · Technical Lead, BNP Paribas
The patterns, trade-offs and vocabulary interviewers expect for microservices, from splitting a domain to running it in production.
Architecture styles
Monolith
Modular monolith
Microservices
Deploy unit
one
one
many, independent
Boundaries
by convention, often eroded
enforced modules, one codebase
network plus separate pipelines
Data
one database
one database, tables owned per module
database per service
Calls
in-process
in-process, through module APIs
network: latency, partial failure
Scaling
whole app
whole app
per service
Consistency
ACID transactions
ACID transactions
sagas, eventual consistency
Ops cost
low
low
high: CI/CD, observability, platform
Fits
small team, new product
most teams by default
many teams needing independent deploys
A microservice is independently deployable, owns its data, and models one business capability. Size follows team ownership, not lines of code.
Buy the benefits (autonomy, independent scaling, fault isolation, tech choice) with distributed-systems costs. Start modular, extract when a boundary has proven itself.
Decomposition & DDD
Split by business capability (ordering, billing, shipping) or by DDD subdomain: core (your edge), supporting, generic (buy it).
Bounded context: a boundary inside which one model and its ubiquitous language stay consistent. “Customer” in Sales is not “Customer” in Support. A service maps to one bounded context, or part of one.
Aggregate: a cluster of entities with one root that forms a consistency boundary. One transaction changes one aggregate; other aggregates are referenced by ID.
Context mapping: anti-corruption layer (translate a legacy or external model), shared kernel, customer/supplier, conformist, open host service.
Good boundaries: high cohesion, few cross-service calls per use case, one owning team; things that change together live together.
Conway’s law: systems mirror the communication structure of the org, so shape teams around the architecture you want (the inverse Conway maneuver).
Strangler fig: put a facade in front of the monolith and route one slice at a time to new services until the old code can go.
Communication
Style
Transport
Strengths
Costs
REST
HTTP + JSON
universal, cacheable, easy to debug
chatty; contracts need OpenAPI
gRPC
HTTP/2 + Protobuf
fast, typed contracts, codegen, streaming
browsers need gRPC-Web; binary payloads
GraphQL
HTTP + JSON
clients pick fields; good for BFFs
caching, N+1 resolvers
Messaging
Kafka, RabbitMQ, SQS
temporal decoupling, buffering, fan-out
eventual consistency, duplicates, harder tracing
Sync calls create temporal coupling: availabilities multiply, so five dependencies at 99.9% give about 99.5%.
Commands (“reserve stock”) have one receiver; events (“OrderPlaced”, past tense) have any number of subscribers.
A queue delivers each message to one of the competing consumers; a topic gives every subscriber its own copy.
Kafka: a partitioned log. Order is guaranteed only within a partition (choose the key, e.g. orderId); each partition goes to one consumer per group; retention allows replay.
Evolve contracts additively; readers ignore unknown fields (tolerant reader). Breaking changes get a new version (/v2, media type, new topic) and both run until consumers migrate.
Gateway, BFF & discovery
API gateway: the single entry point, handling routing, TLS termination, token validation, rate limiting, response aggregation and caching. Keep business logic out of it.
Backend for frontend (BFF): one gateway per client type (web, mobile), owned by that frontend team and shaped for its screens.
Gateway vs service mesh: north-south (client to cluster) vs east-west (service to service). A mesh (Istio, Linkerd) puts sidecar proxies next to each service for mTLS, retries, timeouts, traffic splitting and telemetry without code changes.
A service registry (Eureka, Consul, etcd) tracks healthy instances. Instances register themselves with heartbeats, or the platform registers them (Kubernetes endpoints).
Client-side discovery
Server-side discovery
Who picks the instance
the client queries the registry and load-balances
a router or load balancer
Example
Eureka + Spring Cloud LoadBalancer
Kubernetes Service, cloud load balancer
Trade-off
no extra hop, but logic in every client
simple clients, but an extra hop to keep highly available
In Kubernetes, call http://orders (a Service DNS name); readiness probes decide which pods get traffic.
Data ownership & sagas
Database per service: private tables, schema or server; others go through the API or events. A shared database couples schemas and release schedules.
Cross-service queries: API composition (call and join in memory) or a CQRS read model fed by events.
Avoid 2PC/XA: it blocks while the coordinator is down, hurts availability, and many brokers and NoSQL stores don’t support it.
Saga: a sequence of local transactions. When a step fails, compensating transactions undo earlier steps semantically (refund, release stock), not by rollback.
Choreography
Orchestration
Coordination
services react to each other’s events
an orchestrator sends commands and tracks state
Coupling
services know events, not each other
participants know only their commands
Visibility
flow spread across services, hard to follow
flow in one place, easy to monitor
Risk
cyclic dependencies, hard to change
orchestrator grows into a god service
Fits
a few simple steps
many steps, branching, timeouts
Orchestrators are often built on a workflow engine (Temporal, Camunda, AWS Step Functions).
Sagas are ACD without the I: other transactions can see intermediate state. Countermeasures: a semantic lock (PENDING status), commutative updates, rereading a value before overwriting it, and ordering steps so risky ones come last.
Step order: compensatable steps, then the pivot (the go/no-go point), then retriable steps that must eventually succeed.
Outbox, CDC, CQRS & ES
Dual-write problem: save() then publish() can half-fail, losing events or announcing changes that never committed.
BEGIN;INSERT INTO orders (id, status) VALUES (42, 'PENDING');INSERT INTO outbox (id, aggregate_id, type, payload)VALUES (gen_random_uuid(), 42, 'OrderCreated', '{"orderId": 42}');COMMIT;-- a relay publishes new outbox rows to the broker, then marks or deletes them
sql
Transactional outbox: the event row commits atomically with the business change. The relay is a polling publisher or transaction-log tailing. Delivery is at-least-once, so consumers must be idempotent.
CDC (change data capture, e.g. Debezium) streams row changes from the database log (Postgres WAL, MySQL binlog) to Kafka without touching app code. Capture the outbox table, not internal tables, to avoid leaking your schema.
CQRS: separate the write model (commands, normalized) from read models (queries, denormalized, possibly another store such as Elasticsearch), updated from events. Reads are eventually consistent; use it when read and write needs differ sharply.
Event sourcing: store every state change as an append-only event and rebuild state by replaying (snapshots speed it up). It gives a full audit trail, time travel and events for free; it costs event schema evolution, projections for queries, and harder deletes.
Event sourcing is a persistence choice; event-driven architecture is a communication style. You can have either without the other.
Resilience patterns
Pattern
Protects against
Key idea
Timeout
hung calls holding threads
connect and read timeouts on every remote call, shorter than the caller’s
Retry
transient faults
idempotent calls only; few attempts; exponential backoff with jitter
Circuit breaker
hammering a failing dependency
fail fast while it’s down, probe to detect recovery
Bulkhead
one slow dependency using all threads
separate pools or semaphores per dependency
Fallback
a feature failing completely
cached, default or degraded response
Rate limiting
overload, abuse
token bucket or sliding window; reply 429 with Retry-After
Load shedding
queues growing without bound
bounded queues; reject early when saturated
Circuit breaker states: CLOSED (calls flow, failures counted), OPEN (after the failure rate crosses a threshold, calls fail immediately), HALF_OPEN (after a wait, a few trial calls go through; success closes, failure reopens).
long backoffMillis(int attempt) { // exponential backoff, full jitter long cap = 10_000, base = 100; long ceiling = Math.min(cap, base * (1L << attempt)); return ThreadLocalRandom.current().nextLong(ceiling + 1); // random in [0, ceiling]}
java
Jitter spreads retries so clients don’t stampede in sync. Retry at one layer: 3 retries at each of 3 layers means up to 4 × 4 × 4 = 64 calls to the bottom service.
Propagate a deadline (remaining time budget) downstream instead of fixed per-hop timeouts.
Liveness probe: “restart me if I’m stuck”. Readiness probe: “don’t send me traffic yet”. Never check dependencies in liveness, or one database outage restarts everything.
Token bucket allows bursts up to the bucket size; a fixed window allows double bursts at window edges; a sliding window fixes that.
Idempotency & delivery
Guarantee
How
Risk
At-most-once
no retries; ack or commit before processing
lost messages
At-least-once
retry until acked; commit after processing
duplicates (the usual default)
Effectively exactly-once
at-least-once plus deduplication or idempotent handlers
extra state; Kafka transactions cover only Kafka-to-Kafka
HTTP: GET, HEAD, PUT, DELETE, OPTIONS are idempotent; POST and PATCH are not by definition.
Retry-safe POST: the client sends an Idempotency-Key header; the server stores key and response under a unique constraint and replays the stored response for repeats.
Idempotent consumer: record the message ID in the same transaction as the side effect.
INSERT INTO processed_messages (message_id) VALUES (:id)ON CONFLICT (message_id) DO NOTHING; -- 0 rows inserted: duplicate, skip the handler
sql
Prefer naturally idempotent writes (SET status = 'PAID') over increments (balance = balance - 10), or guard with a version: WHERE version = 7.
Ordering holds only per partition or key: carry a version or sequence number and drop stale events.
Poison messages: bounded retries, then a dead-letter queue, an alert and a replay tool.
Observability
Logs: structured JSON to stdout, with trace ID, service and level; no secrets or PII; centralized (ELK, Loki).
Metrics: RED for services (rate, errors, duration), USE for resources (utilization, saturation, errors), or the four golden signals (latency, traffic, errors, saturation). Alert on p95/p99 latency, not averages.
Traces: a trace is a tree of spans sharing one trace ID; each span records its parent, timing and attributes. Backends: Jaeger, Zipkin, Tempo.
Correlation ID: created at the edge if missing, forwarded in every HTTP and message header, and put in the logging MDC so one ID finds every log line.
OpenTelemetry: vendor-neutral APIs, SDKs and auto-instrumentation (such as the Java agent) for traces, metrics and logs, sent over OTLP to the Collector (receivers, processors, exporters) and on to any backend.
Sampling: head-based decides at the first span (cheap); tail-based decides after the trace finishes, so it can keep every error and slow trace.
SLI is what you measure, SLO the target (“99.9% of requests under 300 ms over 30 days”), SLA the contract with penalties. Error budget = 1 − SLO.
Debug flow: dashboard shows the spike, a trace shows the slow span, logs filtered by trace ID show why.
Security
Authenticate at the edge, but every service still validates tokens (zero trust): signature via JWKS, iss, aud, exp, scopes.
OAuth 2 handles delegated authorization; OpenID Connect adds authentication with an ID token.
Flow
Use for
Authorization code + PKCE
users in web, single-page and mobile apps
Client credentials
service-to-service calls with no user
Refresh token
a new access token without logging in again
Token exchange
swap a token for one scoped to a downstream service
Implicit, password
legacy; current best practice drops them
JWT: header.payload.signature, Base64URL-encoded, signed but not encrypted. Prefer asymmetric keys (RS256, ES256) so services verify with the public key. Keep expiry short; early revocation needs a denylist, or opaque tokens with introspection.
Propagate identity by forwarding or exchanging the token; never trust a raw X-User-Id header from outside the gateway.
mTLS: both sides present certificates, giving encryption plus service identity. A mesh automates issuing and rotation.
Secrets live in Vault or a cloud secret manager. Kubernetes Secrets are only base64-encoded unless encryption at rest is on. Never bake secrets into images or git.
Also: a least-privilege database user per service, network policies, input validation and rate limits.
Deployment strategies
Strategy
How
Pros
Cons
Recreate
stop old, start new
simple
downtime
Rolling
replace instances in batches (Kubernetes default)
no extra capacity
two versions live at once
Blue-green
full parallel environment, switch traffic
instant switch and rollback
double capacity; DB must suit both
Canary
small % of traffic first, watch metrics, ramp up
small blast radius
needs traffic splitting and good metrics
Shadow
mirror live traffic, discard responses
no user impact
stub side effects; extra cost
Feature flags
ship dark, release by toggle
deploy separate from release; kill switch
flag debt
Schema changes use expand/contract: add the new column, write both, backfill, switch reads, then drop the old one. Every release must work alongside the previous version.
One service, one pipeline, one immutable image promoted through environments with config injected at runtime.
A/B testing measures a business metric on user groups; canaries measure release health.
Testing
Test
Scope
Tools
Unit
a class or function
JUnit, Mockito
Integration
service plus real DB or broker
Testcontainers
Component
one service, dependencies stubbed
WireMock, in-memory broker
Contract
the API agreement between consumer and provider
Pact, Spring Cloud Contract
End-to-end
the whole system
a few critical journeys only (slow, flaky)
Consumer-driven contracts: consumer tests record the requests and responses they rely on as a contract; the provider’s build replays them against the real provider. A broker shares contracts and can block a deploy that would break a consumer.
Spring Cloud Contract starts from provider-side contracts and generates both provider tests and consumer stubs.
Test in production too: canaries, synthetic checks, chaos experiments (kill instances, inject latency).
CAP & PACELC
CAP: during a network partition, a replicated store must choose consistency (every read sees the latest write, i.e. linearizability) or availability (every live node answers). Partitions happen anyway, so the real choice is C vs A while partitioned.
CP systems refuse or fail some requests during a partition (ZooKeeper, etcd, HBase); AP systems answer with possibly stale data (Cassandra, CouchDB). Many stores let you tune this per query.
PACELC: if Partitioned, choose A or C; Else choose Latency or Consistency. Dynamo-style stores (Cassandra, Riak) are PA/EL; HBase and Spanner-style databases are PC/EC.
CAP’s C is not ACID’s C (which means invariants hold).
Quorums: with N replicas, reads of R and writes of W where R + W > N overlap, so reads see the latest acknowledged write.
BASE (basically available, soft state, eventually consistent) is the usual trade for availability.
The twelve factors
#
Factor
In one line
I
Codebase
one repo per app, deployed to many environments
II
Dependencies
declare every dependency explicitly; assume nothing on the host
III
Config
settings that vary per deploy come from the environment, not code
IV
Backing services
databases, queues and caches are attached resources swapped by URL
V
Build, release, run
build once, combine with config into an immutable release, run it
VI
Processes
stateless, share-nothing processes; state lives in backing services
VII
Port binding
self-contained app serves HTTP on its own port (embedded server)
VIII
Concurrency
scale out by running more processes
IX
Disposability
fast startup, graceful shutdown on SIGTERM
X
Dev/prod parity
keep dev, staging and prod as alike as possible
XI
Logs
write an event stream to stdout; the platform collects it
XII
Admin processes
run one-off tasks (migrations) as separate processes of the same release
Quick answers
When not to use microservices? A small team, an unclear domain or an early product: start with a modular monolith.
How do you split a monolith? Strangler fig along bounded contexts, starting where seams are clear and change is frequent.
Distributed transactions? Sagas with compensations plus the outbox; avoid 2PC.
Choreography or orchestration? Choreography for short, simple flows; orchestration once flows branch or grow.
How do services find each other? A registry, or platform DNS such as Kubernetes Services.
Gateway vs service mesh? Edge concerns for client traffic vs sidecars for service-to-service traffic.
Preventing cascading failures? Timeouts, retries with backoff and jitter, circuit breakers, bulkheads, fallbacks.
Duplicate messages? Assume at-least-once and make consumers idempotent with dedup keys.
Tracing a request across services? OpenTelemetry with W3C traceparent propagation, and the trace ID in every log line.
Securing service-to-service calls? mTLS for transport identity plus client-credentials tokens for authorization.
Can two services share a database? Avoid it: it couples schemas and releases. Share data through APIs or events.
How do you keep data consistent? Eventual consistency via events, with explicit pending states in the API and UI.
Anti-patterns
Distributed monolith: services that must deploy in lockstep, often through shared domain libraries or a shared database.
Nanoservices and entity services (CRUD per table): chatty calls whose overhead outweighs any benefit.
Long synchronous chains: latency adds up and availability multiplies down. Go async or merge services.
Dual writes without an outbox: lost or phantom events.
Retries without timeouts, backoff or idempotency: retry storms and duplicate side effects.
Smart gateway or ESB full of business logic: a new monolith in the middle.
Big-bang rewrite instead of strangling the old system piece by piece.
No correlation IDs or tracing: debugging means grepping twenty log streams.
Breaking API changes with no versioning or contract tests.
Interview tip
For any design question, name the trade-off out loud: “this buys independent deploys and costs eventual consistency, so here is how I handle it”.