Ch. 10 · Microservices

Microservices Health Check Endpoints

Separate liveness from readiness, keep deep dependency checks off the restart path, and avoid the restart loop.

~2 min readintermediateupdated Oct 5, 2026

Health endpoints tell the platform whether an instance is alive and whether it can serve traffic. They are easy to get wrong: a liveness check that fails on a dependency outage restarts every instance, turning a partial degradation into a full outage.

Before you start

You should understand load balancing and basic orchestration. This article uses Kubernetes liveness and readiness probes as the example.

Step-by-step walkthrough

Step 1: Liveness means the process is stuck

A liveness probe answers “should this instance be restarted”. It should check that the process responds, not that every dependency is healthy, because a dependency outage is not fixed by restarting the app. Keep it shallow and cheap.

Step 2: Readiness means ready for traffic

A readiness probe answers “should traffic go here”. It can check that the instance is warmed up and can reach critical dependencies, and when it fails the load balancer stops sending traffic without restarting the instance. This is where dependency awareness belongs.

Step 3: Put deep checks off the hot path

A thorough dependency check is valuable for diagnostics but too heavy for a probe that runs every few seconds. Expose it on a separate endpoint used by dashboards, or cache the result, so probing stays cheap and does not add load to dependencies.

Worked scenario

The liveness probe is shallow and the readiness probe is bounded.

GET /health/live  -> 200 if the process answers     (liveness, no dependency checks)
GET /health/ready -> 200 if warmed up and deps ok   (readiness)
GET /health/deep  -> detailed dependency status     (diagnostics, not a probe)
Text

Walk through the example

If the database is down, /health/live still returns 200, so the platform does not restart healthy instances. /health/ready returns 503, so traffic is diverted until the dependency recovers. The deep endpoint gives engineers detail without driving restarts.

Common mistake

Failing liveness on a dependency check, which restarts every instance during a dependency outage and multiplies the damage. Another is a readiness check that always passes, so traffic reaches an instance that cannot serve it.

Verify the behavior

Take a dependency down and confirm liveness stays 200 while readiness fails. Confirm the load balancer stops routing to the unready instance. Measure probe latency and confirm it stays low and does not hammer the dependency.

Interview exercise

Why should a dependency outage fail readiness but not liveness?

Answer and reasoning

Readiness controls traffic, so diverting requests away from an instance that cannot reach its dependency is correct and reversible; when the dependency returns, readiness passes and traffic resumes. Liveness controls restarts, and restarting does not fix a dependency outage — it only removes capacity, so failing liveness would turn a partial problem into a wider one.

Continue learning

Compare probe configuration in Kubernetes probe separation and startup readiness in Startup readiness. Read the Kubernetes liveness and readiness documentation and try the Microservices interview questions.

More in Microservices

read ✓Microservices · hard

Microservices Anti-Corruption Layer

Protect a service's domain model from a foreign or legacy model with a translation layer at the boundary.

~2 min readread →
read ✓Microservices · hard

Microservices API Versioning and Evolution

Evolve service APIs without breaking consumers using additive changes, explicit versioning and consumer-driven contracts.

~2 min readread →
read ✓Microservices · hard

Microservices Backend for Frontend

Use a per-client BFF to aggregate services and shape responses, without letting it become a shared god service.

~2 min readread →
esc