Everything a Docker round tends to touch, from how an image is built to why docker stop hangs for ten seconds.
Containers vs VMs
- A container is an ordinary Linux process, isolated by namespaces (what it can see:
pid,net,mnt,uts,ipc,user,cgroup) and limited by cgroups (what it can use: CPU, memory, PIDs, I/O). - Image = read-only layers + config (entrypoint, env, user). Container = image + a thin writable layer + runtime state. One image, many containers.
- Engine stack:
dockerCLI →dockerd(REST API, builds) →containerd(images, lifecycle) →runc(OCI runtime that creates the namespaces and cgroups). - On macOS and Windows, Docker Desktop runs Linux containers inside a lightweight Linux VM.
| Container | Virtual machine | |
|---|---|---|
| Isolates | processes (namespaces, cgroups) | hardware (hypervisor) |
| Kernel | shared with the host | own guest kernel |
| Size / startup | MBs, usually sub-second | GBs, usually tens of seconds |
| Security boundary | weaker: a kernel exploit can escape | stronger |
| Other OS kernels | no (Linux images need a Linux kernel) | yes |
Images, layers & build cache
- Each
RUN,COPYandADDadds a filesystem layer;ENV,CMD,EXPOSE,LABELand friends only change metadata. - Layers are content-addressed and shared between images. A container’s writes go to a copy-on-write layer that is deleted with the container.
- Deleting a file in a later layer hides it but doesn’t shrink the image: clean up in the same
RUN. - Cache rule: a step is reused if the instruction and its inputs are unchanged.
COPY/ADDcompare file checksums (not mtimes);RUNcompares only the command string. After the first miss, every later step rebuilds. - Tags (
nginx:1.27) are mutable pointers; digests (nginx@sha256:…) are immutable.latestis only the default tag, not “the newest”. - Inspect with
docker history imganddocker image inspect img; rebuild with--no-cache, refresh the base with--pull.
Interview tip
Order a Dockerfile from least to most frequently changed: base image → OS packages → dependency manifest (
package.json,go.mod) → install → source. Then a code change only rebuilds the last layers.
Dockerfile instructions
| Instruction | Does | Remember |
|---|---|---|
FROM img AS build |
base image, starts a stage | ARG before the first FROM works only in FROM lines |
RUN |
runs at build time, commits a layer | chain with &&; --mount=type=cache or type=secret |
COPY |
files from the context or --from=stage |
the default choice; --chown, --chmod |
ADD |
COPY + URLs, git repos, auto-extracts local tars | remote tarballs are not extracted |
CMD |
default command or default args | only the last CMD counts; docker run img x replaces it |
ENTRYPOINT |
the fixed executable | run args append to it; override with --entrypoint |
ENV |
variable at build and run time | baked into the image config |
ARG |
build-time variable (--build-arg) |
gone at run time, but visible in docker history |
WORKDIR |
cwd for later steps, created if missing | use it instead of RUN cd … |
USER |
user for later RUN, CMD, ENTRYPOINT |
default is root; a numeric UID lets Kubernetes verify runAsNonRoot |
EXPOSE |
documents a port (TCP by default) | does not publish; -p/-P does |
HEALTHCHECK |
periodic check: exit 0 healthy, 1 unhealthy | defaults: interval 30s, timeout 30s, retries 3; HEALTHCHECK NONE disables |
VOLUME |
declares a mount point | creates an anonymous volume per container; child images can’t undo it |
CMD vs ENTRYPOINT, exec vs shell form
- Exec form
["node", "server.js"]: a JSON array (double quotes), no shell, no$VARexpansion; your process is PID 1 and receives signals. - Shell form
node server.js: runs/bin/sh -c "node server.js"; variables expand, but the shell sits in front of your app and may not forwardSIGTERM.
| ENTRYPOINT | CMD | docker run img |
docker run img x |
|---|---|---|---|
| none | ["nginx", "-g", "daemon off;"] |
the CMD | x |
["python", "app.py"] |
["--port", "8000"] |
python app.py --port 8000 |
python app.py x |
python app.py (shell) |
anything | /bin/sh -c 'python app.py' |
same: CMD and x are ignored |
Multi-stage build & .dockerignore
FROM golang:1.25 AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 go build -o /out/app ./cmd/app
FROM gcr.io/distroless/static-debian12:nonroot
COPY --from=build /out/app /app
USER 65532:65532
ENTRYPOINT ["/app"]- Compilers, sources and caches stay in
build; only the binary ships. Final stage options:scratch, distroless,-slim,alpine. docker build --target build .stops at a stage (useful for tests); BuildKit skips stages the target doesn’t need.- Build secrets:
RUN --mount=type=secret,id=npmrc,target=/root/.npmrc npm ciwithdocker build --secret id=npmrc,src=$HOME/.npmrc .. The secret never lands in a layer. - Dependency cache across builds:
RUN --mount=type=cache,target=/root/.cache/go-build go build …. .dockerignore(context root,.gitignore-like syntax) keeps.git,node_modules,.env, build output and logs out of the context: smaller uploads, fewer cache misses, no leaked secrets.
CLI cheat table
| Task | Command |
|---|---|
| Build | docker build -t app:1.0 . (-f, --target, --no-cache, --build-arg) |
| Run | docker run -d --name web -p 8080:80 nginx |
| Shell in | docker exec -it web sh (running containers only) |
| Logs | docker logs -f --tail 100 web (also works after it exited) |
| List | docker ps (running), docker ps -a (all), docker images |
| Inspect | docker inspect -f '{{.State.ExitCode}}' web |
| Live usage | docker stats, docker top web, docker events |
| Stop | docker stop web (SIGTERM, then SIGKILL after 10 s); docker kill web (SIGKILL now) |
| Remove | docker rm -f web, docker rmi app:1.0 |
| Copy files | docker cp web:/etc/nginx/nginx.conf . |
| Publish | docker tag app:1.0 registry.example.com/team/app:1.0 then docker push it |
| Multi-arch | docker buildx build --platform linux/amd64,linux/arm64 -t … --push . |
| Clean up | docker system prune: stopped containers, unused networks, dangling images, unused build cache |
| Clean harder | add -a (all unused images) and --volumes (unused anonymous volumes) |
docker run flags
| Flag | Meaning |
|---|---|
-d / -it |
detached / interactive with a TTY |
--name web |
fixed name, also its DNS name on user-defined networks |
-p 8080:80 |
publish host:container on all interfaces; -p 127.0.0.1:8080:80 for local only; -P maps every EXPOSEd port to a random host port |
-v data:/data, -v "$PWD":/app:ro |
named volume, bind mount (read-only); --mount type=…,src=…,dst=… is the explicit form |
--rm |
delete the container when it exits (not allowed with --restart) |
-e KEY=val, --env-file .env |
environment; bare -e KEY copies the host’s value |
--network app-net |
attach to a network (or host, none, container:<name>) |
--restart |
no (default), on-failure[:N], always, unless-stopped |
-m 512m, --cpus 1.5, --pids-limit 200 |
memory hard cap (OOM kill above it), CPU quota, process cap |
-u 1000:1000, -w /app |
user, working directory |
--entrypoint sh, --init |
override ENTRYPOINT; run a tiny init as PID 1 |
--read-only, --cap-drop ALL, --security-opt no-new-privileges |
hardening |
alwaysvsunless-stopped: both restart on crash; after a manualdocker stop,alwayscomes back when the daemon restarts,unless-stoppeddoesn’t.--cpusis a hard quota;--cpu-sharesis only a relative weight under contention.
Volumes, bind mounts & tmpfs
| Named volume | Bind mount | tmpfs | |
|---|---|---|---|
| Lives in | Docker-managed storage | any host path | host memory |
| Syntax | -v pgdata:/var/lib/postgresql/data |
-v "$PWD":/app |
--tmpfs /tmp |
| Lifetime | survives docker rm until docker volume rm |
owned by the host | ends when the container stops |
| Use for | databases, persistent state | dev live reload, config files | scratch data that must not hit disk |
| Portability | high; volume drivers (NFS, cloud) | tied to the host layout | Linux containers only |
- A new, empty named volume is pre-filled with the image’s files at that path; a bind mount hides whatever the image had there.
-vsilently creates a missing host path (as a directory);--mount type=bindfails instead.- Anonymous volumes (from
VOLUMEor-v /path) pile up:docker volume prune(add-afor unused named ones too).
Networking
| Mode | What you get |
|---|---|
bridge (default) |
private subnet behind NAT; reach it from outside with -p |
| user-defined bridge | docker network create app-net: built-in DNS by container name, isolation per app |
host |
shares the host’s network stack; no port mapping, no network isolation |
none |
loopback only |
container:<name> |
joins another container’s network namespace (like a Kubernetes pod) |
overlay |
spans several hosts (Swarm) |
macvlan / ipvlan |
container gets its own address on the physical LAN |
- On the default bridge, containers can only reach each other by IP; name resolution needs a user-defined network (
--linkis legacy). localhostinside a container is the container itself. Reach the host viahost.docker.internal(Docker Desktop; on Linux add--add-host=host.docker.internal:host-gateway).- A server must listen on
0.0.0.0, not127.0.0.1, or-pforwards to nothing.
Docker Compose
services:
api:
build: .
ports: ["8080:8080"]
environment: { DB_HOST: db }
depends_on: { db: { condition: service_healthy } }
db:
image: postgres:16-alpine
environment: { POSTGRES_PASSWORD: example }
volumes: [pgdata:/var/lib/postgresql/data]
healthcheck: { test: ["CMD-SHELL", "pg_isready -U postgres"], interval: 5s }
volumes: { pgdata: {} }- Commands:
docker compose up -d --build,ps,logs -f api,exec api sh,config(render the merged file),down(containers + network),down -v(also volumes). - Each project gets its own network; services find each other by service name (
db:5432). depends_onalone only orders startup;condition: service_healthyplus ahealthcheckwaits for readiness.compose.yamlis the preferred file name,compose.override.yamlmerges automatically, and the top-levelversion:key is obsolete..envin the project directory feeds${VAR}interpolation in the file;env_file:sets variables inside the container.docker compose(V2, a CLI plugin) replaced the old Pythondocker-compose.
Smaller, safer images
- Size: multi-stage builds; slim, distroless or
scratchbases;apt-get update && apt-get install -y --no-install-recommends … && rm -rf /var/lib/apt/lists/*in oneRUN; a tight.dockerignore. - Alpine is tiny but uses musl libc: glibc-built binaries and some wheels break or run slower.
- Run as a non-root
USER; never use--privilegedcasually; mounting/var/run/docker.sockgives root on the host. - Pin base images by version (or digest for reproducibility), rebuild often to pick up patches, and scan (Docker Scout, Trivy).
- At run time:
--read-onlywith--tmpfs /tmp,--cap-drop ALLthen add back only what’s needed,no-new-privileges, memory/CPU/PID limits, rootless mode or user namespaces where possible.
Gotcha
Secrets in
ENV,ARGor aCOPYed file persist in the image:ARGvalues show up indocker history, and a file deleted in a later layer is still in the earlier one. UseRUN --mount=type=secretat build time and inject secrets at run time.
PID 1 & signals
docker stopsendsSIGTERM(or the image’sSTOPSIGNAL) to PID 1, waits 10 s by default (-t,--stop-timeout), then sendsSIGKILL.- The kernel skips default signal actions for PID 1: without a handler,
SIGTERMis simply ignored, so every stop waits the full timeout and ends in exit 137. - PID 1 must also reap orphaned children, or zombies accumulate.
- Fixes: exec form;
exec "$@"at the end of entrypoint scripts; handleSIGTERMin the app (stop accepting work, drain, close connections);docker run --initor tini to forward signals and reap zombies. - Wrappers like
npm startorsh -cadd a parent process that may not pass signals on; runnode server.jsdirectly.
#!/bin/sh
set -e
# one-time setup, then hand PID 1 to the real process
[ -n "$DB_HOST" ] || { echo "DB_HOST is required" >&2; exit 1; }
exec "$@" # used with ENTRYPOINT ["/entrypoint.sh"] and CMD ["node", "server.js"]Debugging a crashing container
docker ps -afor the status and exit code, thendocker logs --tail 50 app.docker inspect -f '{{.State.ExitCode}} {{.State.OOMKilled}} {{.State.Error}}' app.- Check what actually runs:
docker image inspect -f '{{.Config.Entrypoint}} {{.Config.Cmd}} {{.Config.User}}' app:1.0. - Reproduce by hand:
docker run -it --rm --entrypoint sh app:1.0, then start the command yourself. Distroless has no shell: use its:debugvariant ordocker cpfiles out. - Watch
docker stats(memory near the limit?) anddocker events;docker diff applists files changed in the container.
| Exit code | Meaning | Usual cause |
|---|---|---|
| 0 | main process finished | app daemonized itself or CMD is a one-shot; keep it in the foreground |
| 1 | application error | read the logs |
| 125 | Docker itself failed | bad flag, name conflict, daemon error |
| 126 | not executable | missing chmod +x, permission denied |
| 127 | command not found | typo, binary missing from a slim image, PATH |
| 137 | SIGKILL (128 + 9) | OOM kill (OOMKilled=true) or SIGTERM ignored at stop |
| 139 | SIGSEGV (128 + 11) | native crash, glibc/musl mismatch |
| 143 | SIGTERM (128 + 15) | graceful docker stop |
Quick answers
- Image vs container? An image is the read-only template; a container is an instance of it with its own writable layer and process.
- CMD vs ENTRYPOINT? ENTRYPOINT is the executable; CMD gives default arguments (or the default command);
docker runargs replace CMD. - COPY vs ADD? COPY for local files; ADD only when you need local tar extraction or a URL/git source.
- ARG vs ENV? ARG exists only during the build; ENV is stored in the image and set in every container.
- Why did my data disappear? It was written to the container layer, which
docker rmdeletes; use a volume. - How do two containers talk? Put them on the same user-defined network (or Compose project) and use the container or service name, not
localhost. - EXPOSE vs
-p? EXPOSE is documentation (and feeds-P); only-p/-Popens a port on the host. - stop vs kill?
stopsends SIGTERM and waits (10 s default) before SIGKILL;killsends SIGKILL right away. - save/load vs export/import?
savewrites an image with all layers and metadata;exportflattens one container’s filesystem with no history or config. - Dangling image? An untagged
<none>:<none>image, usually left when a tag is rebuilt;docker image pruneremoves them. - Why is the build cache useless?
COPY . .comes before the dependency install, or a huge unignored context changes on every build. - Is a container as isolated as a VM? No: it shares the host kernel, so harden it (non-root, dropped capabilities, seccomp, limits).
- Why does my container exit immediately? PID 1 finished: the process forked into the background or the CMD was a one-off command.
Gotchas & traps
RUN apt-get updateon its own line is cached forever, so later installs use stale package lists: combine it with the install.- Exec form doesn’t expand variables:
CMD ["echo", "$HOME"]prints$HOME. Use["sh", "-c", "echo $HOME"]when you need a shell. - Single quotes in exec form are invalid JSON, so Docker silently treats the line as shell form.
- Windows CRLF line endings in an entrypoint script give
exec /entrypoint.sh: no such file or directory. - A bind mount over
/apphides the image’snode_modules; add an anonymous volume for/app/node_modulesin dev setups. - Bind-mounted files keep host UIDs, so a non-root container user can get
permission denied. - The default
json-filelog driver doesn’t rotate logs: setmax-size/max-fileor use thelocaldriver before the disk fills. - On Linux, published ports are opened through Docker’s own iptables rules and can bypass firewalls such as ufw.
- An
unhealthystatus doesn’t restart anything in plain Docker (Swarm replaces the task; Kubernetes ignoresHEALTHCHECKand uses its own probes). docker compose down -vdeletes named volumes, including your database.