Jobs manage completion of batch tasks and may retry failed executions. The task itself must tolerate repeated starts and partial work.
Before you start
You should understand Pods, Deployments and Services. Read desired configuration separately from observed cluster state. Use a development cluster when trying changes, and inspect events and status rather than assuming that an accepted manifest means the workload is ready to serve traffic.
The practical goal is to reason through this situation: An import records progress and uses stable item identity across retries. Read the walkthrough first, then try the interview exercise before opening its answer. The important part is explaining the decision and its consequences, rather than remembering a definition alone.
Step-by-step walkthrough
Step 1: Assume repeated execution
A failed Pod can cause another attempt after partial progress.
Step 2: Make writes repeatable
Use stable item IDs and idempotent updates or checkpoints.
Step 3: Verify completion meaning
Job success should correspond to the business work’s recorded result.
Worked scenario
An import records progress and uses stable item identity across retries.
An importer writes 500 records then crashes before finishing. Its retry starts from the checkpoint or upserts by stable record ID rather than creating duplicates. A successfully completed container is not enough if errors were swallowed and the expected records were never written.
Common mistake
Assuming a Job executes its business action exactly once can duplicate effects.
Verify the behavior
Crash mid-batch, retry and compare logical records and progress markers.
Interview exercise
Make a batch task safe.
Answer and reasoning
Use idempotent writes, checkpoints and an explicit failure policy, including cancellation and completion verification.
Continue learning
Compare the scenario with the Kubernetes interview questions and test your understanding with the Kubernetes MCQs. For terminology and implementation details, consult the reference material.