“Could you reproduce the model you deployed six months ago?” is a favourite MLOps interview question because most teams honestly cannot. It tests whether you understand that a trained model depends on more than code, and whether you have the discipline to record those dependencies before anyone asks. In regulated industries (lending, insurance, healthcare) it is not hypothetical: auditors ask exactly this, and “we retrained it and got something similar” is not an acceptable answer.
Before you start
You should know git, basic pandas and how to train a scikit-learn model. Code was run with Python 3.14, scikit-learn 1.9, pandas 3.0 and MLflow 3.16; outputs in comments are real. Lineage means being able to trace an artifact back to everything that produced it. A content hash is a fingerprint, such as SHA-256, computed from the bytes of a file or the values in a table; if one bit changes, the hash changes.
The short answer
To reproduce a model you need five things linked to the registered model version: the code (git commit of a clean working tree), the data (an immutable snapshot or table version plus a content hash, including how labels were defined), the configuration (hyperparameters, feature list, random seeds), the environment (container image digest or lockfile, plus CUDA and driver versions for GPU work), and the artifact itself with its checksum. Data is the part teams forget, because tables are overwritten in place. Even with all five, bit-for-bit retraining can be impossible on GPUs, so you keep the original artifact and define a tolerance for retrained results.
How it works
Reproducibility comes from capturing a manifest at training time and refusing to train when something cannot be captured. A small helper shows the idea:
import hashlib, subprocess, sys
from importlib.metadata import version
def sha256_file(path: str) -> str:
h = hashlib.sha256()
with open(path, "rb") as f:
for chunk in iter(lambda: f.read(1 << 20), b""):
h.update(chunk)
return h.hexdigest()
def git_commit() -> str:
if subprocess.run(["git", "status", "--porcelain"], capture_output=True, text=True).stdout.strip():
raise SystemExit("refusing to train: uncommitted changes")
return subprocess.run(["git", "rev-parse", "HEAD"], capture_output=True, text=True).stdout.strip()
def build_manifest(data_path: str, params: dict, seed: int) -> dict:
return {
"git_commit": git_commit(),
"data_path": data_path,
"data_sha256": sha256_file(data_path),
"params": params,
"seed": seed,
"python": sys.version.split()[0],
"packages": {p: version(p) for p in ["scikit-learn", "numpy", "pandas"]},
}Editing the training script without committing makes the next run stop with refusing to train: uncommitted changes, which is the point: a model trained from code that exists nowhere else can never be explained.
Seeds deserve care. A seed makes random choices repeatable, but only for identical inputs in identical order:
from sklearn.linear_model import SGDClassifier
def train(df, seed):
X, y = df.drop(columns="label").to_numpy(), df["label"].to_numpy()
return SGDClassifier(random_state=seed).fit(X, y)
a, b = train(df, 7), train(df, 7)
print(np.array_equal(a.coef_, b.coef_)) # True
c = train(df.sample(frac=1, random_state=1), 7) # same rows, different order
print(np.array_equal(a.coef_, c.coef_)) # False (max difference 0.145)Same data, same seed, shuffled rows: a different model. A query without ORDER BY, or files read in directory-listing order, is enough to break reproducibility.
Step-by-step walkthrough
Step 1: Train only from immutable data
Write each training set to a path that is never modified, such as s3://ml/churn/train/2026-10-01/, or use a table format with time travel (Delta Lake, Apache Iceberg) and record the version number. DVC gives files git-like versioning by storing a small pointer file in git and the data in remote storage:
dvc init
dvc add data/train.parquet # writes data/train.parquet.dvc with the hash
git add data/train.parquet.dvc .gitignore
git commit -m "Training data 2026-10-01"
dvc push # uploads the data to the configured remote
# months later:
git checkout <commit> && dvc checkout # restores the exact data for that commitStep 2: Hash content, not just names
A path can be overwritten; a hash cannot lie. Hash the file bytes for exact identity. Be aware that file hashes change when the encoding changes: writing the same DataFrame twice to Parquet gave identical SHA-256 values in our run, but switching compression to zstd changed the file hash while pd.util.hash_pandas_object(df, index=False).sum() stayed equal, an order-insensitive fingerprint of the values themselves. Record whichever matches the question you need to answer.
Step 3: Pin the environment
Record library versions at minimum; better, train in a container and record the image digest (sha256:...), which pins the OS, Python and every library. MLflow writes requirements.txt, conda.yaml and python_env.yaml next to each logged model. Pickled or serialised models can fail or change behaviour when loaded under a different library version, which is another reason to pin.
Step 4: Attach everything to the run and the registry
Log the manifest with the run (mlflow.log_dict(manifest, "manifest.json"), plus a data_sha256 tag for searching), register the model from that run, and keep the artifact for as long as the model might be questioned. The registry version then leads to the run, and the run leads to code, data, config and environment.
Worked scenario
A lender’s auditor asked for the exact credit model used in March, and evidence of what data it learned from. The team had the pickle file in a bucket. Retraining from the repository produced an AUC of 0.781 instead of the recorded 0.774, and different decisions for about 3% of past applicants. Three causes emerged. The training query read applications_latest, a table rebuilt nightly, so seven months of late corrections and deleted records had changed the data. The query had no ORDER BY, and the gradient-boosting model sampled rows during training, so row order mattered. And the environment had moved from scikit-learn 1.7 to 1.9, with a different default in one estimator.
The fix changed the process rather than the model. Training sets became dated, immutable Parquet snapshots with a SHA-256 recorded on the run. Queries got explicit ordering. Training ran in a container whose digest was logged, the pipeline refused to start on a dirty git tree, and the registered artifact and its manifest were retained for seven years, matching the lender’s record-keeping policy. For the March model, the team could now show the auditor the original artifact plus everything still recoverable, and say honestly which parts had not been captured at the time.
Common mistake
- “We version the code, so we can reproduce it.” Code is one of four inputs; data changes most often.
- Treating table names as versions.
customers_latestortrain_finaldescribe a role, not a state. - Believing seeds guarantee identical results. Row order, thread counts, library versions and non-deterministic GPU kernels all change outcomes.
- Promising bit-exact GPU retraining. Parallel floating-point reductions can sum in different orders. PyTorch offers
torch.use_deterministic_algorithms(True)at a speed cost, and some operations have no deterministic version. - Deleting old artifacts to save space. The artifact is the cheapest and strongest evidence of what actually ran.
Verify the behavior
Make reproducibility a test. In CI, train twice on a small fixture dataset with the same seed and assert the coefficients or predictions are identical (np.array_equal returned True for us), then shuffle the rows and confirm your pipeline sorts them back, so the assertion still passes. Add a test that modifies a tracked file and checks that training refuses to start. Periodically, pick an old registered version, restore its data, commit and image, retrain, and compare its metrics with the recorded ones within an agreed tolerance.
Follow-up questions
- How do you version a multi-terabyte dataset cheaply? Do not copy it: use a table format with time travel or lakeFS-style branching, which stores only changed files, and record the version id.
- How do deletion requests (for example under GDPR) interact with data snapshots? Snapshots must still honour deletion. Keep snapshots for a limited period, record which records were removed, and accept that later retraining cannot exactly match.
- What is the difference between reproducibility and replicability? Reproducibility gets the same result from the same inputs; replicability gets a consistent conclusion from new data or a new implementation.
- How much of this does MLflow do for you? It tags the git commit, records params, metrics and environment files and links registry versions to runs. Data versioning and the clean-tree rule are yours.
Interview exercise
Your team trains a demand forecast weekly from a warehouse table that is updated hourly. Leadership wants to know why this week’s forecast differs from what last week’s model would have predicted. What do you put in place so this question can be answered from now on?
Answer and reasoning
Separate the two possible causes: a different model or different inputs. For the model, every weekly run records the git commit, image digest, params and seeds, and registers the model with a link to the run, keeping last week’s artifact. For the data, each run trains from a snapshot taken at a fixed cut-off time, either a time-travel version of the table or a dated export with a hash, and inference inputs are logged too. Then I can run last week’s artifact on this week’s inputs and this week’s model on last week’s inputs; the differences attribute the change to the model, the data or both. Without the snapshot and the retained artifact, that comparison is impossible.
Continue learning
- Practise in the MLOps chapter and the MLOps MCQs.
- See how runs and registry versions record lineage in MLflow tracking and model registry, and how to pin the environment in Docker multi-stage builds.
- Reference: the DVC documentation and the MLflow Tracking guide.