Ch. 31 · MLOps

MLflow Experiment Tracking and Model Registry Explained

Log params, metrics and models with MLflow 3, compare runs with search_runs, and promote or roll back models by moving registry aliases.

~7 min readbeginnerupdated Oct 6, 2026

“How do you track experiments, and how does a model get from a training run into production?” Almost every MLOps interview asks some form of this, and MLflow is the tool most candidates name. Interviewers are not testing whether you memorised the API. They want to hear that you can answer “which data and code produced the model that is serving right now?” in under a minute, and that promotion and rollback are boring, auditable operations rather than someone copying a file.

Before you start

You should be able to train a scikit-learn model and know what a validation metric is. The examples use MLflow 3.16 with scikit-learn 1.9 on Python 3.14, run inside a git repository, with a local SQLite tracking store (sqlite:///mlflow.db). A team would point mlflow.set_tracking_uri at a shared tracking server instead; the code is otherwise identical. Outputs in comments are from real runs.

The short answer

MLflow Tracking records each training run inside an experiment: parameters, metrics (with history), tags, and artifacts such as the model, plots and environment files. Runs can be searched and compared. The Model Registry gives a model a name and immutable, numbered versions, each linked to the run that produced it. Deployment code loads a model through an alias such as models:/churn-model@champion, so promoting a new version or rolling back means moving the alias, with no code change. Stages (Staging, Production) are deprecated in favour of aliases since MLflow 2.9.

How it works

A run is a context manager. Inside it you log what you want to be able to compare later:

import mlflow
from mlflow.models import infer_signature
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
from sklearn.model_selection import train_test_split

mlflow.set_tracking_uri("sqlite:///mlflow.db")
mlflow.set_experiment("churn")

X, y = make_classification(n_samples=2000, n_features=20, n_informative=5,
                           flip_y=0.1, random_state=0)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=0)

for C in [0.001, 1.0]:
    with mlflow.start_run(run_name=f"logreg-C{C}"):
        mlflow.log_params({"C": C, "max_iter": 500})
        mlflow.set_tag("data_version", "churn-2026-10-01")
        model = LogisticRegression(C=C, max_iter=500).fit(X_tr, y_tr)
        auc = roc_auc_score(y_te, model.predict_proba(X_te)[:, 1])
        mlflow.log_metric("val_auc", auc)
        info = mlflow.sklearn.log_model(
            model, name="model",
            signature=infer_signature(X_tr, model.predict(X_tr)),
            registered_model_name="churn-model",
        )
        print(f"C={C} auc={auc:.4f} version={info.registered_model_version}")
# C=0.001 auc=0.8649 version=1
# C=1.0 auc=0.8836 version=2
python

Four kinds of data are being recorded. Params are inputs and are immutable within a run: logging C again with a different value raises MlflowException: Changing param values is not allowed. Metrics are outputs and can be logged repeatedly with a step, which is how training curves are stored; run.data.metrics returns the latest value of each. Tags are free-form metadata. MLflow adds system tags automatically, including mlflow.source.git.commit when the script runs inside a git repository. Artifacts are files. log_model writes the model plus an MLmodel descriptor, the input signature and pinned environment files (requirements.txt, conda.yaml, python_env.yaml). In this MLflow version the scikit-learn flavour serialised the model with skops (model.skops) rather than a raw pickle.

Passing registered_model_name also creates a registry version. Each version is immutable and remembers its source run, so lineage from a serving model to its params, metrics, data tag and commit is one lookup.

Step-by-step walkthrough

Step 1: Decide what every run must record

Agree on a minimum per run: params, the primary validation metric, the data version, and the model with a signature. Code and environment are captured automatically; data is not, so the data_version tag (a snapshot id, table version or DVC revision) is the line teams most often forget. mlflow.autolog() captures params and metrics for many frameworks, but it cannot know which dataset you meant.

Runs are queryable. Filter keys need a prefix (metrics., params., tags.); a bare val_auc raises an invalid attribute key error:

runs = mlflow.search_runs(experiment_names=["churn"],
                          filter_string="metrics.val_auc > 0.85",
                          order_by=["metrics.val_auc DESC"])
print(runs[["params.C", "metrics.val_auc", "tags.data_version"]].to_string(index=False))
# params.C  metrics.val_auc tags.data_version
#      1.0         0.883603  churn-2026-10-01
#    0.001         0.864903  churn-2026-10-01
python

The same query works in CI, which is how an automated gate finds the current best candidate.

Step 3: Promote by alias

Promotion is a pointer move. Serving code never names a version number:

from mlflow import MlflowClient

client = MlflowClient()
client.set_registered_model_alias("churn-model", "champion", 2)
champ = client.get_model_version_by_alias("churn-model", "champion")
print("champion ->", champ.version)            # champion -> 2

model = mlflow.pyfunc.load_model("models:/churn-model@champion")
python

A common convention uses @challenger for the candidate in shadow or canary and @champion for the live model. Several aliases can exist at once, which the old single “Production” stage could not express.

Step 4: Roll back by moving the alias back

client.set_registered_model_alias("churn-model", "champion", 1)
print(client.get_registered_model("churn-model").aliases)   # {'champion': 1}
python

Serving processes pick up the change when they next resolve the alias: on restart, or through a periodic reload you build into the service. Decide which, because a rollback that waits for the next deploy is not fast.

Worked scenario

A team’s serving container loaded models:/fraud-model/latest at start-up. It had worked for months because only the weekly pipeline registered versions. Then a data scientist ran an experiment with registered_model_name="fraud-model" to share it with a colleague. The next autoscaling event started new pods, which loaded the experimental version 14; older pods still served version 13. For six hours two models handled live traffic, one of them never validated, and the dashboards showed a strange bimodal approval rate.

The fixes were small. Serving switched to models:/fraud-model@champion, so new registrations change nothing until someone moves the alias. Only the pipeline’s service account could set the champion alias, after the validation gate passed, and experiments registered under a separate name. The service logged the resolved version number with every prediction, so a split like this would show up on the first dashboard. Registering version 3 of a model whose champion points at version 2 leaves the alias on 2, which is exactly the property they needed.

Common mistake

  • Treating the registry as file storage. The value is lineage and pointers. A bucket of model_final_v2.pkl files has neither.
  • Serving “latest”. Any registration silently changes what new pods load.
  • Logging the metric you tuned on as the final score. Track validation and test separately, or the run history becomes a record of overfitting.
  • Saying stages are the way to promote. Stages still appear in old tutorials, but aliases and tags replace them in current MLflow.
  • Forgetting the data version. Without it, a run cannot be reproduced, however carefully code is tracked.

Verify the behavior

Run the training script twice, then check three things. mlflow.get_run(run_id).data.tags["mlflow.source.git.commit"] exists, proving code lineage. client.get_model_version_by_alias("churn-model", "champion").run_id returns the run you expect, proving model lineage. Finally, register a third version and confirm champion still resolves to the old one. Running mlflow ui --backend-store-uri sqlite:///mlflow.db shows the same runs, metrics and versions in the browser.

Follow-up questions

  • How would you track a hyperparameter search with hundreds of trials? One parent run with nested child runs (mlflow.start_run(nested=True)), logging the best trial’s params on the parent.
  • How do you share models across dev, staging and production? Either one registry with environment-specific aliases, or separate workspaces where a promotion copies a version (copy_model_version) and keeps lineage.
  • How does a serving process notice an alias move? It resolves the alias on load, so either reload on a schedule or trigger a restart from the promotion job.
  • What does the signature buy you? Input validation at serving time and documentation of the expected schema, catching column-order and type mismatches.

Interview exercise

Your manager asks: “The churn model’s approvals jumped yesterday. Which model is live, what data trained it, and can we go back to last week’s model in five minutes?” Describe how your MLflow setup answers each part.

Answer and reasoning

The service logs the resolved model version with each prediction, and get_model_version_by_alias("churn-model", "champion") confirms which version the alias points to now; the registry’s history shows when it moved and who moved it. That version’s run_id leads to the run, which holds the params, metrics, the data_version tag and the git commit, so I can say exactly what trained it. Rollback is set_registered_model_alias("churn-model", "champion", <previous version>) followed by the service’s reload, a minute or two, provided last week’s model still accepts the features the service sends. Then I investigate the approval jump offline, with the bad version safely out of traffic.

Continue learning

More in MLOps

esc