“Your model scored 0.97 AUC offline and performs much worse in production. Nothing errors. What happened?” Interviewers use this question to test whether you understand training-serving skew, the most common reason a correct-looking model disappoints live. It leads straight into the follow-up “what is a feature store and do we need one?”, which checks whether you know the mechanism a feature store provides or only the product names.
Before you start
You should be comfortable with pandas joins, scikit-learn preprocessing such as StandardScaler, and the idea of training on historical labels. Examples were run on Python 3.14 with pandas 3.0, scikit-learn 1.9 and Feast 0.66. An entity is the thing features describe (a user, a card, a merchant), and an event timestamp is when a feature value or label became true.
The short answer
Training-serving skew is any difference between the feature values or preprocessing a model saw during training and what it receives at prediction time. It comes from three main sources: the same feature implemented twice (SQL for training, application code for serving), stateful preprocessing refitted on live data, and training sets built with values that were not yet known at prediction time. Prevent it with one feature definition shared by both paths, preprocessing shipped inside the model artifact, point-in-time correct training data, and logging of served features so you can compare them with training. A feature store packages those ideas: an offline store of timestamped history for training, an online store of latest values for serving, and point-in-time joins between labels and features.
How it works
The smallest skew bug is refitting preprocessing at serving time. A scaler fitted on training amounts maps 120 to about 3.13. A handler that fits a fresh scaler on each request maps every value to zero:
import numpy as np
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler().fit(np.array([[20.0], [40.0], [60.0], [80.0]]))
print(scaler.transform([[120.0]])) # [[3.13049517]]
print(StandardScaler().fit_transform([[120.0]])) # [[0.]]The model still returns a valid probability, so nothing alerts. The fix is to treat preprocessing as part of the model: a Pipeline or an MLflow pyfunc model is fitted once and serialised whole.
The subtler bug is time. A training row says “on 1 March, did this user churn within 30 days?”. The features for that row must be the values known on 1 March. If the training query joins today’s value of orders_30d, it includes behaviour after the prediction moment, and for churners that behaviour is zero orders, which is the answer. A point-in-time join takes, for each label, the most recent feature value at or before the label’s timestamp. pandas implements it as merge_asof:
import pandas as pd
labels = pd.DataFrame({"user": ["u1", "u1"],
"ts": pd.to_datetime(["2026-03-01 10:00", "2026-03-05 10:00"])})
feats = pd.DataFrame({"user": ["u1", "u1", "u1"],
"ts": pd.to_datetime(["2026-02-28", "2026-03-03", "2026-03-06"]),
"orders_30d": [3, 4, 9]})
pit = pd.merge_asof(labels.sort_values("ts"), feats.sort_values("ts"), on="ts", by="user")
latest = labels.merge(feats.groupby("user", as_index=False).last()[["user", "orders_30d"]], on="user")
print(pit["orders_30d"].tolist(), latest["orders_30d"].tolist()) # [3, 4] [9, 9]A feature store does this join at scale and keeps the two serving paths consistent. Feast, the open-source example, has you declare a feature view once. Training calls get_historical_features, which performs the point-in-time join against the offline store. Serving calls get_online_features, which reads the latest value from a key-value online store that a materialisation job keeps filled.
Step-by-step walkthrough
Step 1: Define the feature once
A Feast repository declares entities, a source with an event timestamp column, and a feature view. ttl bounds how far back the join may look for a value:
from datetime import timedelta
from feast import Entity, FeatureView, Field, FileSource, ValueType
from feast.types import Int64
user = Entity(name="user", join_keys=["user_id"], value_type=ValueType.INT64)
orders_source = FileSource(path="data/user_orders.parquet", timestamp_field="event_timestamp")
user_orders = FeatureView(
name="user_orders",
entities=[user],
ttl=timedelta(days=3),
schema=[Field(name="orders_30d", dtype=Int64)],
source=orders_source,
)feast apply registers it. Both training and serving now refer to user_orders:orders_30d, so there is no second implementation to drift apart.
Step 2: Build training data with a point-in-time join
The entity dataframe holds labels with their timestamps. Feast returns each row joined to the value valid at that moment:
from feast import FeatureStore
store = FeatureStore(repo_path=".")
labels = pd.DataFrame({
"user_id": [1, 1],
"event_timestamp": pd.to_datetime(["2026-03-01 10:00", "2026-03-05 10:00"], utc=True),
"churned": [0, 1],
})
training_df = store.get_historical_features(
entity_df=labels, features=["user_orders:orders_30d"]).to_df()
# user_id event_timestamp churned orders_30d
# 1 2026-03-01 10:00:00+00:00 0 3
# 1 2026-03-05 10:00:00+00:00 1 4Step 3: Materialise and serve the latest values
Materialisation copies the newest value per entity into the online store. At request time the service fetches by key, in milliseconds:
from datetime import datetime, timezone
store.materialize(start_date=datetime(2026, 2, 1, tzinfo=timezone.utc),
end_date=datetime(2026, 3, 7, tzinfo=timezone.utc))
print(store.get_online_features(features=["user_orders:orders_30d"],
entity_rows=[{"user_id": 1}]).to_dict())
# {'user_id': [1], 'orders_30d': [9]}In production materialisation runs on a schedule (feast materialize-incremental) or a streaming job pushes values; its lag is itself a skew source to monitor.
Step 4: Log and compare what was served
Every prediction logs the feature vector it used. A daily job joins those logs to offline values for the same entity and timestamp and reports a mismatch rate per feature. This is the only way to catch skew that already exists.
Worked scenario
A food delivery team added is_weekend to an order-value model. The training pipeline computed it in SQL from UTC timestamps; the online service computed it from the user’s local time in New York. Offline metrics improved, live metrics did not. A comparison of logged online values with the offline recomputation found the gap:
import numpy as np, pandas as pd
rng = np.random.default_rng(3)
ts = pd.Timestamp("2026-09-01", tz="UTC") + pd.to_timedelta(
rng.integers(0, 30 * 24 * 3600, 10_000), unit="s")
served = pd.DataFrame({"ts": ts})
served["online"] = served["ts"].dt.tz_convert("America/New_York").dt.dayofweek >= 5
served["offline"] = served["ts"].dt.dayofweek >= 5
bad = served["online"] != served["offline"]
print(f"mismatch rate: {bad.mean():.1%}") # mismatch rate: 4.4%
print(served.assign(h=served["ts"].dt.hour)[bad]["h"].unique()) # UTC hours 0 to 3 onlyEvery mismatch sat between midnight and 4 a.m. UTC on weekend boundaries, which is Friday and Sunday evening in New York, when ordering peaks. The fix was to define the feature once, in the feature store, with an explicit time zone column, and to add the mismatch rate to the daily monitoring job with an alert above 0.5%.
The leakage version of skew is worse because it inflates offline results. In a synthetic churn set where churners stop ordering, a model trained on “orders in the last 30 days as of today” reached 0.977 offline AUC against 0.681 for the point-in-time feature. At a 0.5 threshold it caught 100% of churners offline and 16.5% on the features it actually receives live.
Common mistake
- Calling skew “drift”. Drift is the world changing; skew is your two code paths disagreeing. Skew exists from day one and retraining does not fix it.
- Saying a feature store is mainly a database. The value is the shared definition, point-in-time joins and the offline-online consistency, not storage.
- Using the latest snapshot to build training sets. It is the classic way to leak the future, and it produces the best offline numbers you will ever see.
- Recommending a feature store for every team. One batch model with features from one warehouse table does not need another system to operate.
Verify the behavior
Write a parity test that runs in CI: take a fixed set of entities and timestamps, compute each feature through the offline path and through the online code path, and assert they match exactly (or within float tolerance). In production, the equivalent is the served-versus-offline comparison from Step 4, tracked as a per-feature mismatch rate. For leakage, retrain once with features shifted a day earlier than the label time; if the offline score barely changes, the original features were not peeking.
Follow-up questions
- What happens if materialisation falls behind? Online values go stale, so live features differ from what a point-in-time training row would have seen. Monitor feature freshness and set a
ttlso very old values come back as missing rather than silently wrong. - How do streaming features fit in? A stream processor such as Flink computes windowed aggregates and pushes them to the online store and the offline log, using the same definition.
- How would you do point-in-time joins without a feature store? Keep feature tables with valid-from timestamps and use
merge_asofor an as-of join in SQL, always filtering on feature time at or before label time. - Is training on logged served features a good idea? Often yes: it guarantees the model trains on exactly what serving produces, at the cost of only having history from when logging began.
Interview exercise
A ride-hailing ETA model uses driver_trips_today. Training builds it with a SQL COUNT over each driver’s completed trips that day. The online service increments a Redis counter when a trip starts. Offline MAE is 2.1 minutes; live MAE is 3.4. What do you suspect and how do you fix it?
Answer and reasoning
Two definitions of one feature. Offline counts completed trips over the whole day, which for a morning prediction includes trips completed later that day: future information. Online counts started trips so far, a different quantity. The model learned from values systematically higher than it sees live, and partly from the future. I would confirm by logging the served value and comparing it with the offline value at the same timestamp, then redefine the feature as completed trips before the prediction time, computed by one pipeline that writes both the online value and the timestamped offline history, and rebuild the training set with point-in-time joins. Offline MAE will probably get worse, and that honest number is the one to compare with live MAE.
Continue learning
- Practise in the MLOps chapter and the MLOps MCQs.
- Leakage starts before serving: see cross-validation and data leakage. For catching skew once a model is live, read data drift vs concept drift.
- Reference: the Feast documentation and Google Cloud’s MLOps continuous delivery guide, which covers feature stores and skew.