“Would you serve this model in batch or in real time?” sounds like a quick architecture question, but interviewers use it to see whether you think about the decision the model supports rather than the model itself. Strong answers start from when the inputs exist, how stale a prediction may be, and what happens when the model is slow or unavailable. Weak answers say “real time is better” and then cannot explain p99 latency or what the service returns when the feature store times out.
Before you start
You should know how to train a model and call predict_proba, and have a rough idea of what a REST API is. Examples use Python 3.14 with scikit-learn 1.9, pandas 3.0 and FastAPI 0.142. Latency figures are from a laptop and will differ on your hardware; the ratios are what matter. p99 latency is the response time that 99% of requests beat, and an SLO is the target the service promises for it.
The short answer
Batch inference scores a known set of entities on a schedule and stores the results, for example nightly churn scores for every customer written to a table. Online inference scores one request at a time behind an API within a latency budget, for example a fraud check during checkout. Choose batch when the decision can use data that is hours old and the entities are known in advance; it is cheaper, easier to validate and easy to rerun. Choose online when the prediction depends on information that only exists at request time (the basket, the query, the transaction) or the space of inputs is too large to precompute. The common hybrid precomputes what it can in batch and scores only the request-dependent part online.
How it works
The two patterns differ in cost profile as much as in architecture. Models are efficient on matrices, and each call carries fixed overhead (input validation, Python dispatch, memory allocation). Scoring 2,000 rows in one call versus 2,000 single-row calls with the same gradient-boosting model shows it:
import time, numpy as np
from sklearn.datasets import make_classification
from sklearn.ensemble import HistGradientBoostingClassifier
X, y = make_classification(n_samples=20_000, n_features=20, random_state=0)
model = HistGradientBoostingClassifier(random_state=0).fit(X, y)
rows = X[:2_000]
t = time.perf_counter(); model.predict_proba(rows); batch = time.perf_counter() - t
lat = []
for r in rows:
t = time.perf_counter(); model.predict_proba(r.reshape(1, -1))
lat.append(time.perf_counter() - t)
lat = np.array(lat) * 1000
print(f"one call: {batch*1000:.1f} ms; single calls: {lat.sum():.0f} ms, "
f"p50 {np.percentile(lat, 50):.2f} ms, p99 {np.percentile(lat, 99):.2f} ms")
# one call: 3.5 ms; single calls: 3409 ms, p50 1.45 ms, p99 5.49 msRow-at-a-time scoring was roughly a thousand times more expensive in total, and that is before network hops and feature lookups. Online serving pays this price for freshness. Notice also that p99 is almost four times p50 even in a tight loop on one machine; in a real service with network calls the tail is usually much wider.
Batch jobs therefore optimise for throughput and correctness: process partitions, validate the output, write it atomically, and make reruns safe. Online services optimise for tail latency and availability: keep the model in memory, fetch features in one round trip, put a timeout on every dependency and decide in advance what to return when something is slow.
Step-by-step walkthrough
Step 1: Ask the four deciding questions
When does each input become known? How old can the prediction be before it hurts the decision? How many entities are there, and are they known in advance? What does the caller do if no prediction is available? A nightly email campaign answers “yesterday’s data, a few million known customers, skip the email”. A checkout fraud check answers “the basket exists only now, unbounded combinations, must still decide”. Those answers pick the pattern.
Step 2: Build the batch job to be rerunnable
A batch job should write each run to its own partition, record the model version, and validate before publishing:
import pandas as pd
def score_partition(df: pd.DataFrame, run_date: str, model_version: str) -> pd.DataFrame:
out = df[["customer_id"]].copy()
out["churn_score"] = model.predict_proba(df[["f1", "f2", "f3", "f4"]].to_numpy())[:, 1]
out["model_version"] = model_version
out["run_date"] = run_date
return out
CHUNK = 2_500
scores = pd.concat(score_partition(customers.iloc[i:i + CHUNK], run_date, "7")
for i in range(0, len(customers), CHUNK))
assert scores["customer_id"].is_unique and scores["churn_score"].between(0, 1).all()
scores.to_parquet(f"churn_scores_{run_date}.parquet", index=False) # rerun overwritesBecause the output for a date is overwritten as a whole, a failed or repeated run cannot double-count. On real volumes the same structure runs on Spark or a warehouse, partitioned by date.
Step 3: Build the online path around a latency budget
Split the budget across steps. If the caller gives the model 50 ms, the feature fetch might get 30 ms, inference 10 ms, and the rest is overhead. Each dependency gets a timeout and a fallback:
import asyncio, time
import numpy as np
from fastapi import FastAPI
from pydantic import BaseModel
FEATURE_DEFAULTS = {"orders_30d": 0.0, "avg_basket": 42.0}
app = FastAPI()
class Request(BaseModel):
user_id: int
amount: float
hour: int
@app.post("/score")
async def score(req: Request):
try:
feats = await asyncio.wait_for(fetch_features(req.user_id), timeout=0.03)
fallback = False
except TimeoutError:
feats, fallback = FEATURE_DEFAULTS, True
x = np.array([[req.amount, req.hour, feats["orders_30d"], feats["avg_basket"]]])
return {"score": float(model.predict_proba(x)[0, 1]),
"model_version": MODEL_VERSION, "feature_fallback": fallback}The response says whether a fallback was used, and the service counts fallbacks as a metric. A rising fallback rate means the model is quietly running on default values, which is a quality incident even when latency looks perfect.
Step 4: Operate each path differently
Batch needs freshness monitoring (did the job finish, is the table from today?), row counts and score distributions per run. Online needs request rate, error rate, p50, p95 and p99 latency, fallback rate, and autoscaling with warm replicas, because loading a model on a cold pod can take seconds to minutes. Both log the model version with every prediction.
Worked scenario
An online shop computed “recommended for you” lists nightly in batch for every registered user and served them from a key-value store. It worked well except for one group: people who signed up today had no row, so they saw a generic bestseller list during their most valuable first session. A team tried to fix this by moving all recommendations online, scoring a large candidate set per page view. The p50 stayed acceptable, but the p99 crossed 900 ms because each request fetched dozens of features from a database, and the page timed out for one visitor in fifty during peaks.
The design that worked was a hybrid. The nightly batch kept producing lists for known users, now with a model_version and run_date on every row. For users without a batch row, a small online model re-ranked 200 popular candidates using only request-time context (the category being browsed, device, session clicks) with no database calls, fitting comfortably in 20 ms. Every path had a fallback to the bestseller list behind a 50 ms timeout, and the fallback rate went on the dashboard. New-user conversion improved and the p99 dropped back under 100 ms.
Common mistake
- Choosing online because it sounds better. If nobody acts on the prediction until tomorrow, a real-time service is cost and on-call load with no benefit.
- Quoting average latency. The mean can look fine while one request in a hundred takes a second. SLOs are written on percentiles.
- No fallback. “The model is down” must have a defined answer: a rule, a cached score or a default decision.
- Forgetting new entities in batch designs. Precomputed scores cover only entities that existed at the last run.
- Recomputing features differently on each path. Hybrid designs invite training-serving skew unless features share one definition.
Verify the behavior
Test the fallback in CI with FastAPI’s TestClient: stub the feature fetch so one user id takes 500 ms, then assert the response arrives quickly with feature_fallback: true. In our run, a normal user returned feature_fallback: False and the slow user True, both under 60 ms. For latency, load test with realistic payloads (k6, Locust or hey) and read p99 at your expected peak, not at idle. For batch, assert row counts, uniqueness and value ranges before publishing, and alert if the newest partition is older than expected.
Follow-up questions
- What is streaming inference? Scoring events as they flow through Kafka or Flink, typically seconds after they happen, with results pushed to a store or a downstream topic. It sits between batch and request-response.
- How does dynamic batching help online GPU serving? The server groups concurrent requests into one batch for better GPU efficiency, at the cost of a small queueing delay that you cap.
- How do you choose the batch schedule? From how fast predictions go stale: evaluate yesterday’s scores against today’s outcomes and see how much accuracy each day of delay costs.
- How would you serve a model to millions of users at low cost? Precompute in batch where possible, cache online results with a short TTL, and use a smaller or quantised model for the online path.
Interview exercise
A bank wants credit-limit increase offers shown in its app. Eligibility depends on account history updated nightly and on the customer’s current balance, which changes in real time. The offer must appear when the customer opens the app, within 200 ms. Propose a serving design.
Answer and reasoning
Most of the signal is in the nightly history, so I would run a batch job that scores every eligible customer and stores a base score and the offer amount, with model version and run date. At app open, an online service looks up that precomputed row and applies a lightweight adjustment or rule using the live balance (for example, suppress the offer if the balance now exceeds a limit), which fits easily in 200 ms. If the lookup fails or the row is missing, the fallback is to show no offer, a safe default for a credit product. This keeps the expensive model in batch, where results can be reviewed before customers see them, and keeps the online path fast and simple. I would log every displayed offer with both the batch score and the live inputs, to train the next model on what was actually shown.
Continue learning
- Practise in the MLOps chapter and the MLOps MCQs.
- Keep features consistent across both paths with training-serving skew and feature stores, and size the serving pods with Kubernetes resource requests.
- Reference: Google Cloud’s MLOps continuous delivery guide and the KServe documentation for online model serving on Kubernetes.