Ch. 28 · Machine Learning

Logistic Regression Interview Questions: Sigmoid, Log Loss, Odds

How logistic regression turns a linear score into a probability, why it uses log loss, how to read odds ratios and what separation breaks.

~8 min readbeginnerupdated Oct 6, 2026

Logistic regression is the classifier interviewers reach for when they want to test fundamentals, because every part of it is a question: why is it called regression, where does the sigmoid come from, why log loss instead of squared error, what does a coefficient of 0.8 mean, and why does scikit-learn add a penalty you did not ask for? It is also still a production workhorse in credit scoring, ads and medicine, where a calibrated probability and an explainable weight matter more than the last point of accuracy.

Before you start

You should know linear regression (a weighted sum of features plus a bias) and the idea of a probability and of odds: a probability of 0.8 is odds of 0.8 / 0.2 = 4, “four to one”. It helps to have read gradient descent explained, because that is how the model is trained here. The from-scratch code is plain Python 3.14; the comparison uses scikit-learn 1.9 with NumPy 2.5. Printed outputs come from real runs.

The short answer

Logistic regression computes a linear score z = w·x + b and maps it to a probability with the sigmoid, p = 1 / (1 + e^-z). Equivalently, it models the log-odds log(p / (1 - p)) as a linear function of the features, which is why it is called regression; it becomes a classifier only when you apply a threshold. It is trained by minimising log loss (cross-entropy), which is convex for this model, so there is one global optimum but no closed-form solution. Each coefficient is interpretable: one unit of feature j multiplies the odds by e^w_j. In scikit-learn it is L2-regularized by default with C=1.0.

How it works

The sigmoid takes any real number and returns a value between 0 and 1: z = 0 gives 0.5, large positive z approaches 1 and large negative z approaches 0. Invert it and you get the log-odds, so the model’s claim is that each feature adds a fixed amount to the log-odds. The decision boundary p = 0.5 is where z = 0, a straight line (or hyperplane) in feature space; curved boundaries require engineered features such as squares or interactions.

Training maximises the likelihood of the observed labels, which is the same as minimising log loss: -[y·log(p) + (1 - y)·log(1 - p)], averaged over rows. A confident wrong prediction is punished hard, since -log(0.01) is 4.6. Its gradient with respect to each weight is remarkably clean:

import math

def sigmoid(z):
    if z >= 0:                    # split avoids overflow in math.exp for large |z|
        return 1 / (1 + math.exp(-z))
    e = math.exp(z)
    return e / (1 + e)

def log_loss_and_grad(w, b, rows, targets, l2=0.0):
    n = len(rows)
    loss, g0, g1, gb = 0.0, 0.0, 0.0, 0.0
    for (h, t), yi in zip(rows, targets):
        p = sigmoid(w[0] * h + w[1] * t + b)
        p = min(max(p, 1e-15), 1 - 1e-15)        # keep log() finite
        loss -= yi * math.log(p) + (1 - yi) * math.log(1 - p)
        g0 += (p - yi) * h                       # gradient is (p - y) * feature
        g1 += (p - yi) * t
        gb += (p - yi)
    loss = loss / n + l2 / 2 * (w[0] ** 2 + w[1] ** 2)
    return loss, [g0 / n + l2 * w[0], g1 / n + l2 * w[1]], gb / n
python

That (p - y) · x gradient is also the answer to “why not use squared error?”. Squared error on top of a sigmoid is non-convex, and its gradient carries an extra factor p(1 - p) that vanishes exactly when the model is confidently wrong. For a positive example predicted at p = 0.01, the log-loss gradient with respect to z is -0.99, a strong push; the squared-error gradient is -0.0196, about fifty times weaker.

Step-by-step walkthrough

We simulate 400 students with hours studied (0 to 10) and practice tests taken (0 to 5), and draw pass or fail from a true model with log-odds -6 + 0.9·hours + 0.5·tests. That gives 187 passes.

Step 1: Train with gradient descent

def train(rows, targets, lr=0.1, epochs=20000, l2=0.0):
    w, b = [0.0, 0.0], 0.0
    for _ in range(epochs):
        _, gw, gb = log_loss_and_grad(w, b, rows, targets, l2)
        w = [w[0] - lr * gw[0], w[1] - lr * gw[1]]
        b -= lr * gb
    return w, b

w, b = train(X, y)
print([round(v, 3) for v in w], round(b, 3), round(log_loss_and_grad(w, b, X, y)[0], 4))
# [0.832, 0.657] -6.183 0.3691
python

The estimates are near the true 0.9, 0.5 and -6; the gap is sampling noise from only 400 rows. Because log loss is convex, any reasonable start reaches the same answer.

Step 2: Read the coefficients as odds ratios

print([round(math.exp(v), 2) for v in w])      # [2.3, 1.93]
p = sigmoid(w[0] * 5 + w[1] * 2 + b)
print(round(p, 3), round(math.log(p / (1 - p)), 3))   # 0.33 -0.707
python

Each extra hour multiplies the odds of passing by 2.3, holding practice tests fixed, and each practice test by 1.93. A student with 5 hours and 2 tests has a 33% chance, log-odds -0.707. Add one hour and the odds become exactly 2.3 times larger. The probability does not rise by a fixed amount, though: the sigmoid is steep in the middle and flat at the ends, so the same odds ratio moves 0.33 much further than it moves 0.95.

Step 3: Compare with scikit-learn, with and without the penalty

import numpy as np
from sklearn.linear_model import LogisticRegression

for C in (np.inf, 1.0, 0.01):
    clf = LogisticRegression(C=C).fit(X, y)
    print(C, clf.coef_.round(3), clf.intercept_.round(3))
# inf  [[0.832 0.657]] [-6.183]
# 1.0  [[0.824 0.647]] [-6.114]
# 0.01 [[0.538 0.309]] [-3.708]
python

With C=np.inf (no penalty) scikit-learn agrees with our gradient descent to three decimals. The default C=1.0 shrinks the weights slightly, and C=0.01 shrinks them hard. C is the inverse of the regularization strength, so smaller means stronger. Since scikit-learn 1.8, the old penalty argument is deprecated; you choose L2, L1 or elastic net with l1_ratio (0, 1, or in between) and turn the penalty off with C=np.inf.

Step 4: Turn probabilities into decisions

predict applies a 0.5 threshold, but that is a business choice, not part of the model. If a false negative costs nine times a false positive, the cost-minimising threshold on a calibrated probability is 1 / (1 + 9) = 0.1. Logistic regression is usually reasonably calibrated when it is well specified and lightly regularized, which is one reason it remains popular for risk scores. How to pick the threshold from a target is covered in precision vs recall and F1.

Worked scenario

A credit team fits an unpenalised logistic regression and one coefficient comes out absurd: a binary feature account_flagged has a weight that keeps growing as they raise max_iter, and its odds ratio is reported as something like e^50. The model also scores borderline applicants with near-certain probabilities.

The cause is perfect separation: in the training sample every flagged account defaulted, so the likelihood keeps improving as the weight grows, and the optimum is at infinity. A one-feature reproduction makes the pattern obvious:

# 40 points, label is 1 exactly when x > 0
# no penalty: epoch 100 w=3.84, epoch 1,000 w=8.43, epoch 10,000 w=17.72, epoch 100,000 w=34.13
# l2=0.01:    epoch 100 w=3.11, epoch 1,000 w=3.43, epoch 10,000 w=3.43, epoch 100,000 w=3.43
python

Without a penalty the weight never settles and the loss creeps towards zero. scikit-learn with C=np.inf on the same data stops at a weight of 50.56 and reports a probability of 0.995 for x = 0.1; the number is an artefact of the solver’s stopping tolerance, not an estimate. With the default C=1.0 the weight is 2.5 and the same point scores 0.562, which is honest about how little data sits near the boundary.

The fix has two parts. Keep regularization on (tune C with LogisticRegressionCV), which guarantees finite weights. Then question the feature: a flag that predicts default perfectly is often set after the outcome, which is leakage, and the model would fail in production even with sensible weights.

Common mistake

  • “Logistic regression outputs a class.” It outputs a probability; the class comes from a threshold you choose.
  • “A weight of 0.7 increases the probability by 0.7.” It increases the log-odds by 0.7, so it multiplies the odds by about 2.
  • Comparing raw coefficient sizes across features in different units. Standardize first, or compare odds ratios per meaningful unit.
  • Forgetting that the penalty is unit-sensitive. With the default L2 penalty, an unscaled feature measured in large numbers is barely penalised while one in small numbers is shrunk hard. Put StandardScaler in the pipeline.
  • Assuming it cannot handle non-linear boundaries. It can, with polynomial features, splines or interactions; the boundary is linear in the features you give it.
  • Treating coefficients as causal. Correlated features share credit unpredictably, and the signs can flip when one is removed.

Verify the behavior

The strongest check on a from-scratch implementation is agreement with a reference solver on the unpenalised problem:

import numpy as np
from sklearn.linear_model import LogisticRegression

w, b = train(X, y)
ref = LogisticRegression(C=np.inf).fit(X, y)
assert np.allclose(w, ref.coef_[0], atol=1e-3)
assert abs(b - ref.intercept_[0]) < 1e-3
print("gradient descent matches scikit-learn")
python

If the assertion fails, first increase the epochs (an unconverged loop is the usual cause), then check the gradient with finite differences.

Follow-up questions

  • How does it handle more than two classes? Multinomial (softmax) regression, which scikit-learn uses by default for multiclass targets with lbfgs; one-vs-rest is the alternative.
  • Does it need feature scaling? Not for the unpenalised optimum, but yes in practice: penalties are unit-sensitive and solvers converge much faster on scaled features.
  • What does class imbalance do to it? Mostly shifts the intercept. Class weights change the effective base rate, so probabilities need recalibration afterwards; see imbalanced classification.
  • Why is there no closed form? Setting the gradient to zero gives equations with the sigmoid inside a sum, which cannot be solved algebraically, so iterative solvers such as L-BFGS or Newton’s method are used.

Interview exercise

A churn model has intercept -3 and a single coefficient of 0.4 on “support tickets in the last month”. What is the churn probability for a customer with 0 tickets and with 5 tickets, by what factor do the odds change between them, and why can you not say “each ticket adds a fixed amount of churn probability”?

Answer and reasoning

With 0 tickets, z = -3 and p = 1 / (1 + e^3) ≈ 0.047. With 5 tickets, z = -3 + 5 × 0.4 = -1 and p ≈ 0.269. The odds go from 0.0498 to 0.368, a factor of e^(5 × 0.4) = e^2 ≈ 7.39, which is the per-ticket odds ratio e^0.4 ≈ 1.49 applied five times. The probability change is not constant because the sigmoid is curved: the first ticket moves a customer from 0.047 to 0.069, about two points, while near the middle of the curve one ticket moves 0.731 to 0.802, about seven points. Constant effects live on the log-odds scale, which is why the coefficient is reported as an odds ratio.

Continue learning

More in Machine Learning

esc