pencils ready ✎

Machine Learning MCQs multiple-choice questions with answers & explanations

All 20 Machine Learning quiz questions on one page. Pick an answer in your head, then open Show answer to check it and read why. Want a score and a timer? Take them as a quiz instead.

20 questions
  1. 1.

    What does this print?

    easy
    from sklearn.metrics import precision_score, recall_score, f1_score, accuracy_score
    y_true = [1, 1, 1, 1, 0, 0, 0, 0, 0, 0]
    y_pred = [1, 1, 0, 0, 1, 0, 0, 0, 0, 0]
    print(precision_score(y_true, y_pred), recall_score(y_true, y_pred),
          round(f1_score(y_true, y_pred), 3), accuracy_score(y_true, y_pred))
    1. A0.5 0.666 0.571 0.7
    2. B0.6666666666666666 0.5 0.571 0.7
    3. C0.6666666666666666 0.5 0.583 0.7
    4. D0.5 0.5 0.5 0.8
    Show answer

    Answer: B (0.6666666666666666 0.5 0.571 0.7)

    There are 2 true positives, 1 false positive and 2 false negatives. Precision is 2 / 3, recall is 2 / 4, and F1 is the harmonic mean 2PR / (P + R) = 0.571, not the arithmetic mean 0.583. Accuracy counts the 7 correct predictions out of 10.

  2. 2.

    A dataset has 10 positives and 990 negatives. What does this baseline print?

    easy
    from sklearn.dummy import DummyClassifier
    from sklearn.metrics import accuracy_score, recall_score
    X = [[0]] * 1000
    y = [1] * 10 + [0] * 990
    pred = DummyClassifier(strategy="most_frequent").fit(X, y).predict(X)
    print(accuracy_score(y, pred), recall_score(y, pred))
    1. A0.5 0.5
    2. B0.99 0.0
    3. C0.99 0.99
    4. D0.01 1.0
    Show answer

    Answer: B (0.99 0.0)

    Always predicting the majority class is right 990 times out of 1,000, so accuracy is 0.99, yet it catches none of the positives, so recall is 0. This is why accuracy is the wrong headline metric for imbalanced problems.

  3. 3.

    The labels are pure random noise. What training accuracy does this print?

    mid
    import numpy as np
    from sklearn.tree import DecisionTreeClassifier
    rng = np.random.default_rng(0)
    X = rng.normal(size=(500, 5))
    y = rng.integers(0, 2, size=500)
    tree = DecisionTreeClassifier(random_state=0).fit(X, y)
    print(tree.score(X, y))
    1. AAbout 0.5
    2. B1.0
    3. CAbout 0.75
    4. DIt raises an error because the labels have no signal
    Show answer

    Answer: B (1.0)

    With max_depth=None the tree keeps splitting until every leaf is pure, and with continuous features no two rows collide, so it memorises all 500 labels. Training accuracy says nothing about generalization; on new data this tree would score about 0.5.

  4. 4.

    What does this print?

    mid
    from sklearn.metrics import roc_auc_score
    y = [0, 0, 1, 1, 0, 1]
    s = [0.1, 0.4, 0.35, 0.8, 0.2, 0.7]
    print(round(roc_auc_score(y, s), 3),
          round(roc_auc_score(y, [v ** 3 for v in s]), 3),
          round(roc_auc_score(y, [-v for v in s]), 3))
    1. A0.889 0.889 0.111
    2. B0.889 0.704 0.111
    3. C0.889 0.889 0.889
    4. D0.889 0.704 -0.889
    Show answer

    Answer: A (0.889 0.889 0.111)

    ROC-AUC depends only on how scores rank positives against negatives: 8 of the 9 positive-negative pairs are ordered correctly. Cubing is monotonic, so the ranking and the AUC are unchanged; negating reverses every pair, giving 1 - 0.889.

  5. 5.

    A binary classifier always predicts probability 0.5. What log loss does it get?

    mid
    from sklearn.metrics import log_loss
    y = [0, 1, 1, 0, 1, 0, 1, 1]
    print(round(log_loss(y, [0.5] * len(y)), 3))
    1. A0.5
    2. B0.0
    3. C0.693
    4. D1.0
    Show answer

    Answer: C (0.693)

    Each prediction costs -ln(0.5) = ln 2, about 0.693, whatever the label, so the average is 0.693. It is the natural baseline for binary log loss; a useful model must score below it.

  6. 6.

    One prediction is badly off. What does this print?

    easy
    from sklearn.metrics import mean_absolute_error, root_mean_squared_error
    y_true = [10, 10, 10, 10, 10]
    y_pred = [11, 9, 11, 9, 30]
    print(mean_absolute_error(y_true, y_pred), round(root_mean_squared_error(y_true, y_pred), 2))
    1. A4.8 4.8
    2. B4.8 8.99
    3. C8.99 4.8
    4. D4.8 80.8
    Show answer

    Answer: B (4.8 8.99)

    MAE averages absolute errors: (1 + 1 + 1 + 1 + 20) / 5 = 4.8. RMSE squares them first, so the single error of 20 dominates: sqrt(404 / 5) is about 8.99. A large gap between RMSE and MAE points to a few big misses; 80.8 is the MSE before the square root.

  7. 7.

    Age is in years and income in dollars. What explained variance ratio does PCA report?

    mid
    import numpy as np
    from sklearn.decomposition import PCA
    rng = np.random.default_rng(0)
    age = rng.normal(40, 10, 1000)
    income = rng.normal(60000, 15000, 1000)
    X = np.column_stack([age, income])
    print(PCA().fit(X).explained_variance_ratio_.round(4))
    1. A[0.5 0.5]
    2. B[1. 0.]
    3. C[0.6 0.4]
    4. D[0. 1.]
    Show answer

    Answer: B ([1. 0.])

    PCA maximises variance, and income has a variance about two million times larger than age only because of its units. The first component is essentially the income axis and explains almost all the variance. Standardize the features first if each one should count.

  8. 8.

    Gradient descent minimises x ** 2 with learning rate 1.1. What happens to x?

    easy
    x, lr = 1.0, 1.1
    for step in range(4):
        grad = 2 * x
        x = x - lr * grad
        print(round(x, 3))
    1. AIt converges to 0: prints 0.1, 0.01, ...
    2. BIt flips sign and grows: -1.2, 1.44, -1.728, 2.074
    3. CIt stays at 1.0
    4. DIt oscillates between -1 and 1
    Show answer

    Answer: B (It flips sign and grows: -1.2, 1.44, -1.728, 2.074)

    Each step computes x - 1.1 * 2x = -1.2x, so the iterate overshoots the minimum and grows by 20% every step: divergence. For this function any learning rate above 1.0 diverges; a rate of 1.0 bounces between 1 and -1 forever.

  9. 9.

    Here y is sorted: 50 zeros followed by 50 ones. Which splitter does cv=5 use?

    hard
    from sklearn.linear_model import LogisticRegression
    from sklearn.model_selection import cross_val_score
    scores = cross_val_score(LogisticRegression(), X, y, cv=5)
    1. AKFold(5) without shuffling, so some folds contain one class only
    2. BStratifiedKFold(5) without shuffling
    3. CKFold(5, shuffle=True)
    4. DShuffleSplit(5)
    Show answer

    Answer: B (StratifiedKFold(5) without shuffling)

    For a classifier with a binary or multiclass target, an integer cv becomes StratifiedKFold, so every fold keeps the 50/50 class ratio even though y is sorted. It does not shuffle by default. For a regressor the same cv=5 would be a plain KFold.

  10. 10.

    What regularization does LogisticRegression() from scikit-learn apply with default arguments?

    mid
    1. ANone: it is plain maximum likelihood
    2. BL2 with C=1.0
    3. CL1 with C=1.0
    4. DL2 with alpha=1.0
    Show answer

    Answer: B (L2 with C=1.0)

    scikit-learn regularizes logistic regression by default: an L2 penalty with inverse strength C=1.0, where smaller C means stronger regularization. Since 1.8 the penalty argument is deprecated in favour of l1_ratio (0 for L2, 1 for L1) and C=np.inf for no penalty.

  11. 11.

    Only the first two of ten features matter. How many coefficients are exactly zero for each model?

    mid
    import numpy as np
    from sklearn.linear_model import Lasso, Ridge
    rng = np.random.default_rng(0)
    X = rng.normal(size=(200, 10))
    y = 3 * X[:, 0] - 2 * X[:, 1] + rng.normal(scale=0.5, size=200)
    print(np.sum(Lasso(alpha=0.1).fit(X, y).coef_ == 0),
          np.sum(Ridge(alpha=10).fit(X, y).coef_ == 0))
    1. A0 0
    2. B8 0
    3. C8 8
    4. D0 8
    Show answer

    Answer: B (8 0)

    The L1 penalty has a constant gradient, so weights on the eight irrelevant features are pushed to exactly zero. Ridge shrinks them toward zero but leaves small non-zero values, which is why L1 doubles as feature selection and L2 does not.

  12. 12.

    What do the two lines print?

    easy
    import numpy as np
    from sklearn.preprocessing import StandardScaler
    X_train = np.array([[1.0], [2.0], [3.0]])
    X_new = np.array([[5.0]])
    print(StandardScaler().fit(X_train).transform(X_new).round(3))
    print(StandardScaler().fit_transform(X_new))
    1. A[[3.674]] then [[0.]]
    2. B[[3.674]] twice
    3. C[[0.]] twice
    4. D[[5.]] then [[0.]]
    Show answer

    Answer: A ([[3.674]] then [[0.]])

    Fitted on the training data (mean 2, standard deviation 0.816), the new value becomes (5 - 2) / 0.816 = 3.674. Calling fit_transform on the new data alone re-centres it on itself and erases the information. Fit preprocessing on training data and only transform everything else.

  13. 13.

    What is the last line this prints?

    mid
    from sklearn.model_selection import TimeSeriesSplit
    X = list(range(6))
    for train, test in TimeSeriesSplit(n_splits=3).split(X):
        print(train.tolist(), test.tolist())
    1. A[0, 1, 2, 3, 4] [5]
    2. B[1, 2, 3, 4, 5] [0]
    3. C[0, 1, 2] [3, 4, 5]
    4. D[3, 4, 5] [0, 1, 2]
    Show answer

    Answer: A ([0, 1, 2, 3, 4] [5])

    Each split trains on everything before the test block and tests on the next block, so the training window grows: [0, 1, 2] [3], then [0, 1, 2, 3] [4], then [0, 1, 2, 3, 4] [5]. The model never sees the future, unlike shuffled k-fold.

  14. 14.

    Labels are random. What training accuracies does this print for k = 1 and k = 50?

    mid
    import numpy as np
    from sklearn.neighbors import KNeighborsClassifier
    rng = np.random.default_rng(0)
    X = rng.normal(size=(300, 2))
    y = rng.integers(0, 2, size=300)
    for k in (1, 50):
        print(k, round(KNeighborsClassifier(n_neighbors=k).fit(X, y).score(X, y), 2))
    1. A1 0.5 and 50 0.5
    2. B1 1.0 and 50 0.54
    3. C1 0.54 and 50 1.0
    4. D1 1.0 and 50 1.0
    Show answer

    Answer: B (1 1.0 and 50 0.54)

    With k = 1 every training point is its own nearest neighbour, so training accuracy is perfect even on noise: maximum variance. With k = 50 the vote averages over many neighbours and the score falls to near chance, which is the honest answer here. Small k means low bias and high variance.

  15. 15.

    row_id is unique per row and carries no signal. What do the two lines show?

    hard
    rf = RandomForestClassifier(random_state=0).fit(X_tr, y_tr)   # columns: signal, row_id
    print(rf.feature_importances_.round(2))
    print(permutation_importance(rf, X_te, y_te, random_state=0).importances_mean.round(2))
    # [0.68 0.32]
    # [ 0.2  -0.01]
    1. ABoth methods agree that row_id matters
    2. BImpurity importance credits row_id for overfit splits; permutation importance on test data shows it is useless
    3. CPermutation importance is broken because it can be negative
    4. Drow_id must be useful because it got 32% of the importance
    Show answer

    Answer: B (Impurity importance credits row_id for overfit splits; permutation importance on test data shows it is useless)

    Impurity-based importance is measured on training data and favours features with many unique values, which deep trees use to memorise rows. Permutation importance on held-out data measures the real drop in score when the feature is shuffled: about zero for row_id. Small negative values are just noise.

  16. 16.

    What happens to k-means inertia (within-cluster sum of squares) as k grows from 1 to 8 on the same data?

    easy
    print([round(KMeans(n_clusters=k, n_init=10, random_state=0).fit(X).inertia_)
           for k in (1, 2, 4, 8)])
    # 200 points of 2-D Gaussian noise
    1. AIt rises, so the best k has the highest inertia
    2. BIt falls steadily: [395, 262, 136, 71]
    3. CIt stays roughly constant
    4. DIt is lowest at k = 2 and rises after
    Show answer

    Answer: B (It falls steadily: [395, 262, 136, 71])

    More centroids can only bring points closer to their nearest centre, so inertia keeps falling and reaches 0 when k equals the number of points. That is why you cannot pick k by minimising inertia; look for an elbow, use the silhouette score or judge usefulness.

  17. 17.

    Both models train on x from 0 to 9 with y = 2x. What do they predict for x = 20?

    hard
    import numpy as np
    from sklearn.linear_model import LinearRegression
    from sklearn.ensemble import RandomForestRegressor
    X = np.arange(0, 10).reshape(-1, 1).astype(float)
    y = 2 * X.ravel()
    lin = LinearRegression().fit(X, y)
    rf = RandomForestRegressor(random_state=0).fit(X, y)
    print(lin.predict([[20]]).round(1), rf.predict([[20]]).round(1))
    1. A[40.] [40.]
    2. B[40.] [17.]
    3. C[18.] [18.]
    4. D[40.] [0.]
    Show answer

    Answer: B ([40.] [17.])

    A tree predicts the average target of a leaf, so for any input beyond the training range the forest returns values close to the largest targets it saw (the bootstrap averages land near 17 here). Linear regression extends the fitted line to 40. Tree ensembles cannot extrapolate trends.

  18. 18.

    The 100 labels are random and there are 5,000 noise features. Roughly what does the first line print?

    hard
    X_sel = SelectKBest(f_classif, k=20).fit_transform(X, y)
    print(round(cross_val_score(LogisticRegression(), X_sel, y, cv=5).mean(), 2))
    
    pipe = make_pipeline(SelectKBest(f_classif, k=20), LogisticRegression())
    print(round(cross_val_score(pipe, X, y, cv=5).mean(), 2))
    1. AAbout 0.5, the same as the second line
    2. BAbout 0.85, while the second line prints about 0.5
    3. CAbout 1.0 for both
    4. DAbout 0.5, while the second line prints about 0.85
    Show answer

    Answer: B (About 0.85, while the second line prints about 0.5)

    Selecting features on all 100 rows picks the 20 noise columns that happen to correlate with every label, including the labels of the future validation folds, so cross-validation reports about 0.85 on pure noise. Inside the pipeline, selection is refitted on each training fold and the score drops to chance (0.48 in this run).

  19. 19.

    In a soft-margin SVM, what does increasing C from 0.01 to 100 usually do?

    mid
    1. AWidens the margin and increases bias
    2. BNarrows the margin, tolerates fewer training violations and increases variance
    3. CChanges the kernel from linear to RBF
    4. DHas no effect when the data is not linearly separable
    Show answer

    Answer: B (Narrows the margin, tolerates fewer training violations and increases variance)

    C is the penalty on margin violations. A large C makes violations expensive, so the optimiser accepts a narrower margin that fits the training points closely, lowering bias and raising variance. A small C is stronger regularization: a wider margin with more points inside it.

  20. 20.

    Columns are counts of the words "free" and "meeting"; 1 means spam. What spam probability is printed for a message with one "free"?

    mid
    import numpy as np
    from sklearn.naive_bayes import MultinomialNB
    X = np.array([[2, 0], [3, 0], [0, 2], [0, 3]])
    y = np.array([1, 1, 0, 0])
    nb = MultinomialNB(alpha=1.0).fit(X, y)
    print(nb.predict_proba([[1, 0]])[0, 1].round(3))
    1. A1.0
    2. B0.857
    3. C0.5
    4. D0.75
    Show answer

    Answer: B (0.857)

    With Laplace smoothing, P(free given spam) = (5 + 1) / (5 + 2) = 6/7 and P(free given ham) = (0 + 1) / (5 + 2) = 1/7. The priors are equal, so the posterior is 6/7 = 0.857. Without smoothing, ham would get probability 0 for any message containing "free" and the output would be 1.0.

esc