AI & ML / Classical ML / 02_logistic_regression.md

Logistic regression

Updated 6 interview angles 5 min read source
On this page10
  1. The model
  2. Why not MSE
  3. Interpreting coefficients
  4. The threshold is not part of the model
  5. Class imbalance
  6. Multi-class
  7. Calibration
  8. Separation: the failure mode nobody warns you about
  9. Why it survives
  10. Interview angle

Logistic regression

The default baseline for classification, and the model most likely to still be in production five years later because someone has to explain it to a regulator.

The model

Linear score, squashed into a probability:

text
z = w . x + b
p = sigmoid(z) = 1 / (1 + exp(-z))
python
from sklearn.linear_model import LogisticRegression

model = LogisticRegression(max_iter=1000, C=1.0).fit(X_train, y_train)
# use this, not .predict()
proba = model.predict_proba(X_test)[:, 1]

Note C is the inverse regularisation strength — smaller C means stronger regularisation. It’s the opposite of alpha in Ridge and Lasso, and it catches people out constantly.

Why not MSE

Two reasons, and the second is the better answer:

  1. Squared error on a sigmoid output is non-convex in the weights, so optimisation can get stuck.
  2. The gradient vanishes exactly when you need it most. With MSE, the gradient contains the sigmoid derivative p(1-p), which goes to zero when the model is confidently wrong. Cross-entropy’s gradient is simply (p - y) * x — the sigmoid derivative cancels — so a confident wrong prediction produces a large corrective step.
text
Cross-entropy:  L = -[y*log(p) + (1-y)*log(1-p)]
Gradient:       dL/dw = (p - y) * x        # clean, and doesn't saturate

That cancellation is a favourite follow-up question. Being able to state it is a clear signal.

Interpreting coefficients

w_i is the change in log-odds per unit change in feature i. Exponentiate for an odds ratio:

python
import numpy as np
odds_ratio = np.exp(model.coef_[0])
# 1.5 -> a one-unit increase multiplies the odds by 1.5 (+50%)

Coefficients are only comparable if features are scaled. And odds ratios are not probability ratios — a common misstatement. Doubling the odds does not double the probability.

The threshold is not part of the model

predict() uses 0.5. That default is almost never the business-optimal choice, and treating it as fixed is a junior mistake.

python
from sklearn.metrics import precision_recall_curve
import numpy as np

precision, recall, thresholds = precision_recall_curve(y_val, proba)
f1 = 2 * precision * recall / (precision + recall + 1e-12)
best_threshold = thresholds[np.argmax(f1[:-1])]

Better still, choose the threshold from the cost matrix rather than from F1. If a missed fraud costs $500 and a false alarm costs $5 of review time, the optimal threshold follows from that ratio, not from a symmetric metric.

Train the model once; tune the threshold separately on validation data. They’re independent decisions.

Class imbalance

python
# weights inversely to class frequency
LogisticRegression(class_weight="balanced")

class_weight="balanced" re-weights the loss so the minority class matters more. It’s usually the first thing to try, and it’s cheaper and safer than resampling.

Important consequence: re-weighting distorts the predicted probabilities. The ranking stays useful, but predict_proba no longer reflects the true base rate. If you need calibrated probabilities — for expected-value calculations, say — either don’t re-weight, or recalibrate afterwards.

Resampling (SMOTE, undersampling) is the alternative, and must happen strictly inside the training fold. See Train / validation / test splits.

Multi-class

Strategy How Note
Multinomial / softmax one model, softmax over classes the default in modern sklearn; usually better
One-vs-rest N binary models probabilities don’t sum to 1 without normalisation
One-vs-one N(N-1)/2 models expensive; rarely worth it

For multi-label (any subset can apply), use independent binary classifiers, not softmax — softmax forces mutual exclusivity.

Calibration

Logistic regression is typically well-calibrated out of the box, which is one of its underrated advantages: a predicted 0.7 really does mean roughly 70% of such cases are positive. Tree ensembles and SVMs generally are not.

python
from sklearn.calibration import CalibratedClassifierCV, calibration_curve

calibrated = CalibratedClassifierCV(base_model, method="isotonic", cv=5)

Use sigmoid (Platt scaling) for small data, isotonic when you have plenty. Check with a reliability diagram before assuming you need it.

Calibration matters whenever the probability feeds a downstream calculation — expected loss, bid pricing, triage ordering. If you only threshold or rank, it doesn’t.

Separation: the failure mode nobody warns you about

If a feature perfectly separates the classes, the likelihood is maximised by pushing that coefficient to infinity. You’ll see a convergence warning and an absurd coefficient.

Causes are almost always leakage — a feature that encodes the label. Regularisation masks the symptom by bounding the coefficient, which is why a strong L2 penalty makes the warning go away without fixing anything. Investigate the feature before you silence the warning.

Why it survives

  • Calibrated probabilities by default.
  • Explainable at the coefficient level, in a form regulators accept.
  • Microsecond inference, kilobytes of model.
  • Convex — one optimum, reproducible runs.
  • A real baseline. If gradient boosting beats it by one point, you’ve learned the problem is mostly linear.

Where it loses: it can’t capture interactions or non-linearities unless you engineer them explicitly. That’s exactly where gradient boosting wins on tabular data — see Gradient boosting.

Interview angle 6

  • “Why is it called regression if it classifies?” — it regresses the log-odds, which is a continuous quantity; the classification comes from thresholding the resulting probability.
  • “Why cross-entropy rather than MSE?” — MSE is non-convex through the sigmoid, and its gradient contains p(1-p), which vanishes precisely when the model is confidently wrong. Cross-entropy’s gradient reduces to (p - y)x, so wrong-and-confident produces a strong correction.
  • “What does a coefficient of 0.7 mean?” — a one-unit increase in that feature raises the log-odds by 0.7, multiplying the odds by exp(0.7) ≈ 2. Only interpretable if features are scaled, and odds are not probabilities.
  • “Your classes are 99:1. What do you change?” — not accuracy as the metric (use PR-AUC), class_weight="balanced" or resampling inside the fold, and a threshold chosen from the cost matrix rather than left at 0.5. Note that re-weighting breaks probability calibration.
  • “When do you need calibration?” — when the probability itself is consumed downstream, such as expected-value or pricing decisions. Logistic regression is usually fine as-is; tree ensembles usually are not.
  • “You get a convergence warning and a coefficient of 40. What happened?” — likely perfect separation, which almost always means a leaked feature. Regularisation will hide it; find the feature instead.