AI & ML / Maths foundations / 02_probability_statistics.md

Probability and statistics for ML

Updated 6 interview angles 5 min read source
On this page8
  1. Distributions you should recognise on sight
  2. Expectation and variance
  3. Bayes’ theorem, and the base rate trap
  4. Central limit theorem
  5. Confidence intervals
  6. Hypothesis testing and p-values
  7. Correlation is not causation, and the ways it bites
  8. Interview angle

Probability and statistics for ML

The part of the maths that shows up in product conversations: is this A/B result real, why is your 99th percentile latency the number that matters, and why does a 99%-accurate fraud model catch almost nothing.

Distributions you should recognise on sight

Distribution Models Shows up as
Bernoulli one yes/no trial a single click, a single conversion
Binomial n independent Bernoulli trials conversions out of n visitors
Normal sums of many small effects measurement noise, weight initialisation
Log-normal products of many small effects latency, file sizes, income
Poisson count of rare events in fixed time requests per second, errors per hour
Exponential time between Poisson events inter-arrival time
Power law scale-free phenomena word frequency, popularity, degree distributions

Latency is log-normal, not normal. That’s why mean latency is a lie and you report p50/p95/p99. A senior answer to “how do you monitor an ML service” mentions percentiles unprompted. See SLO, SLI, SLA, and Error Budgets.

Expectation and variance

python
E[X]     = sum(x * p(x))          # the mean, the "centre of mass"
Var[X]   = E[(X - E[X])**2]       # spread, in squared units
Std[X]   = sqrt(Var[X])           # spread, in original units

Two properties that get used constantly:

  • Linearity of expectation: E[X + Y] = E[X] + E[Y] — always, even when X and Y are dependent. This makes many “expected number of…” puzzles trivial.
  • Variance is not linear: Var[X + Y] = Var[X] + Var[Y] only when independent. Correlated failures are why “we have three replicas so we’re 99.9%^3 safe” is wrong.

Bayes’ theorem, and the base rate trap

text
P(A|B) = P(B|A) * P(A) / P(B)

The classic interview scenario: a fraud detector with 99% sensitivity and 99% specificity, on a population where 0.1% of transactions are fraud.

python
prior_fraud      = 0.001
p_flag_if_fraud  = 0.99      # sensitivity / recall
p_flag_if_clean  = 0.01      # 1 - specificity

p_flag = p_flag_if_fraud * prior_fraud + p_flag_if_clean * (1 - prior_fraud)
precision = p_flag_if_fraud * prior_fraud / p_flag
# 0.0902 -> only ~9% of flagged transactions are actually fraud

91% of your alerts are false positives, from a model that sounds excellent. This is the single most useful piece of statistics for an ML interview: it explains why accuracy is a useless metric on imbalanced data, why precision/recall exist, and why the fraud team hates your model. Continued in Precision, recall and F1.

Central limit theorem

The mean of many independent samples is approximately normal, regardless of the underlying distribution. This is what licenses confidence intervals and t-tests on non-normal data — you’re not assuming the data is normal, you’re relying on the sample mean being normal.

The caveat that matters: it needs independence and finite variance. Heavy-tailed data (power laws) converges slowly or not at all, which is why “average revenue per user” is unstable when a few whales dominate.

Confidence intervals

python
import numpy as np

def mean_ci(sample: np.ndarray, z: float = 1.96) -> tuple[float, float]:
    """95% CI for the mean. z=1.96 for 95%, 2.576 for 99%."""
    m = sample.mean()
    se = sample.std(ddof=1) / np.sqrt(len(sample))
    return m - z * se, m + z * se

The standard error shrinks as 1/sqrt(n). To halve your error bar you need four times the data — the reason experiments take longer than product managers expect.

A 95% CI does not mean “95% probability the true value is in this interval”. It means the procedure produces intervals that contain the true value 95% of the time. Being able to state that correctly is a mild seniority signal.

Hypothesis testing and p-values

A p-value is P(data at least this extreme | null hypothesis true). It is not the probability the null is true, and not the probability your result is a fluke.

Error Meaning Controlled by
Type I (false positive) you ship a change that does nothing significance level alpha
Type II (false negative) you discard a change that worked statistical power 1 - beta

Practical failure modes worth naming in an interview:

  • Peeking. Checking the test daily and stopping when it goes significant inflates false positives badly. Fix: fixed sample size decided in advance, or a sequential-testing method built for continuous monitoring.
  • Multiple comparisons. Test 20 metrics at alpha = 0.05 and you expect one false positive by construction. Fix: Bonferroni (conservative) or Benjamini-Hochberg (controls false discovery rate).
  • Statistical vs practical significance. With ten million users, a 0.01% lift is significant and worthless. Always quote the effect size and its CI, not just the p-value.

For bootstrapping, which sidesteps most distributional assumptions:

python
def bootstrap_ci(sample, statistic=np.mean, n_boot=10_000, seed=0):
    rng = np.random.default_rng(seed)
    boots = [statistic(rng.choice(sample, size=len(sample), replace=True))
             for _ in range(n_boot)]
    return np.percentile(boots, [2.5, 97.5])

Bootstrapping is the pragmatic default when you can’t justify a parametric test — it works for medians, percentiles and weird custom metrics where no closed form exists.

Correlation is not causation, and the ways it bites

  • Confounding: ice cream sales and drownings correlate; temperature causes both.
  • Selection bias: your training data is the users who didn’t churn before you logged them.
  • Simpson’s paradox: a trend present in every subgroup reverses when the groups are pooled. It’s the reason segment-level analysis is mandatory before you believe an aggregate.

The practical version for an ML engineer: a feature that correlates with the label in training may be a consequence of it, not a cause. That’s target leakage, and it produces spectacular offline metrics and a useless production model. See Data leakage.

Interview angle 6

  • “Your fraud model is 99% accurate. Is it good?” — unanswerable without the base rate. If fraud is 0.1% of traffic, predicting “never fraud” is 99.9% accurate. Ask for precision, recall and the class balance; then do the Bayes calculation above to show what fraction of alerts are real.
  • “What’s a p-value?” — the probability of observing data at least this extreme assuming the null hypothesis is true. Not the probability the null is true, and not the probability you’re wrong.
  • “Why can’t we just check the A/B test every morning and stop when it’s significant?” — peeking. Each look is another chance to cross the threshold by luck, so the real false-positive rate is far above the nominal 5%. Fix the sample size in advance or use a sequential test designed for it.
  • “Why report p99 latency rather than the mean?” — latency is right-skewed and roughly log-normal, so the mean sits below most of the pain and is dragged around by outliers. Percentiles describe what users actually experience.
  • “How much more data do you need to halve the error bar?” — four times as much; standard error goes as 1/sqrt(n).
  • “A metric improved in every country but got worse overall. How?” — Simpson’s paradox: the country mix shifted toward a lower-performing segment. Always check segment sizes before trusting a pooled number.