ORF 526 · Princeton · Fall 2026

Probability for
Modern Machine Learning

The Gaussian The semicircle
MW 10:40–12:00 26 sessions Sep 2 – Dec 7 No measure theory

The course

A graduate probability course built out of tools and examples rather than a slow development of the foundations.

Twenty-six sessions, in two modules. Each one is organized around a technique worth owning — Lindeberg swapping, Stein’s method, Wick’s theorem, the resolvent, exponential tilting — and around an example that shows what the technique is for. A substantial part of that second half is modern machine learning: wide networks at initialization, ridge regression, the spectrum of a sample covariance, the geometry of random loss landscapes, stochastic gradient descent, and diffusion models.

No measure theory. Nothing is assumed beyond linear algebra, multivariable calculus and an undergraduate probability course. Where a standard fact from analysis is needed — Fourier inversion, the spectral theorem — it is quoted honestly and marked as such in the notes.

How it works

Homework, assigned every sessionnot graded
Weekly quiz, 15 min, in class20%
Midterm, in class30%
Final50%

The lowest quiz is dropped. Both examinations are closed book.

Grades

A90–100
B80–89
C70–79
D60–69
F0–59

On the homework and the quizzes. The exercises live inside the lecture notes, typed A, B and C: A checks that nothing was missed, B asks whether the argument was understood, and C withholds exactly one idea. They are not collected and not graded — use whatever help you like, including an AI assistant, and if you are short of time do the C problems. The weekly quiz is where that work is cashed in: closed book, by hand, drawn from the week’s assigned exercises. The homework is the exploration layer; the quiz is the verification layer, and it is the only thing that needs policing.

Module I · Gaussians 14 sessions · Sep 2 – Oct 21
1Sep 2
Gaussians: definition and linear structure
  • Definition by the Fourier transform, and why not the density
  • Linear images, existence, rotational invariance
  • Uncorrelated and jointly Gaussian implies independent
  • Maximum entropy at fixed covariance
notes
2Sep 9
Gaussians: conditioning as projection
  • Conditional expectation as orthogonal projection
  • The Schur complement, derived
  • Tower property and total variance, read off the picture
  • Bayesian posteriors and ridge regression; Tweedie's formula
notes
3Sep 14
CLT: Lindeberg's method
  • Maximum entropy, recalled; why a Gaussian at all
  • What convergence in distribution means, and why smooth test functions
  • Lindeberg's swapping argument, proved
  • Lindeberg's and Feller's conditions; the CLT for triangular arrays
  • Records: why random search improves only log n times, and why the Gaussian/Poisson dial is s_n
notes
4Sep 16
Stein's method
  • The Gaussian characterized by an identity, not a limit
  • Stein's equation; the CLT with a rate
  • Chen–Stein for Poisson; dependence
notes
5Sep 21
Wick's theorem, and a Gaussian random matrix
  • Every moment of a Gaussian vector, as a sum over pair partitions
  • Proved from the characteristic function — no density, no invertibility
  • The Gaussian Wigner matrix
  • ‖Mu‖² for a fixed direction: its mean and variance, by Wick
  • Cumulants, in the reading
notes
6Sep 23
The semicircle law by counting paths
  • The theorem, and the strategy: moments are traces
  • E tr M² and E tr M⁴ by hand
  • Matrix powers as sums over paths; from a path to a pair partition
  • Only tree paths survive; non-crossing pairings and the Catalan numbers
  • The fluctuations are small
notes
7Sep 28
The Stieltjes transform
  • The resolvent and mN(z) = (1/N) tr (M − zI)−1
  • Eigenvalues are the poles: counting them in an interval by a contour integral
  • What to compute: limε→0 limN→∞ (1/π) Im mN(x + iε), and why the limits come in that order
  • Near z = ∞ the transform is the generating function of the traces
notes
8Sep 30
The semicircle law by the Stieltjes transform
  • Schur complement: one diagonal entry of a resolvent through a smaller resolvent
  • The self-consistent equation m² + zm + 1 = 0, with explicit error bounds
  • The trace concentrates: a Poincaré inequality and a cancellation of N's
  • Stability of the equation, and the density read off the unit circle
notes
9Oct 5
The Weyl chamber formula and the log gas
  • The exact joint density of the eigenvalues
  • The Vandermonde determinant as a repulsion between levels
  • Eigenvalues as a gas of charges in a confining potential
  • The semicircle law as the equilibrium measure
10Oct 7
Gaussian processes: the geometric picture
  • A process on an index set T as a curve in a Hilbert space
  • Existence from any positive semidefinite kernel
  • The Karhunen–Loève expansion
11Oct 12
Gaussian processes: the canonical metric and maxima
  • The canonical pseudometric is distance in the Hilbert space
  • Suprema as support functions; Gaussian width
  • Sudakov and Dudley, stated
12Oct 14
Kac–Rice and zero sets
  • The area formula, invoked without proof
  • Rice's formula for stationary processes
  • Random trigonometric polynomials; Kac polynomials
13Oct 19
Random landscapes: counting critical points
  • Conditioned on ∇f = 0, the Hessian is a GOE matrix
  • log|det| as a spectral integral
  • Exponentially many, overwhelmingly saddles
14Oct 21
Midterm examination
  • Sessions 1–13, closed book
Module II · Dynamics, Deviations, Diffusion 12 sessions · Oct 26 – Dec 7
1Oct 26
Martingales: conditional expectation and definitions
  • Conditional expectation as projection
  • Filtrations and the tower property
  • The Doob decomposition
2Oct 28
Martingales: limit theorems
  • Doob's maximal inequality; convergence in L²
  • Optional stopping
  • The martingale central limit theorem
3Nov 2
Martingales: concentration
  • Azuma–Hoeffding and Freedman
  • Bounded differences
4Nov 4
SGD: stochastic approximation
  • Drift plus martingale difference
  • Robbins–Monro and the ODE method
5Nov 9
SGD: fluctuations and averaging
  • Asymptotic normality of the iterates
  • Polyak–Ruppert averaging
  • Concentration along the trajectory
6Nov 11
SGD: high-dimensional dynamics
  • A deterministic ODE plus a fluctuation term
  • The data covariance spectrum, through the resolvent
7Nov 16
Large deviations: energy versus entropy
  • Cramér's theorem by exponential tilting
  • The rate function as a competition
8Nov 18
Large deviations: Sanov's theorem
  • The method of types
  • Relative entropy as the rate function
  • The contraction principle
9Nov 23
Large deviations: Varadhan and examples
  • Varadhan's lemma and the Gibbs variational principle
  • Curie–Weiss and its phase transition
  • Hypothesis testing; escape from a basin
10Nov 30
Diffusion models: the forward process
  • Noising as a Gaussian channel
  • Tweedie's formula from integration by parts
  • Denoising as score estimation
11Dec 2
Diffusion models: score matching
  • The reverse chain
  • Denoising score matching
  • The DDPM objective
12Dec 7
Diffusion models: the PDE picture
  • Fokker–Planck and the heat equation
  • The Gaussian as fundamental solution
  • Time reversal; the variational bound