Measure Theory Lesson 31: Weak Convergence of Measures

Introduction

Up to this point, we have studied convergence of functions:

  • Pointwise convergence
  • Uniform convergence
  • Almost everywhere convergence
  • Convergence in measure

Today we ask a different question:

What does it mean for measures themselves to converge?

For example, suppose:

$$\mu_1,\mu_2,\mu_3,\ldots$$

is a sequence of probability measures.

Can we say:

$$\mu_n\to\mu$$

for some measure:

$$\mu$$

If so, what should this mean?

This question turns out to be one of the central problems in probability theory and modern analysis.

The answer leads to Weak Convergence of Measures.

Weak convergence is fundamental in:

  • Probability theory
  • Bayesian statistics
  • Machine learning
  • Stochastic processes
  • Functional analysis

and later:

  • Operator algebras
  • Noncommutative geometry

Motivation

Suppose:

$$X_n\sim\mu_n$$

and:

$$X\sim\mu$$

where:

$$\sim$$

means “has distribution.”

Imagine that the distributions of:

$$X_n$$

become increasingly similar to the distribution of:

$$X$$

How should we formalize this?

Comparing measures directly is difficult.

Instead, we compare their effects on functions.


A Key Idea

Recall the Riesz Representation Theorem.

A measure is determined by its integrals.

If we know:

$$\int f,d\mu$$

for all suitable functions:

$$f$$

then we know the measure.

This suggests a definition.

Instead of comparing measures directly,

compare the integrals they produce.


Definition of Weak Convergence

Let:

$$\mu_n$$

and:

$$\mu$$

be Radon measures on a topological space:

$$X$$

We say:

$$\mu_n \Rightarrow \mu$$

and read:

μₙ converges weakly to μ

if:

$$\int_X f,d\mu_n \to \int_X f,d\mu$$

for every bounded continuous function:

$$f:X\to\mathbb R$$


Why Continuous Functions?

Continuous functions are sensitive enough to detect geometric information.

At the same time, they are regular enough to avoid pathological behavior.

They provide exactly the right testing class.


Intuition

Weak convergence means:

Every continuous measurement produces nearly the same result for μₙ and μ when n is large.

Instead of comparing sets,

we compare averages of continuous functions.


Example 1: Convergence of Point Masses

Consider:

$$\mu_n=\delta_{1/n}$$

and:

$$\mu=\delta_0$$

Does:

$$\mu_n\Rightarrow\mu$$

hold?

Take any bounded continuous function:

$$f$$

Then:

$$\int f,d\delta_{1/n}=f(1/n)$$

Since:

$$f$$

is continuous,

$$f(1/n)\to f(0)$$

Thus:

$$\int f,d\delta_{1/n}\to\int f,d\delta_0$$

Therefore:

$$\delta_{1/n}\Rightarrow\delta_0$$


Why This Makes Sense

The point masses are moving closer and closer to:

$$0$$

Thus the measures themselves should converge toward:

$$\delta_0$$

Weak convergence captures exactly this intuition.


Example 2: Normal Approximations

Consider:

$$N(0,\sigma_n^2)$$

with:

$$\sigma_n\to0$$

The distributions become increasingly concentrated near:

$$0$$

In fact:

$$N(0,\sigma_n^2)\Rightarrow\delta_0$$

This is a fundamental example in probability.


Example 3: Empirical Measures

Suppose:

$$x_1,\ldots,x_n$$

are observations.

Define:

$$\mu_n=\frac1n\sum_{k=1}^{n}\delta_{x_k}$$

This is called the empirical measure.

One of the central goals of statistics is to show:

$$\mu_n\Rightarrow\mu$$

where:

$$\mu$$

is the true population distribution.

Thus weak convergence lies at the heart of statistical inference.


Why Not Use Set Convergence?

One might attempt:

$$\mu_n(A)\to\mu(A)$$

for every measurable set:

$$A$$

Unfortunately this is often too strong.

Weak convergence succeeds because it uses continuous test functions instead of arbitrary sets.


Portmanteau Theorem

One of the most important results in probability theory is the Portmanteau Theorem.

It states that many apparently different definitions of weak convergence are actually equivalent.

For probability measures:

$$\mu_n\Rightarrow\mu$$

if and only if various conditions hold.

For example:

For every bounded continuous:

$$f$$

$$\int f,d\mu_n\to\int f,d\mu$$

or equivalently,

for every closed set:

$$F$$

$$\limsup_{n\to\infty}\mu_n(F)\le\mu(F)$$

or equivalently,

for every open set:

$$U$$

$$\liminf_{n\to\infty}\mu_n(U)\ge\mu(U)$$

These equivalences are extraordinarily useful.


Weak Convergence vs Pointwise Convergence

Measures do not converge pointwise the way functions do.

Instead, convergence is detected through integrals.

This is one of the first examples where:

Objects are understood through their action on test functions.

This viewpoint becomes increasingly important later.


Weak Convergence and Probability

Suppose:

$$X_n\sim\mu_n$$

and:

$$X\sim\mu$$

Then:

$$X_n \Rightarrow X$$

means:

$$E[f(X_n)]\to E[f(X)]$$

for every bounded continuous:

$$f$$

This is called convergence in distribution.

Thus:

Weak convergence of measures = convergence in distribution of random variables.

This is one of the most important facts in probability theory.


Example: Central Limit Theorem

The Central Limit Theorem states:

$$\frac{\overline X_n-\mu}{\sigma/\sqrt n}\Rightarrow N(0,1)$$

This symbol:

$$\Rightarrow$$

means weak convergence.

Therefore the Central Limit Theorem is fundamentally a theorem about weak convergence of measures.


Why Analysts Care

Weak convergence is powerful because it allows complicated objects to be approximated by simpler ones.

Examples:

  • Approximating distributions
  • Approximating solutions to PDEs
  • Approximating measures on manifolds
  • Studying stochastic processes

Many existence proofs rely on weak convergence.


The Topological Viewpoint

Notice something remarkable.

A measure is not converging pointwise.

Instead, convergence is defined through continuous functions.

This means the topology of the underlying space becomes crucial.

Weak convergence is one of the first major interactions between:

  • topology
  • measure theory
  • functional analysis

Why Radon Measures Were Needed

In the previous lesson we introduced Radon measures.

Weak convergence behaves especially well for Radon measures because:

  • compact sets control the measure
  • continuous functions interact naturally with topology

Without Radon measures, the theory becomes much more complicated.


Connection to Bayesian Statistics

Suppose:

$$\Pi_n$$

is a sequence of posterior distributions.

A common goal is to prove:

$$\Pi_n \Rightarrow \delta_{\theta_0}$$

This means the posterior concentrates around the true parameter.

Thus weak convergence is one of the mathematical foundations of Bayesian consistency.


Connection to Functional Analysis

The definition:

$$\int f,d\mu_n\to\int f,d\mu$$

looks very familiar.

We are studying convergence through linear functionals.

This is precisely the viewpoint of functional analysis.

Later we will see that weak convergence of measures is closely related to weak-* convergence in dual spaces.


Connection to Alain Connes

One of the central themes of modern analysis is:

Study an object through how it acts on test objects.

Weak convergence studies measures through their action on continuous functions.

Later:

  • Operators are studied through their action on vectors.
  • States are studied through their action on algebras.
  • Noncommutative spaces are studied through their representations.

This philosophy is everywhere in Connes’ work.

Weak convergence is one of the earliest examples of this perspective.


Key Concepts Learned

By the end of this lesson you should understand:

  • Weak convergence is written:

$$\mu_n\Rightarrow\mu$$

  • It means:

$$\int f,d\mu_n\to\int f,d\mu$$

for every bounded continuous function:

$$f$$

  • Weak convergence compares measures through integrals.
  • Convergence in distribution is weak convergence of probability measures.
  • The Central Limit Theorem is a weak convergence theorem.
  • Empirical measures converge weakly under suitable conditions.
  • Weak convergence connects measure theory, topology, probability, and functional analysis.

Looking Ahead

Measure Theory Lesson 32: Tightness and Prokhorov’s Theorem

Weak convergence is powerful, but it raises a crucial question:

How do we know a sequence of measures has a convergent subsequence at all?

The answer involves tightness, one of the most important compactness concepts in probability theory. Prokhorov’s Theorem provides the bridge between tightness and weak convergence and is one of the foundational results underlying modern probability, Bayesian asymptotics, and stochastic processes.

Leave a Reply

Discover more from nerd-ish

Subscribe now to keep reading and get access to the full archive.

Continue reading