Sum of random variables, Gaussian limits, Lévy limits

A first course in probability often gives us one very powerful mental picture: when many independent random variables are added, the answer should begin to look Gaussian. This is the central limit theorem. In its most familiar form, it says that if X_1,X_2,\ldots are independent and identically distributed, with mean \mathbb E X_j=\mu and variance \mathrm{Var}(X_j)=\sigma^2<\infty , then

\displaystyle \frac{X_1+\cdots+X_n-n\mu}{\sigma\sqrt n}\Rightarrow N(0,1).

This is one of the main organizing facts of probability. But the slogan “sums become Gaussian” is too crude. A sum becomes Gaussian only when we look at it on a scale where no single summand remains visible, while the total quadratic fluctuation remains visible. If that balance fails, other limits can appear. Some limits remember a finite number of rare events. Some remember jumps of many different sizes. Some have infinite variance. Some even have infinite mean. The point of this article is to understand, concretely and through calculation, how these possibilities arise.

The surprising part is that all of this can happen even when every finite-stage random variable is bounded. Suppose that for each n we have independent random variables \displaystyle X_{n,1},X_{n,2},\ldots,X_{n,n}, and suppose that \displaystyle |X_{n,j}|\le 1 for every n and every j . At no finite stage do we use an unbounded random variable. Nevertheless, the sums

\displaystyle S_n=X_{n,1}+\cdots+X_{n,n}

can have Gaussian limits, Poisson limits, compound Poisson limits, truncated stable limits, full stable limits, and positive stable limits with infinite mean. The reason is that the distribution of X_{n,j} is allowed to depend on n . This is a triangular array, not a single fixed distribution repeated forever.

The basic issue is scale. If we divide by a number b_n\to0 , then the bounded variable X_{n,j} becomes X_{n,j}/b_n . Even though |X_{n,j}|\le1 , the normalized variable satisfies \displaystyle \left|\frac{X_{n,j}}{b_n}\right|\le\frac1{b_n}, and 1/b_n\to\infty . Thus the original variable is bounded, but the magnified variable is not bounded uniformly in n . The interval [-1,1] becomes [-1/b_n,1/b_n] , and this expands toward the whole real line. This is the basic mechanism by which bounded triangular arrays can produce infinite-variance or infinite-mean limits.

So the right question is not simply: are the summands bounded? The right question is: after normalization, what does one summand look like? Does every normalized summand become tiny? Does a finite number of jumps remain visible? Do jumps survive at many different scales? Does the maximum matter on the same scale as the whole sum? These are the questions that separate the Gaussian world from the Poisson world and the stable world.

A useful diagnostic throughout the article will be the scale-counting quantity

\displaystyle n\Pr(|X_{n,j}|>x).

This has a direct meaning: among n independent observations, it is approximately the expected number whose magnitude is larger than x . If this number is huge, then many observations are visible above scale x . If it tends to a finite nonzero number, then jumps at that scale survive individually. If it tends to zero for every fixed x>0 , then perhaps we are looking too coarsely, and the interesting structure has moved to a smaller scale. This quantity often sees more than the variance sees.

The Fourier Idea

Before going through examples, let us record the main calculation principle. The characteristic function of a random variable Y is

\displaystyle \phi_Y(t)=\mathbb E e^{itY}.

If S_n=X_{n,1}+\cdots+X_{n,n} and the X_{n,j} are independent, then

\displaystyle \mathbb E e^{itS_n}=\prod_{j=1}^n\mathbb E e^{itX_{n,j}}.

In the identically distributed row case, this is

\displaystyle \mathbb E e^{itS_n}=\left(\mathbb E e^{itX_{n,1}}\right)^n.

Taking logarithms gives

\displaystyle \log\mathbb E e^{itS_n}=n\log\mathbb E e^{itX_{n,1}}.

This is why characteristic functions are so useful for sums. Independence turns sums of random variables into products of characteristic functions, and logarithms turn products into sums. If one summand makes a small contribution to the characteristic function, then multiplying by n tells us whether that contribution disappears, survives, or explodes.

In many triangular-array problems, for fixed t one has \displaystyle \mathbb E e^{itX_{n,1}}=1+\text{a small term}. Then \displaystyle n\log(1+\text{small term})\approx n\cdot\text{small term}, provided the small term is small enough. So the central question becomes: what is the first non-negligible term in the one-step characteristic function, and what happens after multiplying by n ? This is the Fourier-side version of the physical question: what remains visible after adding many independent pieces?

Fair coins

Start with the cleanest model. Let \xi_1,\ldots,\xi_n be independent fair signs, so

\displaystyle \Pr(\xi_j=1)=\Pr(\xi_j=-1)=\frac12,

and put \displaystyle S_n=\xi_1+\cdots+\xi_n. Since \mathbb E\xi_j=0 and \mathrm{Var}(\xi_j)=1 , independence gives \displaystyle \mathrm{Var}(S_n)=n. Thus the natural size of S_n is \sqrt n , not n . The positive and negative signs cancel enough that the typical imbalance is of order \sqrt n . So we look at

\displaystyle \frac{S_n}{\sqrt n}=\frac{\xi_1}{\sqrt n}+\cdots+\frac{\xi_n}{\sqrt n}.

This display already contains the Gaussian geometry. Each normalized summand has size 1/\sqrt n , so

\displaystyle \max_j\left|\frac{\xi_j}{\sqrt n}\right|=\frac1{\sqrt n}\to0.

Thus no individual toss remains visible. But the total variance remains visible, since

\displaystyle \sum_{j=1}^n\mathrm{Var}\left(\frac{\xi_j}{\sqrt n}\right)=n\cdot\frac1n=1.

So every single contribution disappears, but the total quadratic fluctuation remains order one. This is exactly the situation where Gaussian behavior should arise.

Now calculate the characteristic function. For one normalized sign,

\displaystyle \mathbb E e^{it\xi_j/\sqrt n}=\frac12e^{it/\sqrt n}+\frac12e^{-it/\sqrt n}=\cos(t/\sqrt n).

By independence, \displaystyle \mathbb E e^{itS_n/\sqrt n}=\cos(t/\sqrt n)^n. Set u=t/\sqrt n . Since \displaystyle \log\cos u=-\frac{u^2}{2}-\frac{u^4}{12}+O(u^6), we get

\displaystyle n\log\cos(t/\sqrt n)=-\frac{t^2}{2}-\frac{t^4}{12n}+O(n^{-2}).

This line is the whole story. The quadratic term survives because \displaystyle n\left(\frac{t}{\sqrt n}\right)^2=t^2. The quartic term disappears because \displaystyle n\left(\frac{t}{\sqrt n}\right)^4=\frac{t^4}{n}\to0. The sixth-order term disappears even faster. More generally, a term of order u^k contributes roughly \displaystyle n\left(\frac{t}{\sqrt n}\right)^k=t^k n^{1-k/2}. When k=2 , this is order one. When k>2 , it vanishes. The linear term would have size t\sqrt n , but it is absent because the variable is centered. Positive and negative signs cancel at first order.

Therefore

\displaystyle \mathbb E e^{itS_n/\sqrt n}\to e^{-t^2/2},

which is the characteristic function of N(0,1) . So Gaussian behavior comes from a specific balance. The normalization makes each individual summand invisible, but preserves the total variance. If we divided by a much larger scale than \sqrt n , even the variance would disappear and the limit would collapse to zero. If we divided by a much smaller scale, the sum would spread out. The scale \sqrt n is the scale where the quadratic term is neither killed nor blown up.

Poisson behavior

Now change just one feature. Instead of receiving a nonzero contribution on every trial, suppose most trials give zero and only rarely something happens. Let

\displaystyle B_{n,j}=1 \text{ with probability }p_n,\qquad B_{n,j}=0 \text{ with probability }1-p_n.

Define \displaystyle N_n=B_{n,1}+\cdots+B_{n,n}. This counts the number of successes among n trials. By linearity of expectation, \displaystyle \mathbb E N_n=np_n. The quantity np_n is the expected number of visible events. In the fair-coin model, something happened on every trial, so the number of active contributions was exactly n . Here the number of active contributions is random, and its typical size is controlled by np_n .

There are three regimes.

If np_n\to0 , then even one success is unlikely. Indeed, \displaystyle \Pr(N_n=0)=(1-p_n)^n. When p_n is small, (1-p_n)^n\approx e^{-np_n} . If np_n\to0 , then \Pr(N_n=0)\to1 , so N_n\to0 in probability. The events are too rare to survive.

If np_n\to\infty , then many successes occur. After subtracting the mean and dividing by the standard deviation, one is again in a Gaussian-type regime: \displaystyle \frac{N_n-np_n}{\sqrt{np_n(1-p_n)}}\Rightarrow N(0,1), under the usual nondegeneracy assumptions. Many small centered errors accumulate, and no single trial decides the final normalized value.

The new regime is the middle one: \displaystyle np_n\to\lambda\in(0,\infty). Here each event becomes rarer and rarer, but the number of opportunities grows at just the right rate. We should expect a finite random number of visible events. This is the Poisson regime.

For fixed k\ge0 , \displaystyle \Pr(N_n=k)=\binom nkp_n^k(1-p_n)^{n-k}. If np_n\to\lambda , then p_n\to0 , and p_n behaves like \lambda/n . For fixed k , \displaystyle \binom nk\sim\frac{n^k}{k!},\qquad p_n^k\sim\left(\frac{\lambda}{n}\right)^k. Thus \displaystyle \binom nkp_n^k\sim\frac{\lambda^k}{k!}. Also, \displaystyle (1-p_n)^{n-k}=(1-p_n)^n(1-p_n)^{-k}\to e^{-\lambda}. Therefore

\displaystyle \Pr(N_n=k)\to e^{-\lambda}\frac{\lambda^k}{k!}.

So \displaystyle N_n\Rightarrow\mathrm{Poisson}(\lambda). This is already a non-Gaussian limit from bounded variables taking only the values 0 and 1 . The reason is not mysterious. In the Gaussian coin model, many contributions survive but each becomes individually invisible after normalization. In the Poisson regime, only finitely many events survive, and they remain visible. The limit remembers the number of rare events.

Poisson jumps or Gaussian fluctuation

Now let the rare event have a sign. Define

\displaystyle X_{n,j}=1 \text{ with probability }p_n/2,\qquad X_{n,j}=-1 \text{ with probability }p_n/2,\qquad X_{n,j}=0 \text{ with probability }1-p_n.

Most of the time nothing happens. When a jump occurs, it is equally likely to be positive or negative. The variable is still bounded and simple: it takes only the values -1,0,1 . Let \displaystyle S_n=X_{n,1}+\cdots+X_{n,n}. By symmetry, \mathbb E X_{n,j}=0 . Also, X_{n,j}^2=1 exactly when a nonzero jump occurs. Therefore

\displaystyle \mathbb E X_{n,j}^2=p_n,\qquad \mathrm{Var}(X_{n,j})=p_n,

and hence \displaystyle \mathrm{Var}(S_n)=np_n. Again, np_n measures the expected number of active jumps.

If np_n\to\lambda\in(0,\infty) , then only finitely many nonzero jumps survive. The number of positive jumps converges to \mathrm{Poisson}(\lambda/2) , and the number of negative jumps converges to an independent \mathrm{Poisson}(\lambda/2) . Thus

\displaystyle S_n\Rightarrow N_+-N_-,

where N_+ and N_- are independent \mathrm{Poisson}(\lambda/2) random variables. This is a jump limit. The limit remembers the surviving positive and negative jumps.

The characteristic function gives the same result. For one summand,

\displaystyle \mathbb E e^{itX_{n,j}}=1-p_n+\frac{p_n}{2}e^{it}+\frac{p_n}{2}e^{-it}=1+p_n(\cos t-1).

Therefore \displaystyle \mathbb E e^{itS_n}=\left(1+p_n(\cos t-1)\right)^n.

If np_n\to\lambda , this tends to \displaystyle \exp(\lambda(\cos t-1)).

This is the characteristic function of the difference of two independent \mathrm{Poisson}(\lambda/2) variables.

If instead np_n\to\infty , then the number of nonzero jumps diverges. The natural scale is the standard deviation \sqrt{np_n} . Under this scaling, one nonzero jump has size \displaystyle \frac1{\sqrt{np_n}}\to0. Now individual jumps are no longer visible. We are back in the Gaussian geometry. Indeed,

\displaystyle \mathbb E e^{itS_n/\sqrt{np_n}}=\left(1+p_n\left(\cos\left(\frac{t}{\sqrt{np_n}}\right)-1\right)\right)^n.

Using \cos u-1=-u^2/2+O(u^4) with u=t/\sqrt{np_n} , we get

\displaystyle p_n\left(\cos\left(\frac{t}{\sqrt{np_n}}\right)-1\right)=-\frac{t^2}{2n}+O\left(\frac{1}{n^2p_n}\right).

Multiplying by n , the main term becomes -t^2/2 , and the error vanishes if np_n\to\infty . Thus

\displaystyle \frac{S_n}{\sqrt{np_n}}\Rightarrow N(0,1).

The same bounded variables produce two different limits. If np_n stays finite, the limit is made of finitely many visible jumps. If np_n tends to infinity, the number of jumps grows, each normalized jump disappears, and the Gaussian mechanism returns.

Compound Poisson limits

The rare signed coin has only one nonzero magnitude: every jump has size 1 . To get a richer jump limit, allow the size of the jump to be random. Let

\displaystyle X_{n,j}=B_{n,j}U_j,

where B_{n,j}\sim\mathrm{Bernoulli}(p_n) , U_j\sim U[-1,1] , and all variables are independent. The Bernoulli variable decides whether a jump occurs. If B_{n,j}=0 , then X_{n,j}=0 . If B_{n,j}=1 , then the jump size is U_j , uniformly distributed in [-1,1] .

Assume np_n\to\lambda\in(0,\infty) . Then the number of jumps converges to a Poisson random variable N\sim\mathrm{Poisson}(\lambda) . Conditional on N=k , the limiting contribution is

\displaystyle U_1+\cdots+U_k.

Thus the limit has the form \displaystyle Y=\sum_{j=1}^N U_j,\qquad N\sim\mathrm{Poisson}(\lambda),\qquad U_j\sim U[-1,1]. This is a compound Poisson random variable. The word “Poisson” records the random number of jumps. The word “compound” records that each jump carries a random size.

The characteristic function confirms this. For U\sim U[-1,1] ,

\displaystyle \mathbb E e^{itU}=\frac12\int_{-1}^{1}e^{itx} dx=\frac{\sin t}{t}.

For one summand,

\displaystyle \mathbb E e^{itX_{n,j}}=(1-p_n)\cdot 1+p_n\frac{\sin t}{t}=1+p_n\left(\frac{\sin t}{t}-1\right).

Therefore

\displaystyle \mathbb E e^{itS_n}=\left(1+p_n\left(\frac{\sin t}{t}-1\right)\right)^n.

If np_n\to\lambda , then

\displaystyle n\log\left(1+p_n\left(\frac{\sin t}{t}-1\right)\right)\to\lambda\left(\frac{\sin t}{t}-1\right),

and hence

\displaystyle \mathbb E e^{itS_n}\to\exp\left(\lambda\left(\frac{\sin t}{t}-1\right)\right).

This is exactly the characteristic function of \sum_{j=1}^NU_j . If we condition on N , then

\displaystyle \mathbb E(e^{itY}\mid N)=\left(\frac{\sin t}{t}\right)^N.

Averaging over N\sim\mathrm{Poisson}(\lambda) gives the same exponential.

So when finitely many jumps survive, the limit remembers their sizes. This suggests the next question: how complicated can the surviving jump-size distribution become?

A first attempt is to change the width of a familiar distribution. Suppose \displaystyle X_{n,j}\sim U[-a_n,a_n]. This changes the physical size of the summand, but it does not create a hierarchy of scales. A typical observation is of size comparable to a_n . There are not rare jumps at size a_n , more common jumps at size a_n/10 , and even more common jumps at a_n/100 . There is one scale.

Indeed, \displaystyle \mathrm{Var}(X_{n,j})=\frac{a_n^2}{3}. The standard normalized variable is

\displaystyle Z_n=\frac{\sqrt3 S_n}{a_n\sqrt n}.

The factor a_n disappears after scaling. The characteristic function is

\displaystyle \phi_{Z_n}(t)=\left(\frac{\sin(\sqrt3 t/\sqrt n)}{\sqrt3 t/\sqrt n}\right)^n.

Since \displaystyle \log\left(\frac{\sin u}{u}\right)=-\frac{u^2}{6}+O(u^4), we get

\displaystyle \log\phi_{Z_n}(t)=n\left(-\frac{(\sqrt3,t/\sqrt n)^2}{6}+O(n^{-2})\right)=-\frac{t^2}{2}+O(n^{-1}).

Thus \displaystyle Z_n\Rightarrow N(0,1). Changing the width of a uniform distribution changes the unit of measurement, not the limiting mechanism. It does not produce a new jump geometry.

Now try a slightly richer model. Suppose a rare jump can have one of two magnitudes, say 1 or 1/10 . Let a jump occur with probability p_n . Conditional on a jump, suppose it has signed size 1 with probability q , and signed size 1/10 with probability 1-q . Then one summand has characteristic function

\displaystyle 1+p_n\left[q(\cos t-1)+(1-q)(\cos(t/10)-1)\right].

The sum has characteristic function

\displaystyle \left(1+p_n\left[q(\cos t-1)+(1-q)(\cos(t/10)-1)\right]\right)^n.

If np_n\to\lambda , this tends to

\displaystyle \exp\left(\lambda q(\cos t-1)+\lambda(1-q)(\cos(t/10)-1)\right).

This is still compound Poisson. We have not created a fundamentally new kind of limit. We have only allowed the finitely many surviving jumps to choose between two possible sizes.

The same calculation works for any fixed finite list of jump sizes. Suppose possible magnitudes are a_1,\ldots,a_m , and a signed jump of size a_r occurs with probability p_{n,r} . Then one summand has characteristic function

\displaystyle 1+\sum_{r=1}^m p_{n,r}(\cos(ta_r)-1).

The sum has characteristic function

\displaystyle \left(1+\sum_{r=1}^m p_{n,r}(\cos(ta_r)-1)\right)^n.

If np_{n,r}\to\lambda_r for each r , then the limit is

\displaystyle \exp\left(\sum_{r=1}^m\lambda_r(\cos(ta_r)-1)\right).

This is again compound Poisson, now with finitely many jump classes. A finite list of jump sizes gives a finite sum in the exponent. To get something more like a stable law, we need more and more possible scales.

Now allow the number of possible scales to grow with n . Imagine possible magnitudes \displaystyle 1,\frac12,\frac14,\ldots,2^{-K_n}, where K_n\to\infty . Suppose a signed jump of size 2^{-r} occurs with probability p_{n,r} . Then one summand has characteristic function

\displaystyle 1+\sum_{r=0}^{K_n}p_{n,r}(\cos(t2^{-r})-1),

and the sum has characteristic function

\displaystyle \left(1+\sum_{r=0}^{K_n}p_{n,r}(\cos(t2^{-r})-1)\right)^n.

The important quantities are \displaystyle \lambda_{n,r}=np_{n,r}. This is the expected number of jumps at scale 2^{-r} among n observations. If only finitely many \lambda_{n,r} remain significant, we fall back to finite compound Poisson behavior. If many tiny scales contribute and individual jumps become negligible, a Gaussian component may appear. If jumps survive across a growing range of scales, we get an infinitely divisible jump law.

Instead of tracking each p_{n,r} , ask for the expected number of jumps larger than a threshold x :

\displaystyle n\Pr(|X_{n,j}|>x).

For the dyadic model, if x=2^{-m} , then jumps larger than x are the jumps at levels r=0,1,\ldots,m . Therefore

\displaystyle n\Pr(|X_{n,j}|>2^{-m})\approx\sum_{r=0}^m np_{n,r}=\sum_{r=0}^m\lambda_{n,r}.

So the many-scale problem becomes: how do these cumulative counts grow as we move down the scales?

If the cumulative count stays bounded, only finitely many jumps survive. If it grows very fast but only at tiny scales, those tiny jumps may average into Gaussian behavior. If it grows regularly from scale to scale, then a stable-type hierarchy appears.

Power laws enter as the cleanest scale-regular choice. Suppose that each time we reduce the scale by a factor of 2 , the expected number of visible jumps is multiplied by 2^\alpha . Then the expected number of jumps above scale 2^{-m} grows like \displaystyle 2^{\alpha m}=(2^{-m})^{-\alpha}. In continuous notation, this becomes

\displaystyle n\Pr(|X_{n,j}|>x)\approx Cx^{-\alpha}.

This is the scale-counting reason for power laws. We are not inserting a power law just because stable laws are famous. We are asking for jump counts that transform regularly as we zoom in. Reducing x by a factor of 10 multiplies the expected number of visible jumps by 10^\alpha . To realize this with bounded finite-stage variables, we use a cutoff power-law density. Fix 0<\alpha<2 , and let \varepsilon_n\to0 . Define a symmetric density on [-1,1] by

\displaystyle f_n(x)=\frac{\alpha\varepsilon_n^\alpha}{2(1-\varepsilon_n^\alpha)}|x|^{-1-\alpha}\mathbf 1_{{\varepsilon_n\le |x|\le1}}.

The upper cutoff 1 keeps all variables bounded. The lower cutoff \varepsilon_n says that at stage n we include only finitely many scales. As n grows, \varepsilon_n decreases, so more small scales enter.

Let us compute the tail. For \varepsilon_n\le u\le1 , \displaystyle \Pr(|X_{n,j}|>u)=2\int_u^1 f_n(x) dx. Substituting the density,

\displaystyle \Pr(|X_{n,j}|>u)=\frac{\alpha\varepsilon_n^\alpha}{1-\varepsilon_n^\alpha}\int_u^1 x^{-1-\alpha}dx.

Since \displaystyle \int_u^1 x^{-1-\alpha} dx=\frac{u^{-\alpha}-1}{\alpha}, we get

\displaystyle \Pr(|X_{n,j}|>u)=\frac{\varepsilon_n^\alpha}{1-\varepsilon_n^\alpha}(u^{-\alpha}-1).

Thus for fixed u\in(0,1) ,

\displaystyle n\Pr(|X_{n,j}|>u)\sim n\varepsilon_n^\alpha(u^{-\alpha}-1).

This identifies the central parameter \displaystyle \theta_n=n\varepsilon_n^\alpha.

For rare coins, np_n counted how many jumps of one available size survived. Here \theta_n counts how many macroscopic jumps survive in a many-scale distribution.

Moments also reflect the same parameter. For q>\alpha ,

\displaystyle \mathbb E|X_{n,j}|^q=2\int_{\varepsilon_n}^1 x^q f_n(x) dx=\frac{\alpha\varepsilon_n^\alpha}{1-\varepsilon_n^\alpha}\int_{\varepsilon_n}^1 x^{q-1-\alpha} dx.

Since \displaystyle \int_{\varepsilon_n}^1 x^{q-1-\alpha} dx=\frac{1-\varepsilon_n^{q-\alpha}}{q-\alpha}, we have

\displaystyle \mathbb E|X_{n,j}|^q=\frac{\alpha\varepsilon_n^\alpha}{1-\varepsilon_n^\alpha}\cdot\frac{1-\varepsilon_n^{q-\alpha}}{q-\alpha}\sim\frac{\alpha}{q-\alpha}\varepsilon_n^\alpha.

In particular, for q=2 and 0<\alpha<2 ,

\displaystyle \mathrm{Var}(X_{n,j})=\mathbb E X_{n,j}^2\sim\frac{\alpha}{2-\alpha}\varepsilon_n^\alpha.

Therefore

\displaystyle \mathrm{Var}(S_n)\sim\frac{\alpha}{2-\alpha}n\varepsilon_n^\alpha=\frac{\alpha}{2-\alpha}\theta_n.

Thus the same quantity \theta_n controls both macroscopic jump counts and total variance.

The dense regime: many jumps, Gaussian after standardization

First assume \displaystyle \theta_n\to\infty. Then for every fixed u\in(0,1) ,

\displaystyle n\Pr(|X_{n,j}|>u)\sim\theta_n(u^{-\alpha}-1)\to\infty.

So there are many jumps above every fixed visible scale. At first this may look very non-Gaussian. But the total variance is also going to infinity:

\displaystyle \mathrm{Var}(S_n)\sim\frac{\alpha}{2-\alpha}\theta_n\to\infty.

Let \displaystyle b_n^2=\mathrm{Var}(S_n). Then b_n\to\infty . Since |X_{n,j}|\le1 , \displaystyle \frac{|X_{n,j}|}{b_n}\le\frac1{b_n}\to0. At the standard-deviation scale of the sum, every single jump becomes invisible. The entire interval [-1,1] collapses to [-1/b_n,1/b_n] .

Now compute the characteristic function of S_n/b_n . Since the variables are centered and symmetric,

\displaystyle \mathbb E e^{itX_{n,j}/b_n}=1-\frac{t^2\mathbb E X_{n,j}^2}{2b_n^2}+R_n(t),

where the remainder can be bounded by a constant times \displaystyle \frac{|t|^3\mathbb E|X_{n,j}|^3}{b_n^3}.

Multiplying by n , the quadratic term gives

\displaystyle -\frac{t^2}{2}\frac{n\mathbb E X_{n,j}^2}{b_n^2}=-\frac{t^2}{2}.

For the remainder, since n\mathbb E|X_{n,j}|^3 is of order \theta_n , while b_n^3 is of order \theta_n^{3/2} , the total error is of order \displaystyle \frac{\theta_n}{\theta_n^{3/2}}=\theta_n^{-1/2}\to0. Therefore

\displaystyle \frac{S_n}{\sqrt{\mathrm{Var}(S_n)}}\Rightarrow N(0,1).

In the dense regime, the many-scale structure exists at the raw level, but the standard-deviation normalization crushes every individual jump. Only the total quadratic variation remains. That is why the Gaussian mechanism returns.

The critical regime: finitely many visible jumps at each fixed scale

Now assume \displaystyle \theta_n\to\lambda\in(0,\infty). Then for fixed u\in(0,1) ,

\displaystyle n\Pr(|X_{n,j}|>u)\to\lambda(u^{-\alpha}-1).

Above each fixed threshold u , only finitely many jumps survive. So we should not divide by a scale going to infinity, because doing so would crush the visible jumps. We look at the raw sum. By symmetry,

\displaystyle \mathbb E e^{itX_{n,j}}-1=\frac{\alpha\varepsilon_n^\alpha}{1-\varepsilon_n^\alpha}\int_{\varepsilon_n}^1(\cos(tx)-1)x^{-1-\alpha} dx.

Multiplying by n and using n\varepsilon_n^\alpha\to\lambda gives

\displaystyle n(\mathbb E e^{itX_{n,j}}-1)\to\alpha\lambda\int_0^1(\cos(tx)-1)x^{-1-\alpha} dx.

Hence

\displaystyle \mathbb E e^{itS_n}\to\exp\left(\alpha\lambda\int_0^1(\cos(tx)-1)x^{-1-\alpha} dx\right).

This is not Gaussian. A Gaussian exponent is quadratic in t . Here the exponent is an integral over jump sizes 0<x\le1 . The finite-scale model gave finite sums such as

\displaystyle \sum_r\lambda_r(\cos(ta_r)-1).

The many-scale model replaces the finite sum by an integral. The corresponding jump intensity is

\displaystyle \nu(dx)=\frac{\alpha\lambda}{2}|x|^{-1-\alpha}\mathbf 1_{{0<|x|\le1}} dx.

This is a Lévy measure. Concretely, it says how many jumps of each size survive. The factor |x|^{-1-\alpha} says that smaller jumps are more numerous. The cutoff |x|\le1 remembers the original bounded support. This is why the limit is naturally called a truncated stable law: it has stable-like behavior at small scales, but no jump larger than the original cutoff.

The sparse regime:

Now assume \displaystyle \theta_n\to0. Then for every fixed u>0 ,

\displaystyle n\Pr(|X_{n,j}|>u)\to0.

No jump survives at any fixed macroscopic scale. Also \displaystyle \mathrm{Var}(S_n)\sim\frac{\alpha}{2-\alpha}\theta_n\to0, so S_n\to0 in probability. If we look at the original scale, the limit is trivial. But the sum vanished because we were looking too coarsely. The interesting activity has moved below every fixed scale. To find the correct scale, solve for b_n such that order-one many jumps exceed b_n :

\displaystyle n\Pr(|X_{n,j}|>b_n)\asymp1.

Using \Pr(|X_{n,j}|>x)\approx\varepsilon_n^\alpha x^{-\alpha} , we get \displaystyle n\varepsilon_n^\alpha b_n^{-\alpha}\approx1. Thus \displaystyle b_n=\varepsilon_n n^{1/\alpha}. Since \displaystyle b_n^\alpha=n\varepsilon_n^\alpha=\theta_n\to0, we have b_n\to0 . So we magnify the sum and study S_n/b_n . The support changes dramatically. Originally active magnitudes lie in [\varepsilon_n,1] . After division by b_n , they lie in \displaystyle \left[\frac{\varepsilon_n}{b_n},\frac1{b_n}\right]. But

\displaystyle \frac{\varepsilon_n}{b_n}=n^{-1/\alpha}\to0,\qquad \frac1{b_n}\to\infty.

So under the stable magnification, the bounded interval of possible magnitudes expands toward (0,\infty) . This is how a full unbounded stable law emerges from bounded finite-stage variables.

At the magnified scale, for fixed x>0 , \displaystyle n\Pr(|X_{n,j}|>b_nx)\to x^{-\alpha}. Thus the number of normalized jumps larger than x has the stable scale-counting form. Now compute the characteristic function. We have

\displaystyle n(\mathbb E e^{itX_{n,j}/b_n}-1)=n\frac{\alpha\varepsilon_n^\alpha}{1-\varepsilon_n^\alpha}\int_{\varepsilon_n}^{1}\left(\cos\left(\frac{tx}{b_n}\right)-1\right)x^{-1-\alpha} dx.

Change variables x=b_ny . Then dx=b_n dy , and

\displaystyle x^{-1-\alpha} dx=(b_ny)^{-1-\alpha}b_n,dy=b_n^{-\alpha}y^{-1-\alpha} dy.

So the expression becomes approximately

\displaystyle n\alpha\varepsilon_n^\alpha b_n^{-\alpha}\int_{\varepsilon_n/b_n}^{1/b_n}(\cos(ty)-1)y^{-1-\alpha} dy.

But b_n^\alpha=n\varepsilon_n^\alpha , so the prefactor is \alpha . The lower limit tends to 0 and the upper limit tends to \infty . Therefore

\displaystyle n(\mathbb E e^{itX_{n,j}/b_n}-1)\to \alpha\int_0^\infty(\cos(ty)-1)y^{-1-\alpha} dy.

The integral has the scaling form of a stable law. Indeed, with z=|t|y ,

\displaystyle \int_0^\infty(1-\cos(ty))y^{-1-\alpha},dy=|t|^\alpha\int_0^\infty(1-\cos z)z^{-1-\alpha} dz.

Thus

\displaystyle \mathbb E e^{itS_n/b_n}\to e^{-c_\alpha |t|^\alpha},\qquad c_\alpha=\alpha\int_0^\infty(1-\cos z)z^{-1-\alpha} dz.

So \displaystyle \frac{S_n}{b_n}\Rightarrow Z_\alpha, where Z_\alpha is a symmetric \alpha -stable random variable.

The three regimes can now be summarized cleanly. If \theta_n\to\infty , many jumps survive at fixed scales, but the standard-deviation scale grows so much that individual jumps disappear; the limit is Gaussian. If \theta_n\to\lambda , finitely many jumps survive at fixed scales; the raw sum has a truncated stable jump limit. If \theta_n\to0 , no fixed-scale jumps survive; the raw sum vanishes, but a smaller scale reveals a full stable law.

What convergence can hide

At finite n , every X_{n,j} is bounded. Yet in the sparse stable regime, the normalized sums S_n/b_n may have variances tending to infinity. This is not a contradiction. Convergence in distribution does not automatically imply convergence of moments.

In the sparse stable regime, \displaystyle \mathrm{Var}(S_n)\asymp\theta_n,\qquad b_n^2=\theta_n^{2/\alpha}. Therefore

\displaystyle \mathrm{Var}(S_n/b_n)\asymp \frac{\theta_n}{\theta_n^{2/\alpha}}=\theta_n^{1-2/\alpha}.

If \alpha<2 , then 1-2/\alpha<0 . Since \theta_n\to0 , the variance diverges. So S_n/b_n can converge in distribution to a stable random variable while its variances blow up.

A toy example shows the danger. Let Y_n=n with probability 1/n , and Y_n=0 otherwise. Then Y_n\to0 in probability, because the exceptional event has probability 1/n\to0 . But \displaystyle \mathbb E Y_n=n\cdot\frac1n=1. If instead Y_n=n^2 with probability 1/n , then still Y_n\to0 in probability, but \displaystyle \mathbb E Y_n=n^2\cdot\frac1n=n\to\infty. Rare events can disappear from ordinary probability convergence while controlling expectations and higher moments. Stable limits are a systematic version of this phenomenon.

For positive stable laws, infinite mean appears similarly. Suppose 0<\alpha<1 and use the one-sided density \displaystyle f_n(x)=\frac{\alpha\varepsilon_n^\alpha}{1-\varepsilon_n^\alpha}x^{-1-\alpha}\mathbf 1_{{\varepsilon_n\le x\le1}}. Each variable satisfies 0<X_{n,j}\le1 . In the sparse regime \theta_n=n\varepsilon_n^\alpha\to0 , with b_n=\varepsilon_n n^{1/\alpha} , the normalized sum S_n/b_n converges to a positive \alpha -stable law. The raw mean satisfies \mathbb E S_n\asymp\theta_n\to0 . But the normalized mean behaves like \displaystyle \frac{\mathbb E S_n}{b_n}\asymp\frac{\theta_n}{b_n}. Since b_n=\theta_n^{1/\alpha} , this is \displaystyle \theta_n^{1-1/\alpha}.

For \alpha<1 , the exponent 1-1/\alpha is negative, so this tends to infinity. Thus the raw sums vanish, while the normalized sums converge to a finite random variable with infinite mean. The infinite mean comes from rare normalized extremes.

Largest Term

Another way to distinguish Gaussian and stable behavior is to look at the maximum

\displaystyle M_n=\max_{1\le j\le n}|X_{n,j}|.

In the dense Gaussian regime, with b_n=\sqrt{\mathrm{Var}(S_n)}\to\infty ,

\displaystyle \frac{M_n}{b_n}\le\frac1{b_n}\to0.

The largest observation is negligible compared with the scale of the whole sum. This is one of the signatures of Gaussian behavior.

In the sparse stable regime, with b_n=\varepsilon_n n^{1/\alpha} , the maximum is on the same scale as the sum. For fixed x>0 ,

\displaystyle \Pr(M_n/b_n\le x)=\Pr(|X_{n,1}|\le b_nx)^n.

This equals \displaystyle \left(1-\Pr(|X_{n,1}|>b_nx)\right)^n. Since \displaystyle n\Pr(|X_{n,1}|>b_nx)\to x^{-\alpha}, we get

\displaystyle \Pr(M_n/b_n\le x)\to e^{-x^{-\alpha}}.

So in the stable regime, the largest observation remains visible. It is not necessarily the whole sum, because several large jumps may contribute, but it lives on the same scale. This is a major difference between Gaussian and stable addition. In the Gaussian world, the maximum is negligible. In the stable world, extreme terms are part of the main structure.

The Lévy–Khintchine picture

We have now seen several mechanisms: Gaussian behavior appears when many tiny centered contributions accumulate, with no individual term visible. Poisson behavior appears when finitely many visible events survive. Compound Poisson behavior appears when finitely many visible jumps survive, and their sizes are random. Truncated stable behavior appears when jumps survive across infinitely many small scales, but the original upper cutoff remains visible. Full stable behavior appears when the correct normalization magnifies the bounded support so much that the upper cutoff disappears.

The Lévy–Khintchine formula is the general bookkeeping system for all of these mechanisms. Let us build it from examples rather than state it abstractly. Suppose a summand has four independent pieces:

\displaystyle X_{n,j}=\frac{a}{n}+\frac{\sigma\xi_j}{\sqrt n}+B_{n,j}J_j+Y_{n,j}.

Here a/n is a tiny deterministic displacement, \sigma\xi_j/\sqrt n is a tiny centered fluctuation, B_{n,j}J_j is a rare visible jump with B_{n,j}\sim\mathrm{Bernoulli}(\lambda/n) , and Y_{n,j} is a many-scale small-jump variable.

Because the four pieces are independent, the one-step characteristic function approximately factors into four pieces:

\displaystyle \mathbb E e^{itX_{n,j}}\approx e^{ita/n}\cdot\cos(\sigma t/\sqrt n)\cdot\left(1+\frac{\lambda}{n}(\mathbb E e^{itJ}-1)\right)\cdot\left(1+\text{small-jump contribution}\right).

Now raise this to the n th power and take logarithms. The deterministic part gives

\displaystyle n\cdot\frac{iat}{n}=iat.

The Gaussian part gives

\displaystyle n\log\cos(\sigma t/\sqrt n)\to-\frac{\sigma^2t^2}{2}.

The rare visible jumps give

\displaystyle n\log\left(1+\frac{\lambda}{n}(\mathbb E e^{itJ}-1)\right)\to\lambda(\mathbb E e^{itJ}-1).

The many-scale jumps give an integral over jump sizes. If the limiting jump intensity is a measure \nu , then the jump contribution has the form

\displaystyle \int\left(e^{itx}-1-itx\mathbf 1_{{|x|\le1}}\right)\nu(dx),

with a possible drift correction if the small jumps are asymmetric.

Thus a general limiting exponent has the form

\displaystyle i\gamma t-\frac{\sigma^2t^2}{2}+\int_{\mathbb R\setminus{0}}\left(e^{itx}-1-itx\mathbf 1_{{|x|\le1}}\right)\nu(dx).

This is the Lévy–Khintchine formula. The characteristic function is

\displaystyle \mathbb E e^{itY}=\exp\left(i\gamma t-\frac{\sigma^2t^2}{2}+\int_{\mathbb R\setminus{0}}\left(e^{itx}-1-itx\mathbf 1_{{|x|\le1}}\right)\nu(dx)\right).

Each term has a concrete meaning.

The term i\gamma t is drift. In the simplest case, if every summand contributes a/n , then the sum contributes a , and the characteristic function is e^{iat} . That is pure deterministic motion. In more complicated asymmetric jump models, drift also includes the deterministic first-order part left after small-jump compensation. So drift is not simply “the mean.” It is the net deterministic motion after we have separated visible jumps from their compensated small-jump part. The term -\sigma^2t^2/2 is the Gaussian part. It comes from many individually invisible centered fluctuations whose total variance converges to \sigma^2 . The measure \nu is the Lévy measure. It records the intensity of jumps of different sizes. If A is a set away from zero, then \nu(A) can be read as the limiting expected number of jumps with sizes in A .

For compound Poisson jumps with jump distribution \mu and rate \lambda , \displaystyle \nu(dx)=\lambda\mu(dx). Then

\displaystyle \int(e^{itx}-1)\nu(dx)=\lambda(\mathbb E e^{itJ}-1),

which is exactly the compound Poisson exponent. If there are finitely many jump sizes a_1,\ldots,a_m , then \displaystyle \nu=\lambda_1\delta_{a_1}+\cdots+\lambda_m\delta_{a_m}, and the jump exponent becomes

\displaystyle \sum_{r=1}^m\lambda_r(e^{ita_r}-1).

If the jumps are symmetric, the imaginary parts cancel and we get

\displaystyle \sum_{r=1}^m\lambda_r(\cos(ta_r)-1).

This is exactly the finite-scale calculation.

In the truncated stable regime, the finite sum becomes an integral. The Lévy measure is

\displaystyle \nu(dx)=\frac{\alpha\lambda}{2}|x|^{-1-\alpha}\mathbf 1_{{0<|x|\le1}} dx.

Since the measure is symmetric,

\displaystyle \int_{\mathbb R}(e^{itx}-1)\nu(dx)=\int_{\mathbb R}(\cos(tx)-1)\nu(dx).

Thus

\displaystyle \int_{\mathbb R}(\cos(tx)-1)\nu(dx)=\alpha\lambda\int_0^1(\cos(tx)-1)x^{-1-\alpha} dx.

This is the continuum version of the finite jump-size sum. The finite-scale model has a finite list of surviving jump sizes. The stable-type model has a continuum of surviving jump sizes.

There is one important subtlety. In the compound Poisson case, \nu(\mathbb R)<\infty , so the total number of jumps is finite. In the truncated stable case,

\displaystyle \int_0^1 x^{-1-\alpha} dx=\infty.

So the total mass of \nu near zero is infinite. This means infinitely many tiny jumps accumulate near zero. However, for every fixed u>0 , the mass of {|x|>u} is finite. There are only finitely many jumps larger than any fixed threshold. The infinity is hidden near zero.

This is why the Lévy–Khintchine formula subtracts the linear term itx\mathbf 1_{{|x|\le1}} . For small x ,

\displaystyle e^{itx}-1=itx-\frac{t^2x^2}{2}+O(x^3).

If there are infinitely many tiny jumps, the integral of the linear term itx may not converge. The compensated expression

\displaystyle e^{itx}-1-itx\mathbf 1_{{|x|\le1}}

behaves like -t^2x^2/2 near zero, which is much more integrable. The defining condition on a Lévy measure is

\displaystyle \int_{\mathbb R\setminus{0}}(1\wedge x^2)\nu(dx)<\infty.

This says: there cannot be too many large jumps, and the infinitely many small jumps must be square-summable enough for the characteristic exponent to make sense after compensation.

The drift term is clearest from this compensation viewpoint. In symmetric examples, drift usually vanishes because positive and negative first-order effects cancel. But in asymmetric examples, infinitely many small jumps may have a systematic average direction. The compensation subtracts their linear part from the integral, and the leftover deterministic correction is placed into the drift \gamma . Thus drift is the deterministic first-order motion left after centering and small-jump compensation. It is not the same thing as replacing visible jumps by their average. Visible rare jumps remain random and belong in the jump measure.

The Lévy–Khintchine formula is therefore not mysterious. It is a compact record of what can survive after adding many independent pieces:

\displaystyle \text{drift}+\text{Gaussian fluctuation}+\text{jumps}.

A pure Gaussian limit occurs when every individual normalized summand is negligible and only quadratic variation remains. A compound Poisson limit occurs when finitely many jumps survive. A truncated stable limit occurs when jumps survive at infinitely many small scales but the upper cutoff remains visible. A full stable limit occurs when normalization magnifies the support so that the cutoff disappears. A drift appears when deterministic first-order motion survives.

A triangular array does not have to choose one clean regime along the whole sequence. The parameter \displaystyle \theta_n=n\varepsilon_n^\alpha may oscillate. Along subsequences where \theta_n\to\infty , one may see Gaussian behavior. Along subsequences where \theta_n\to\lambda\in(0,\infty) , one may see truncated stable behavior. Along subsequences where \theta_n\to0 , one may see stable behavior after magnification. Limits depend on parameter regimes, not merely on the formula for one X_{n,j} .

There are also borderline corrections. A tail may behave like \displaystyle x^{-\alpha}L(x), where L varies slowly with scale. Then the correct normalization may include logarithmic or slowly varying factors. The philosophy remains the same: solve

\displaystyle n\Pr(|X_{n,j}|>b_n)\asymp1

to find the scale where visible extremes live, then ask whether the maximum is negligible, finite, or dominant, and ask what happens to the small jumps below that scale.

Conclusion

The central limit theorem is not simply the statement that sums of independent variables become Gaussian. More precisely, Gaussian behavior arises when, after normalization, every individual summand becomes negligible and only the aggregate quadratic variation remains. Poisson behavior arises when only finitely many visible events survive. Compound Poisson behavior arises when finitely many visible events survive and carry random sizes. Truncated stable behavior arises when jumps survive at a continuum of small scales, but the original upper cutoff remains visible. Full stable behavior arises when the raw sum vanishes at fixed scale, but after magnifying to the scale where order-one jumps reappear, the bounded support expands to an unbounded range.

The power-law construction was chosen because we wanted jump counts that behave regularly across scales:

\displaystyle n\Pr(|X_{n,j}|>x)\approx Cx^{-\alpha}.

This leads to densities proportional to |x|^{-1-\alpha} , with a lower cutoff \varepsilon_n and an upper cutoff 1 . The phase diagram is governed by \displaystyle \theta_n=n\varepsilon_n^\alpha. If \theta_n\to\infty , Gaussian behavior returns after standardization. If \theta_n\to\lambda\in(0,\infty) , the raw sum has a truncated stable jump limit. If \theta_n\to0 , the raw sum vanishes, but the magnified sum has a full stable limit.

The final moral is this: boundedness alone does not decide the limiting behavior of sums. What matters is what remains visible at the scale where we observe the sum. The normalization decides which events are crushed to zero, which jumps remain visible, which rare tail events are magnified, and whether the limit records drift, Gaussian fluctuation, jumps, or some combination of all three.

.

Leave a Reply