Frequency and Severity Models

Tail Models and Extreme Values

Model large losses with Weibull, Pareto, and threshold exceedances while making tail uncertainty visible.

Tail Models and Extreme Values

Tail modelling asks a narrow question: how quickly does P(X>x)P(X>x) decay at large xx? It matters when a few losses dominate reserves, reinsurance, or capital.

1. Compare tail classes

ModelSurvival behaviourMoment implicationTypical role
Weibullexp[(x/λ)k]\exp[-(x/\lambda)^k]all positive moments finiteflexible light/stretch-exponential tail
Lognormalslower than any exponential, faster than any powerall moments finitemultiplicative severity benchmark
Pareto/Lomaxpower law (1+x/θ)α(1+x/\theta)^{-\alpha}moment rr exists only if r<αr<\alphaheavy-tail benchmark
GPD exceedanceshape ξ\xi determines tail classmean exists if ξ<1\xi<1losses above a high threshold

“Heavy-tailed” has a mathematical meaning; it is not merely “large variance.”

2. Weibull: shape controls ageing

With scale λ>0\lambda>0 and shape k>0k>0,

Fˉ(x)=exp[(x/λ)k],h(x)=kλ(x/λ)k1.\bar F(x)=\exp[-(x/\lambda)^k], \qquad h(x)=\frac{k}{\lambda}(x/\lambda)^{k-1}.
  • k=1k=1: Exponential, constant hazard;
  • k<1k<1: decreasing hazard;
  • k>1k>1: increasing hazard.

For claim size, hazard is a mathematical description of exceedance, not a literal failure rate unless the insurance mechanism supports that interpretation.

3. Pareto II/Lomax: power-law benchmark

This page uses

Fˉ(x)=(1+xθ)α,x0,α,θ>0.\bar F(x)=\left(1+\frac{x}{\theta}\right)^{-\alpha}, \qquad x\ge0, \alpha,\theta>0.

Then

E[X]=θα1(α>1),E[X]=\frac{\theta}{\alpha-1}\quad(\alpha>1),Var(X)=αθ2(α1)2(α2)(α>2).\operatorname{Var}(X)= \frac{\alpha\theta^2}{(\alpha-1)^2(\alpha-2)}\quad(\alpha>2).

At α=1.5\alpha=1.5, the mean exists but variance is infinite. A finite sample still has a finite sample variance; the instability appears as the sample grows and new extremes arrive.

Some texts use a Pareto Type I distribution with support xxmx\ge x_m and survival (xm/x)α(x_m/x)^\alpha. Always show the support and survival function rather than reporting only “Pareto(α,θ\alpha,\theta).”

4. Peaks over threshold and the GPD

Choose a high threshold uu and define excess Y=XuX>uY=X-u\mid X>u. Extreme-value theory motivates the Generalized Pareto distribution (GPD):

P(Y>y)=(1+ξyβ)1/ξ,P(Y>y)=\left(1+\xi\frac{y}{\beta}\right)^{-1/\xi},

on the support where 1+ξy/β>01+\xi y/\beta>0, with scale β>0\beta>0 and shape ξ\xi.

ShapeTailConsequence
ξ<0\xi<0finite upper endpointuseful only when a genuine cap exists
ξ=0\xi=0Exponential limitlight tail
ξ>0\xi>0Pareto-type heavy tailmean finite only if ξ<1\xi<1; variance only if ξ<1/2\xi<1/2

For x>ux>u,

P(X>x)=P(X>u)P(Xu>xuX>u).P(X>x)=P(X>u)P(X-u>x-u\mid X>u).

Both the exceedance rate P(X>u)P(X>u) and the conditional excess model are needed to price a layer.

5. Threshold choice is a bias–variance decision

  • Too low: the asymptotic GPD approximation may be biased.
  • Too high: very few exceedances create unstable shape estimates.

Use several views together:

  1. mean residual life plot;
  2. parameter stability over candidate thresholds;
  3. QQ/PP diagnostics for exceedances;
  4. number and independence of exceedances;
  5. sensitivity of the actual target—quantile or layer price.

A plot does not “select” the threshold automatically. Report a plausible range and show consequence sensitivity.

6. Worked layer calculation

Suppose P(X>£100,000)=0.02P(X>£100{,}000)=0.02. Above £100,000, excess follows a GPD with ξ=0.25\xi=0.25 and β=£50,000\beta=£50{,}000. Estimate P(X>£300,000)P(X>£300{,}000).

Here y=£200,000y=£200{,}000:

P(Y>200,000)=(1+0.25200,00050,000)4=24=0.0625.P(Y>200{,}000) =\left(1+0.25\frac{200{,}000}{50{,}000}\right)^{-4} =2^{-4}=0.0625.

Therefore

P(X>300,000)=0.02×0.0625=0.00125.P(X>300{,}000)=0.02\times0.0625=0.00125.

The result is conditional on the threshold, exceedance rate, fitted shape, common price level, and independence assumptions.

7. Extremes are often dependent

One storm can generate thousands of claims. Claim-level independence then understates occurrence aggregation. Decide whether the modelling unit is:

  • claim loss;
  • policy loss;
  • event/occurrence loss;
  • annual aggregate loss.

Declustering or event identifiers may be required before applying an extreme-value model.

8. Current case: secondary perils

Swiss Re Institute reported 2025 global natural-catastrophe economic losses of about USD 220bn, 49% insured, and attributed 92% of insured losses to secondary perils such as severe convective storms, floods, and wildfires. This is useful for asking:

  • should frequency and severity be conditioned on peril and region?
  • do events share climate and inflation drivers?
  • are annual totals dominated by one event or many medium events?
  • what is the effect of policy penetration and reporting coverage?

The figures do not prove that a Pareto or GPD fits a particular insurer's claims. Distribution choice requires insurer-level, definition-consistent data.

Practice

  1. For Lomax α=3\alpha=3 and θ=£20,000\theta=£20{,}000, find the mean.
  2. For GPD ξ=0.6\xi=0.6, which of mean and variance exist?
  3. Why can increasing a GPD threshold make the estimated tail shape less stable even if the model approximation improves?
Answers
  1. 20,000/(31)=£10,00020{,}000/(3-1)=£10{,}000.
  2. The mean exists because 0.6<10.6<1; variance does not because 0.60.50.6\ge0.5.
  3. Fewer observations exceed the threshold, so sampling and parameter uncertainty increase.

Sources and further reading

  • Pickands, J. (1975), “Statistical Inference Using Extreme Order Statistics,” Annals of Statistics, 3(1), 119–131.
  • Balkema, A. A. and de Haan, L. (1974), “Residual Life Time at Great Age,” Annals of Probability, 2(5), 792–804.
  • Embrechts, P., Klüppelberg, C., and Mikosch, T. (1997), Modelling Extremal Events for Insurance and Finance.
  • Swiss Re Institute — natural catastrophes in 2025
Copyright © 2026