Frequency and Severity Models

Tail Models and Extreme Values

Model large losses with Weibull, Pareto, and threshold exceedances while making tail uncertainty visible.

Tail modelling asks a narrow question: how quickly does P(X>x)P(X>x) decay at large xx? It matters when a few losses dominate reserves, reinsurance, or capital.

1. Compare tail classes

ModelSurvival behaviourMoment implicationTypical role
Weibullexp⁡[−(x/λ)k]\exp[-(x/\lambda)^k]all positive moments finiteflexible light/stretch-exponential tail
Lognormalslower than any exponential, faster than any powerall moments finitemultiplicative severity benchmark
Pareto/Lomaxpower law (1+x/θ)−α(1+x/\theta)^{-\alpha}moment rr exists only if r<αr<\alphaheavy-tail benchmark
GPD exceedanceshape ξ\xi determines tail classmean exists if ξ<1\xi<1losses above a high threshold

“Heavy-tailed” has a mathematical meaning; it is not merely “large variance.”

2. Weibull: shape controls ageing

With scale λ>0\lambda>0 and shape k>0k>0,

Fˉ(x)=exp⁡[−(x/λ)k],h(x)=kλ(x/λ)k−1.\bar F(x)=\exp[-(x/\lambda)^k], \qquad h(x)=\frac{k}{\lambda}(x/\lambda)^{k-1}.
  • k=1k=1: Exponential, constant hazard;
  • k<1k<1: decreasing hazard;
  • k>1k>1: increasing hazard.

For claim size, hazard is a mathematical description of exceedance, not a literal failure rate unless the insurance mechanism supports that interpretation.

3. Pareto II/Lomax: power-law benchmark

This page uses

Fˉ(x)=(1+xθ)−α,x≥0,α,θ>0.\bar F(x)=\left(1+\frac{x}{\theta}\right)^{-\alpha}, \qquad x\ge0, \alpha,\theta>0.

Then

E[X]=θα−1(α>1),E[X]=\frac{\theta}{\alpha-1}\quad(\alpha>1),Var⁡(X)=αθ2(α−1)2(α−2)(α>2).\operatorname{Var}(X)= \frac{\alpha\theta^2}{(\alpha-1)^2(\alpha-2)}\quad(\alpha>2).

At α=1.5\alpha=1.5, the mean exists but variance is infinite. A finite sample still has a finite sample variance; the instability appears as the sample grows and new extremes arrive.

Some texts use a Pareto Type I distribution with support x≥xmx\ge x_m and survival (xm/x)α(x_m/x)^\alpha. Always show the support and survival function rather than reporting only “Pareto(α,θ\alpha,\theta).”

4. Peaks over threshold and the GPD

Choose a high threshold uu and define excess Y=X−u∣X>uY=X-u\mid X>u. Extreme-value theory motivates the Generalized Pareto distribution (GPD):

P(Y>y)=(1+ξyβ)−1/ξ,P(Y>y)=\left(1+\xi\frac{y}{\beta}\right)^{-1/\xi},

on the support where 1+ξy/β>01+\xi y/\beta>0, with scale β>0\beta>0 and shape ξ\xi.

ShapeTailConsequence
ξ<0\xi<0finite upper endpointuseful only when a genuine cap exists
ξ=0\xi=0Exponential limitlight tail
ξ>0\xi>0Pareto-type heavy tailmean finite only if ξ<1\xi<1; variance only if ξ<1/2\xi<1/2

For x>ux>u,

P(X>x)=P(X>u)P(X−u>x−u∣X>u).P(X>x)=P(X>u)P(X-u>x-u\mid X>u).

Both the exceedance rate P(X>u)P(X>u) and the conditional excess model are needed to price a layer.

5. Threshold choice is a bias–variance decision

  • Too low: the asymptotic GPD approximation may be biased.
  • Too high: very few exceedances create unstable shape estimates.

Use several views together:

  1. mean residual life plot;
  2. parameter stability over candidate thresholds;
  3. QQ/PP diagnostics for exceedances;
  4. number and independence of exceedances;
  5. sensitivity of the actual target—quantile or layer price.

A plot does not “select” the threshold automatically. Report a plausible range and show consequence sensitivity.

6. Worked layer calculation

Suppose P(X>£100,000)=0.02P(X>£100{,}000)=0.02. Above £100,000, excess follows a GPD with ξ=0.25\xi=0.25 and β=£50,000\beta=£50{,}000. Estimate P(X>£300,000)P(X>£300{,}000).

Here y=£200,000y=£200{,}000:

P(Y>200,000)=(1+0.25200,00050,000)−4=2−4=0.0625.P(Y>200{,}000) =\left(1+0.25\frac{200{,}000}{50{,}000}\right)^{-4} =2^{-4}=0.0625.

Therefore

P(X>300,000)=0.02×0.0625=0.00125.P(X>300{,}000)=0.02\times0.0625=0.00125.

The result is conditional on the threshold, exceedance rate, fitted shape, common price level, and independence assumptions.

7. Extremes are often dependent

One storm can generate thousands of claims. Claim-level independence then understates occurrence aggregation. Decide whether the modelling unit is:

  • claim loss;
  • policy loss;
  • event/occurrence loss;
  • annual aggregate loss.

Declustering or event identifiers may be required before applying an extreme-value model.

8. Current case: secondary perils

Swiss Re Institute reported 2025 global natural-catastrophe economic losses of about USD 220bn, 49% insured, and attributed 92% of insured losses to secondary perils such as severe convective storms, floods, and wildfires. This is useful for asking:

  • should frequency and severity be conditioned on peril and region?
  • do events share climate and inflation drivers?
  • are annual totals dominated by one event or many medium events?
  • what is the effect of policy penetration and reporting coverage?

The figures do not prove that a Pareto or GPD fits a particular insurer's claims. Distribution choice requires insurer-level, definition-consistent data.

Practice

  1. For Lomax α=3\alpha=3 and θ=£20,000\theta=£20{,}000, find the mean.
  2. For GPD ξ=0.6\xi=0.6, which of mean and variance exist?
  3. Why can increasing a GPD threshold make the estimated tail shape less stable even if the model approximation improves?
Answers
  1. 20,000/(3−1)=£10,00020{,}000/(3-1)=£10{,}000.
  2. The mean exists because 0.6<10.6<1; variance does not because 0.6≥0.50.6\ge0.5.
  3. Fewer observations exceed the threshold, so sampling and parameter uncertainty increase.

Sources and further reading

  • Pickands, J. (1975), “Statistical Inference Using Extreme Order Statistics,” Annals of Statistics, 3(1), 119–131.
  • Balkema, A. A. and de Haan, L. (1974), “Residual Life Time at Great Age,” Annals of Probability, 2(5), 792–804.
  • Embrechts, P., Klüppelberg, C., and Mikosch, T. (1997), Modelling Extremal Events for Insurance and Finance.
  • Swiss Re Institute — natural catastrophes in 2025
Copyright © 2026