How Actuaries Use Probability to Price Insurance

Insurers don't predict the future, they predict averages. A visual, intuition-first walk through expected value, the law of large numbers, and risk pooling, the three ideas behind every insurance premium you've ever paid.

By Petrus Sheya

July 27, 2026 · 7 min read

How does an insurance company know what to charge you, when nobody knows what's going to happen to you?

They don't know if your house will flood next year. They don't know if you'll total your car. Nobody does. But somehow, every year, insurance companies price these unknowable events down to the dollar, collect premiums, and stay in business.

Here's the trick: they're not predicting you. They're predicting a crowd. And a crowd, it turns out, is far more predictable than any one person in it.


The weighted coin

Picture your insurance policy as a coin flip. Not a fair one, a weighted one.

Flip it, and there's some probability pp it lands on "claim", and probability 1p1-p it lands on "no claim". If it lands on claim, the insurer pays out some amount LL. If not, they pay nothing.

You don't know which way your coin lands. Neither does the insurer. But here's the thing: if you had to guess the average cost of insuring you, before the coin is even flipped, there's a very natural answer.

Multiply each outcome by how likely it is, and add them up:

E[loss]=pL+(1p)0=pLE[\text{loss}] = p \cdot L + (1-p) \cdot 0 = p \cdot L

This is called the expected value, and it's not a prediction of what will happen. It's the average of what would happen if you replayed the coin flip thousands of times. Think of it as a balancing point, like a seesaw with a small weight sitting at "no claim, zero cost" and a bigger or smaller weight sitting at "claim, cost L", depending on how likely each outcome is.

Drag the claim weight along the line — the fair premium is wherever the lever balances.

E[loss] = $90085%no claim, $015%claim, $6000$0$20,000
Claim probability15%
Payout if claim$6,000
Fair premium$900

Drag the claim weight along the line, or slide the probability. Notice the fulcrum, that's the fair premium, and it always sits closer to whichever side is heavier. Crank the probability up and the balance point slides right toward the payout. Drop it near zero and the balance point collapses back toward $0.

That number, pLp \cdot L, is the actuary's starting point. It's called the pure premium. It's exactly what the insurer expects to pay out, per policy, on average. Everything else in the premium (overhead, profit, a buffer for bad luck) gets added on top of this number. But the pure premium is the foundation, and it's just a weighted average.


One person is a mystery. A thousand people aren't

Here's the part that should bother you. Expected value is an average. Averages are only meaningful over many repeats. But your insurance policy only "happens" once. You either file a claim this year or you don't. So how does an average help price a single, unrepeated event?

It doesn't, for you alone. It works because you're never priced alone.

An insurer isn't betting on your coin. They're holding a jar of ten thousand coins, all weighted the same way, and they only care about the total. And here's the remarkable fact: even though each individual coin flip is completely unpredictable, the total number of "claim" outcomes across ten thousand flips is almost eerily predictable. This is the law of large numbers, and it's the entire reason insurance is a viable business.

Each tick is one more random policyholder — watch the noisy average settle onto the true expected cost.

true E[loss] = $5000600 policies
Policies so far0
Running average$0
Error vs true mean

Hit play. Each tick simulates one more random policyholder, claim or no claim, and plots the running average cost per person so far. Early on, the line jumps around wildly, one unlucky claim can swing it hard. But watch what happens as the pool grows: the swings get smaller and smaller, and the average locks onto the true expected value, the dashed line, almost exactly.

Try dragging the probability slider to a new value and replaying. The destination changes, but the pattern doesn't: individual randomness always gets tamed by scale.

This is the opposite of gambling. A gambler wants variance, the chance of a huge win. An insurer wants the absence of variance. They're not trying to win big on any one policy. They're trying to make the total, across everyone, as boring and predictable as possible.


Why size is the whole game

So a bigger pool makes the average more predictable. But how much more, and how fast?

If you have NN independent policies, each with claim probability pp, the number of claims you'll actually see follows a well-known shape, a bell curve centered at NpN \cdot p. What changes as NN grows isn't the center, it's the width. The standard deviation of the claim rate shrinks like:

σ=p(1p)N\sigma = \sqrt{\frac{p(1-p)}{N}}

Notice the NN is inside a square root in the denominator. Quadruple your pool size, and the uncertainty only halves. That's a real cost, but it's still a guarantee: the relative uncertainty always shrinks as the pool grows, no matter what pp is.

Same 8% claim rate every time — only the pool size changes. Watch the uncertainty collapse.

0.0%8.0%33.8%
Pool size501
90% range6.010.0%
Relative uncertainty±15.2%

Slide the pool size up. The claim rate always centers on 8%, that never moves. What collapses is the spread around it. At a pool of 10, the actual claim rate could reasonably land anywhere from 0% to 25%, wildly unpredictable. At a pool of 100,000, it's pinned to within a fraction of a percent.

This is why insurance barely works for tiny mutual aid groups and works extremely well for national insurers. It's not a different kind of math, it's the same math with a bigger NN. Scale is the product.


But not everyone is the same risk

There's a catch we've been quietly ignoring. We've been assuming every policyholder has the same probability pp. In reality, a 20-year-old driver and a 50-year-old driver don't have the same crash probability. A house on a floodplain and a house on a hill don't have the same flood probability.

If an insurer charges one flat price to everyone, based on the average risk across the whole mixed pool, something quietly breaks: low-risk customers end up subsidizing high-risk ones.

Toggle the pricing model and watch what low-risk customers pay versus what they actually cost.

true costpricedLow risktrue costpricedHigh risk
Low-risk pays$348
Low-risk true cost$120
Total subsidy to high-risk$159,600

Start with "Flat price for everyone" and slide the high-risk share up. Watch the gap between what low-risk customers pay and what they actually cost, that gap is money flowing from safe customers to risky ones, whether either group asked for it or not.

Now click "Risk-adjusted pricing." Each group gets priced at its own true expected cost. The subsidy disappears.

This matters for a reason beyond fairness. If safe customers can tell they're overpaying, some of them will leave, buy a cheaper policy elsewhere, or self-insure. When they leave, the pool that remains skews riskier, which pushes the flat price up even further, which drives out the next safest layer of customers. Round and round it goes. Actuaries call this adverse selection, and it's the reason insurers spend so much effort classifying risk instead of just charging everyone the pool average. Accurate pricing per risk class isn't just about fairness, it's what keeps the whole pool from unraveling.


Putting it together

Every premium you've ever paid is built from the same three ideas, stacked on top of each other.

Expected value turns an unknown, one-time event into a single number, the probability-weighted average cost. That's the pure premium.

The law of large numbers is what makes that average trustworthy. One person's outcome is noise. Ten thousand people's average is signal.

Risk classification makes sure that number is computed on people who actually resemble each other, so the average one group pays reflects the risk that group actually carries, not someone else's.

None of this requires the insurer to know what will happen to you. It only requires them to know, with high confidence, what will happen to enough people like you, added together. That's the whole business: turning unpredictable individuals into a predictable crowd, and charging accordingly.


All simulations on this page run entirely in the browser. The risk pool distribution uses a normal approximation to the binomial, valid here since NpNp and N(1p)N(1-p) stay comfortably above 5 across the slider's range.