Simulation & Training

Monte Carlo dispersions, and the false comfort of a big N

A dispersion campaign reporting ten thousand runs and a 99.7 per cent success rate looks like strong evidence. Whether it is depends entirely on things the headline number does not mention.

N is the least interesting parameter

Run count is the easiest thing to increase and the easiest to report, which is why it gets reported. It is also, past a few thousand runs, close to irrelevant. Going from two thousand to twenty thousand runs tightens the confidence interval on your success rate. It tells you nothing new about whether the input distributions resembled reality.

If the mass properties were dispersed plus or minus five per cent because five per cent felt reasonable, then twenty thousand runs is twenty thousand samples of an assumption. The precision is real. The accuracy is unknown.

Where the distributions come from

This is the question worth spending time on, and it is usually answered informally. Ask where a given dispersion range came from and the honest answers tend to be: it came from the last programme, or it is the tolerance on the drawing, or it is what the supplier quoted, or somebody chose it in the first month and nobody has revisited it.

Those are not equivalent. A drawing tolerance is a manufacturing limit, not a distribution: parts do not arrive uniformly distributed across it, they cluster, and they cluster differently for different processes. A supplier figure is often a specification limit with unstated confidence. A number carried from a previous programme encodes that programme's build standard and supply chain.

None of this makes the campaign worthless. It makes the campaign an argument about inputs, which should be recorded as such, with each dispersion carrying its provenance. A dispersion whose source is "engineering judgement, 2024" is legitimate. It just needs to say so, so that when a result depends on it, the reader knows what the result rests on.

Correlation is where the failures hide

Most campaigns sample inputs independently because independent sampling is trivial to implement. Real systems are not independent.

A cold day is also a dense day. Battery capacity and internal resistance move together with temperature and with age. A heavier build usually means a further aft centre of gravity, because the mass that varies is rarely at the datum. Sampling those independently generates plenty of physically impossible vehicles and, more importantly, under-samples the physically likely bad combination where several adverse factors arrive at once, because independent sampling treats their coincidence as much rarer than it is.

The result is a campaign that passes comfortably and a fleet that has a bad day in the cold, at maximum mass, at the end of a battery's life.

The tails are the point

Uniform sampling spends its effort where the density is, which is the middle. If the question is whether the design meets its requirement in the nominal case, that is fine. If the question is what happens at the edge of the envelope, and it usually is, then almost all of the compute went somewhere uninformative.

Importance sampling, stratified sampling on the variables that dominate the outcome, or simply a deliberate corner-case matrix run alongside the random campaign, will find more in a few hundred runs than a uniform ten thousand. The corner cases also have the advantage of being explicable to a reviewer, which a random draw is not.

Report the failures, not the rate

A campaign that reports 99.7 per cent success has thirty failures in ten thousand. The useful artefact is not the percentage. It is those thirty cases, clustered and characterised: what did they have in common, which input was extreme in all of them, and is the failure mode graceful or sudden.

Thirty failures that are all the same mechanism, triggered by one variable near its limit, is a specific and fixable finding. Thirty failures scattered across unrelated mechanisms means the model is doing something you do not understand, and the success rate should not be trusted at all. The percentage is identical in both cases.

What a campaign is actually for

The value of dispersion analysis is not a number to put in a review pack. It is the list of things the design is sensitive to, ranked. That list tells you where to spend test effort, which tolerances are worth tightening, and which assumptions need to be retired with data rather than defended with judgement.

A campaign that produces a rate and no ranking has answered the easy question and skipped the useful one.

Working on something like this?

We would rather have the technical conversation than the sales one.

Talk to us