You already read the research overview. Now you want the layer under it: what the model actually does to a distribution.

The discard

Why does Gaussian VaR discard the part you care about?

Because it assumes the shape before it sees the data. A Gaussian fixes thin tails by construction, so crash density beyond |r| > 5% is discarded at modeling time — never represented. On the S&P 500 the real return kurtosis is 16.70; a Gaussian/GBM baseline collapses to 2.90 (Fukunishi, Qiu & Takahashi, 2026). That gap is the unmodeled tail — the mass your VaR pretends cannot exist.

The cost

What does that missing mass cost a concentrated book?

It costs you the exact scenario you hold the position against. Real S&P 500 Expected Shortfall at the 1% level is -5.28%; the Gaussian baseline reports -3.16% — a 40% understatement of the average loss inside the worst 1% of sessions. On a position at 60–90% of net worth, that error is the difference between a hedge sized for the crash and one sized for the brochure.

The true tail

What does reading the true tail look like?

A failure surface with the crash density left in. You stop assuming the shape and reconstruct it from twenty years of data — 2008 and 2020 included — so the extreme quantile is observed, not back-solved from a formula.

How it works

How it works

The Sandbox Engine — diffusion sampling over fat-tailed returns

A diffusion model learns the distribution backwards. The forward process adds Gaussian noise to real returns in steps; the engine trains a network to reverse it, denoising noise back into a return. Because that reverse process — its variance learned, not fixed — is fit to the data, it allocates probability mass directly to the heavy tails instead of regressing to the mean. The Sandbox Engine runs this denoising over 50,000 paths against your allocation, the same surface that prices convex tail-risk options downstream.

Model: tail-aware diffusion (learned variance, v-prediction, T=500 steps) · Data: 20y S&P/FTSE/TOPIX · Coverage: 5%/1% VaR + ES · Published false-comfort rate: 4.2%

This is the recovery, measured: the tail-aware model rebuilds S&P 500 kurtosis to 13.38 against the real 16.70, where Gaussian sits at 2.90 (Fukunishi, Qiu & Takahashi, 2026). The 4.2% false-comfort rate is the share of runs still reporting a tail smaller than realized — printed because a model that hides its error is the failure mode.

The objection

“Isn’t diffusion just a heavier Monte Carlo?”

No — and the distinction is where your money lives. Classical Monte Carlo samples from a distribution you specify; specify Gaussian and you get 50,000 thin-tailed paths and a confident lie. Diffusion learns the distribution’s shape — the tail decay included — before it samples, so the crash is in the generated paths because it was in the data, not because you added it. We test the same logic against a language model’s read of the tail, then surface the recovered numbers in your portfolio dashboard.

3 fields · 48-hour document · no call, no sequence.

Frequently asked questions

How does a diffusion model generate synthetic stock returns?

It learns the return distribution backwards. A forward process adds Gaussian noise to twenty years of real returns in steps; a trained network then reverses it, denoising pure noise into a return whose shape — fat tails included — was fit to the data rather than assumed. Fukunishi, Qiu & Takahashi (2026) showed this reconstructs heavy-tailed structure that parametric models cannot represent.

Why does diffusion sampling beat Monte Carlo for tail risk?

Classical Monte Carlo samples from a distribution you specify in advance, so a Gaussian assumption produces 50,000 thin-tailed paths and a confident understatement of the crash. Diffusion learns the distribution’s shape before it samples, so the tail appears in the generated paths because it was in the historical data, not because you injected it.

What does kurtosis measure in this context?

Kurtosis measures how much probability mass sits in the extreme tails — how crash-prone the distribution is. Real S&P 500 return kurtosis is 16.70; the tail-aware diffusion model rebuilds it to 13.38, while a Gaussian baseline collapses to 2.90 (Fukunishi, Qiu & Takahashi, 2026). The gap between 2.90 and 16.70 is the unmodeled crash mass.

What is the model’s main limitation?

It still understates the tail on a measurable share of runs. The published false-comfort rate is 4.2% — the fraction of runs reporting a tail smaller than realized — and recovered kurtosis of 13.38 remains below the real 16.70. We print these failure rates because a model that hides its error is the failure mode.

How many simulation paths does the engine run?

The Sandbox Engine runs the denoising process over 50,000 paths against your specific allocation, fit to twenty years of S&P, FTSE and TOPIX data including 2008 and 2020. That sample size resolves the 5% and 1% VaR and Expected Shortfall quantiles the tail decision turns on.

Entail Capital — The Risk Atelier

The crash is a distribution.
We compute its shape.

48-hour turnaround · a document, not a pitch · if your tail is smaller than you feared, the document will say so.

Run my diagnostic →