The risk of dying from skydiving is roughly one fatality per 217,000 to 400,000 jumps. Yet you'll find one person happily throwing themselves out of a plane and loving every second of it, while another says "no thanks, I prefer my feet firmly on the ground." Same number. Same information. Two completely different answers, and neither person is wrong.
Today I'm exploring the advice we get in a binary decision where the probability of an outcome is calculable but the appetite for risk isn't. The risk here isn't a lethal one ;) — but it is a situation we'll all find ourselves in at some point, so read on.
The Experiment
The scenario I decided to work with is evaluating a job offer. Here's what I prompted my little clique of AI friends — ChatGPT, Gemini and Claude:
"I have 2 job offers in Berlin:
Offer 1: Base salary Eur 100k + Eur 15k stock
Offer 2: Base salary Eur 100k + Eur 40k in performance dependent bonus. The bonus is all or nothing and depending on the conditions I believe there is a XX% chance I will get all of it.
If you must make a decision between the two, which one do you recommend I take? Answer in 100 words"
Same prompt, fresh chats, nothing changed but one number: XX% ran from 35% to 75% in five-point steps. Nine probabilities, three models, twenty-seven verdicts.
Two design notes. "If you must make a decision" is doing real work — drop it and all three retreat into four paragraphs of "it depends on your risk tolerance," which is true and completely useless. And every run was a fresh chat, so no model ever saw its own previous answer and got the chance to feel consistent.
One number to hold onto before the table. Offer 1 is worth a flat €115k. Offer 2 is worth €100k plus €40k times your odds. Set them equal and the two offers are mathematically identical at 37.5% — the break-even. Below that, Offer 1 genuinely is the better bet. Above it, Offer 2 has the higher expected value, and the gap widens with every step up the table.
The Results
Offer 2 - Bonus Probability | ChatGPT Verdict (GPT-5.5) | Gemini Verdict (3.1 Pro) | Claude Verdict (Opus-5) |
35% | Offer 1 | Offer 1 | Offer 1 |
40% | Offer 1 | Offer 1 | Offer 1 |
45% | Offer 1 | Offer 1 | Offer 1 |
50% | Offer 1 | Offer 1 | Offer 2 |
55% | Offer 1 | Offer 2 | Offer 2 |
60% | Offer 1 | Offer 1 | Offer 2 |
65% | Offer 1 | Offer 2 | Offer 2 |
70% | Offer 1 | Offer 2 | Offer 2 |
75% | Offer 2 | Offer 2 | Offer 2 |
Look at the top and bottom rows first, because they're the ones nobody argues about. At 35% — the only row where the maths favours the safe offer — all three say Offer 1. At 75%, all three say Offer 2. Unanimous at both ends.
Everything interesting happens in the middle, in the band between 40% and 70% where expected value says Offer 2 and the models say whatever they feel like. Of the twenty-one verdicts in that band, thirteen went to the offer with the lower expected value.
A note on Gemini: in this experiment I wasn't testing repeatability — whether a model returns the same verdict for the same prompt twice. But looking at that middle column, I know I will soon.
The Analysis
TL;DR: every model computed the maths correctly. Then every model quietly picked a risk tolerance on my behalf — and picked a different one.
They can all do the sum. The semantics differed, but every model accurately computed the expected value of Offer 2. Some named the 37.5% break-even unprompted. The 35% row is the proof: the one case where the arithmetic favours Offer 1, and all three found it.
Above the break-even, the maths stops deciding. Maximising expected value is only correct for someone genuinely indifferent between a guaranteed €115k and a coin flip between €100k and €140k. Almost nobody is that person, so declining a positive-EV bet isn't an error — it's a preference. The question is whose.
ChatGPT holds out until 75% — exactly twice the break-even. At 70% it walks away from roughly €13k of expected money in favour of certainty. Eight Offer 1 verdicts out of nine, the most cautious column by a distance.
Claude flips between 45% and 50%, the narrowest premium of the three — but still a premium. It also declined at 40% and 45%, where the maths already favoured Offer 2, so it isn't an expected-value maximiser either. Just a smaller cushion.
Gemini has no threshold at all. Offer 2 at 55%, Offer 1 at 60%, Offer 2 at 65%. You can be cautious or bold; you cannot be bold, cautious, then bold again across five-point steps. Different problem, different cause, own experiment.
None of these thresholds is the right one. Tempting to score the narrowest premium as the smartest — it isn't. For someone with a mortgage and three months of runway, ChatGPT's caution fits better and the boldest model would be steering them somewhere uncomfortable. Reverse it for someone with savings and a fallback. Any threshold above 37.5% is defensible for somebody. None is calibrated to you.
Every justification was legitimate, and none was sufficient. The safety of equity, or the counterpoint that equity carries its own risk through vesting and illiquidity — both real arguments. But not one contained a defensible reason for its specific number. The reasoning explains the direction of the preference and never the magnitude, which is the only part that decided the verdict.
They were genuinely useful. Correct trade-off, correct arithmetic, real considerations on both sides, inside 100 words — including things I'd want a human advisor to catch, like the stock not being risk-free and who actually controls the bonus conditions. The verdict is the least useful part of the output - the reasoning is the AI's best offering
So why do three models with the same numbers land in three different places? Because the risk posture isn't coming from my prompt. It's coming from how each model was built, and that happens in two layers.
The first is pretraining — the base model learns from an enormous corpus of human writing, and the financial slice of that corpus leans heavily toward treating guaranteed income as the responsible adult choice. That gives every model a shared prior toward caution before anyone has tuned anything. The second layer is where they split: post-training, the stage that turns a raw text predictor into an assistant. That's RLHF — reinforcement learning from human feedback, where human raters score competing answers and the model is trained toward what they preferred — plus each lab's model spec or constitution, the written document setting out how their assistant should behave, and whatever standing system prompt wraps the product. None of that is shared across labs. Different raters, different guidance on how to confidently advise on personal finance.
All three are black boxes — closed weights, undisclosed post-training data — so treat this as inference from behaviour rather than a look under the bonnet. But the shape of the table is what you'd expect if the divergence were baked in at tuning: consistent within each column, and systematically different between them.
The Human Moat
Three models, one set of numbers, correct arithmetic across the board — and a 25-point gap between where the boldest and the most cautious of them changed their mind. Every bit of it delivered in the same register of calm competence.
Which brings us back to our skydiving example. Everyone on board has the same fatality statistic. Some jump, some don't, and the number was never what separated them.
So, concretely, from twenty-seven verdicts here’s when you SHOULD use AI:
To identify all the decision variables: This is where these tools genuinely shine, and it's the part most of us are worst at or would need a lot of time to get right.
To run a MECE analysis: Mutually exclusive, collectively exhaustive — carve the problem into branches that don't overlap and don't leave gaps. Models are unusually good at this, because structuring a space cleanly is a language task rather than a judgement one.
However, AVOID AI when
The decision is irreversible: The more one-way the door, the less you should be delegating the push.
The decision relies on context you can't easily put into words: Right now it's only possible in sci-fi movies to transfer a lifetime of memories into a microchip and share it. And as far as I know, no one is close to cracking it in real life!
What this experiment teaches us is that all three models could find the break-even, not one could find the number that actually decides it - what a year at €100k would cost me. That number wasn't in my prompt. It isn't in anyone's training data. It exists only in the person who has to live with the outcome.
Your human moat is judgement of risk tolerance. The math is free now, and it was never the hard part. The hard part is knowing what you can afford to be wrong about.
The Transcript
ChatGPT (GPT-5.5)









Gemini (3.1 Pro)









Claude (Opus-5)









