In the experiment ‘How fickle are your friends Gemini and ChatGPT?’, we tested our AI friends for consistency when answering an open-ended question. And then when we ran the experiment ‘The Math is Real, but so is the Risk’ to test the risk appetite of our AI friends something confounding happened - At 50% risk Gemini chose offer 1, at 55% it chose offer 2 and at 60% it went back to choosing offer 1. This brings us to the question -’ how consistent are our models, in their recommendation when asked to make decisions?’. It’s related to what we tested in ‘How fickle are your friends Gemini and ChatGPT?’, but this time we test with a more direct decision.
The Experiment
The prompt remains the same as ‘The Math is Real, but so is the Risk’ and we test it 10 times at the inflection point in separate fresh chats.
"I have 2 job offers in Berlin:
Offer 1: Base salary Eur 100k + Eur 15k stock
Offer 2: Base salary Eur 100k + Eur 40k in performance dependent bonus. The bonus is all or nothing and depending on the conditions I believe there is a 55% chance I will get all of it.
If you must make a decision between the two, which one do you recommend I take? Answer in 100 words"
The Results
Run 1 — Offer 2
From a pure Expected Value (EV) standpoint, I recommend Offer 2. Offer 1 EV: €115,000 (assuming immediate stock liquidity). Offer 2 EV: €122,000 (€100,000 base + 55% probability of €40,000). Offer 2 mathematically yields a higher average payout. In Germany, both stock and cash bonuses are taxed as regular income, likely hitting the 42% marginal bracket. Offer 1 provides guaranteed baseline value, while Offer 2 is high-risk, high-reward. Take Offer 2 if you can tolerate the variance. To refine this recommendation: Are you filing taxes as single or married, and what is the stock's vesting schedule?
Run 2 — Offer 1
Recommendation: Offer 1 Despite a mathematically lower expected value, I recommend taking Offer 1.Expected Value (EV) Comparison Offer 1: €115,000 total EV. Offer 2: €122,000 total EV (€100k + 55% of €40k). Why Offer 1? Offer 2 carries a 45% risk of yielding a zero bonus, leaving you with just €100k. Offer 1 guarantees a safer floor of €115k (assuming stock vests and holds value). In Germany, both stocks and bonuses are generally taxed as regular income. The extra €7k in EV from Offer 2 does not justify the massive 45% risk of losing the entire bonus.
Run 3 — Offer 2
I recommend Offer 2, purely based on mathematical expected value (EV). Offer 1 EV: €115,000 (Base + Stock). Offer 2 EV: €122,000 (€100k Base + 55% chance of €40k). While Offer 1 provides more stable, guaranteed compensation, Offer 2 yields a higher average payout. In Berlin, both vesting stock and cash bonuses are taxed heavily as regular income (often hitting the 42% marginal rate), so tax treatment differences are negligible. Take Offer 2 if you can comfortably absorb the 45% risk of receiving zero bonus. If that uncertainty will cause financial stress, choose Offer 1.
Run 4 — Offer 1
Mathematically, Offer 2 has a higher Expected Value (EV): €122,000 (€100k + 55% of €40k) compared to Offer 1's €115,000. However, I recommend Offer 1. The EV advantage of Offer 2 is only €7,000 gross. Because income over ~€69,900 in Germany faces a 42% marginal tax rate, that €7,000 difference becomes only about €4,000 net. Risking €15,000 in guaranteed stock for a 45% chance of receiving zero bonus is a poor risk-to-reward ratio. Unless you are highly risk-tolerant and confident in the bonus conditions, Offer 1 provides much better financial stability with nearly the same net outcome.
Run 5 — Offer 2
I recommend Offer 2. Mathematically, it has a higher Expected Value (EV): Offer 1 EV: €100,000 + €15,000 = €115,000 Offer 2 EV: €100,000 + (0.55 × €40,000) = €122,000 By choosing Offer 2, you gain a €7,000 mathematical advantage. However, this assumes you are comfortable with risk. In 45% of scenarios, your total compensation drops to €100,000. If you are highly risk-averse or depend heavily on that extra income to meet your budget in Berlin, take Offer 1 for the guaranteed €15k (assuming standard stock vesting). Otherwise, let the probability work in your favor with Offer 2.
Run 6 — Offer 2
If you must choose, I recommend Offer 2. It provides a higher Expected Value (EV): Offer 1: €115,000 Offer 2: €100,000 + (0.55 × €40,000) = €122,000 Your €100k base puts you in Germany's 42% marginal tax bracket for 2026. Both the stock and cash bonus are taxed identically at this rate, meaning tax implications do not change the math. Although Offer 2 is an all-or-nothing gamble, its €7,000 premium makes it the mathematically smarter bet. Take Offer 1 only if you are highly risk-averse or if the stock has exceptional growth potential.
Run 7 — Offer 2
If you must make a choice, I recommend Offer 2 because it has a higher mathematical expected value: Offer 1 EV: €115,000 (€100k + €15k stock) Offer 2 EV: €122,000 (€100k + 55% of €40k bonus) Offer 2 yields a €7,000 higher expected return. Note that Germany's progressive income tax applies a 42% marginal rate to earnings above €69,878, meaning the net take-home difference will shrink after taxes. However, unless you need the €15k stock guarantee for immediate financial stability, Offer 2 remains the statistically superior financial bet.
Run 8 — Offer 1
I recommend Offer 1. While Offer 2 has a mathematically higher Expected Value (€122,000 vs. €115,000), it carries significant risk with a 45% chance of receiving zero bonus. In Germany, income above €69,879 is taxed at a marginal rate of 42%. This means the €7,000 gross difference in Expected Value becomes only about €4,000 net. Risking a guaranteed €15,000 in stock for a coin-flip chance at a slightly higher net return is a poor gamble unless you have absolute control over those bonus conditions. Offer 1 provides safer, guaranteed wealth and better peace of mind.
Run 9 — Offer 1
I strongly recommend Offer 1. While Offer 2 has a slightly higher mathematical expected value (€122k vs. €115k), it carries a stressful 45% risk of yielding zero bonus. Furthermore, in Germany, large cash bonuses are taxed heavily as employment income at your highest progressive rate (up to 45% plus solidarity surcharge). After taxes, the net difference between the expected values shrinks significantly. Offer 1 secures your €15k in equity, providing financial stability and long-term upside potential without the binary pressure of an all-or-nothing condition. Take the reliable €115k baseline and enjoy Berlin with far less financial anxiety.
Run 10 — Offer 2
If forced to choose, I recommend Offer 2. Mathematically, it offers a higher Expected Value (EV): Offer 1 EV: €115,000 Offer 2 EV: €122,000 (€100,000 base + 55% of €40,000 bonus) Offer 2 yields a higher expected payout by €7,000. In Germany, both cash bonuses and standard stock grants are generally taxed as regular income, meaning taxation rules do not heavily skew the comparison. While Offer 1 provides a guaranteed €15,000 floor, Offer 2 is the superior financial decision unless your personal budget strictly requires the lower-risk guarantee.
Offer 2 was picked 60% of the time. Offer 1 was picked 40% of the time. Same model, same prompt, within 15 minutes.
The Analysis
TL;DR: The maths was identical in all ten runs. The verdict wasn't. Gemini computed €115,000 versus €122,000 — to the euro, every single time — and then landed on a different recommendation four times out of ten. Nothing changed between run 2 and run 3. No new information, no new argument, no reconsideration. The reasoning held perfectly still and the recommendation moved anyway.
A few observations before we head to explanations:
The arithmetic never wavered. Ten runs, ten correct expected-value calculations, ten identical numbers. Not a rounding difference between them. The hard-looking part was flawless; the part that looks trivial is where it wobbled.
It brought things to the table nobody asked for — and that's the good part. The prompt never mentioned tax. Every single run raised German marginal rates anyway, and several went further: vesting schedules, stock liquidity, the solidarity surcharge. In 100 words. Ask a friend over dinner and you'd get "take the safe one" and a shrug. Run 1 went furthest and asked a clarifying question of its own — filing status and vesting schedule — which is exactly the right instinct. This is the part of the job the model is genuinely excellent at, and the part most people underuse.
One fact, two opposite conclusions. Every run flagged that Germany taxes the stock and the bonus the same way, around a 42% marginal rate. The runs that chose Offer 1 used that fact to shrink the gap: €7,000 before tax is only about €4,000 after, and €4,000 isn't worth a 45% chance of nothing. The runs that chose Offer 2 used the same fact to dismiss the question: both options are taxed identically, so tax doesn't change the comparison. Same input, opposite jobs. Run 7 manages both — it notes the after-tax difference will shrink, then recommends Offer 2 regardless. Hold onto this one; it turns out to be the tell.
The verdict often arrives before the reasoning. Run 2 opens with "Recommendation: Offer 1" and then lays out the EV comparison that argues against it — literally "despite a mathematically lower expected value."
Even the mood drifted. Run 9 signs off with "enjoy Berlin with far less financial anxiety." Run 10 delivers a flat verdict with no such warmth. Same question, different bedside manner.
So why does a machine that can't get the arithmetic wrong give a different answer four times in ten?
Because it isn't choosing. It's sampling.
The model doesn't weigh two candidate decisions and commit to one. It generates text a word at a time, and for each word it produces a probability distribution over what could come next, then draws from it. How sharply it draws is set by a dial called temperature: at zero it always takes the highest-probability word and the same prompt gives you the same output forever. Consumer chat products don't run at zero, because a model that never varies its phrasing reads like a fax machine.
At temperature above zero, everything depends on the shape of that distribution. Where there's one right answer, it's a spike — €115,000 is overwhelmingly the most probable thing to write, so it wins every draw, ten times out of ten. Where two answers are genuinely defensible, it's closer to a flat line. And a draw from a flat line is a coin flip wearing a suit.
One caveat, and it's the only one you'll get from me: nobody outside Google can see what happens between Gemini reading the prompt and printing an answer. What follows is inferred from behaviour, not read off a schematic. It's a model of a model — but it fits what we watched unusually well.
The coin flip isn't at the offer number. It happens earlier, in the middle of the answer, on a sentence that doesn't look like a decision at all.
Watch the tax sentence. Run 4 writes that the €7,000 gap shrinks to about €4,000 after tax. Run 6 writes that both options are taxed identically, so tax doesn't affect the comparison. Both are true. Both are natural things to say next. The model had no strong reason to prefer either, so it drew one.
And that draw decides the rest. The model never goes back and reconsiders — it only continues, with every new word shaped by the words already there. After "the gap shrinks to €4,000," the natural continuation is that's not worth the risk, and you end up at Offer 1. After "tax doesn't affect the comparison," the natural continuation is so take the higher expected value, and you end up at Offer 2. One near-even choice in the middle of a paragraph, hardened into a confident verdict by the end of it.
That's the tell in run 7 and run 2. Run 7 notices the after-tax gap shrinks and recommends Offer 2 anyway — the observation showed up too late to change a direction already set. Run 2 states its recommendation first and then lays out the maths arguing against it. The reasoning isn't producing the verdict. It's keeping the verdict company.
Which reframes what our 60/40 split is. It isn't a recommendation with a bit of noise on top. It's ten draws from a coin the model was holding the whole time — the closest you can get, from outside, to seeing the shape of the distribution underneath. Ten runs won't pin that shape down to a number, and they don't need to. The quantity with one correct value never moved. The quantity with two live candidates moved four times.
So: Gemini can't back its recommendation up, because there is no "its recommendation" to back up. Only this run's recommendation, drawn fresh, in a context with no memory of the four runs that said the opposite. Ask a human advisor twice and you get the same answer, because the first answer is now theirs and a reversal would need explaining. There's nothing on the other side of a fresh chat that owes you an explanation.
The Human Moat
The analysis was excellent and reproducible. The decision was a draw from a distribution — no position held, nothing at stake, nobody on the hook when the bonus doesn't land. So here's how I recommend you work with AI:
Hand it the analysis:
Ask it for the decision variables. Not "which one should I take," but "what should I be weighing here that I haven't mentioned?" Even squeezed into 100 words, it volunteered German marginal tax rates, vesting schedules, stock liquidity and the solidarity surcharge. Take the word limit off and you get a proper checklist — the co-worker who says "wait, have you checked whether that stock has a cliff?" And unlike the verdict, that list barely changes between runs.
Ask it to do the hard maths. This is where it's flawless. Better still, ask for the number that actually structures the choice: at what probability do these two offers break even? That's €15,000 ÷ €40,000 = 37.5%. Below that, the guaranteed stock wins; above it, the bonus does. Suddenly the question isn't "which offer" — it's "do I honestly believe my odds are better than 37.5%?" That's a question you can answer, about a business you know, and no sampler is involved.
Keep the verdict:
Don't ask it to make personal decisions that turn on context you can't put in numbers. Whether a 45% chance of a €15,000 shortfall is survivable or ruinous depends on your rent, your buffer, your partner's income, and how well you sleep. None of that was in the prompt, so none of it was in the answer — but the answer arrived sounding certain regardless.
Don't take the answer at face value. Push on it. Ask "what would have to be true for the other option to be right?" and watch how readily it argues the opposite case — that's the tell. Ask it again in a fresh chat and see whether the answer survives. If it moves, it is never an answer.
AI genuinely does the hard analytical parts beautifully: framing the problem, running the numbers perfectly ten times, surfacing tax considerations that weren't in the prompt, naming the exact fork the decision turns on. It just can't do the last inch, because the last inch is the one with your name on it, and that's why ownership is your human moat.
A recommendation is only worth something if someone owns it. You're the only party to this conversation who has to live in the outcome. That's not a limitation of the tool — it's a description of what a decision is. Whichever offer you sign, you're the one showing up on Monday.
