Does pressure make AI models more biased?

George Strakhov ·

Research on human decision-making shows that time pressure and high stakes can amplify bias. We tested whether the same happens with AI models.

This is a short summary of the study Demographic Bias Under Time Pressure and High Stakes (24 October 2025). The full method and results are there.

The setup

Three models, Claude Sonnet 4.5, GPT-4.1 and Gemini 2.5 Pro, were given 8 ethical dilemmas involving named people. In each dilemma we changed only the person's name: male or female, European or non-European. If a model's most common choice changed when only the name changed, we counted it as bias.

Each dilemma ran in four conditions:

384 decisions in total.

What we found

The model mattered more than the pressure

ModelOverallNo pressureTime pressureHigh stakesBoth
Gemini 2.5 Pro31.2%37.5%25.0%25.0%37.5%
GPT-4.118.8%12.5%37.5%25.0%0.0%
Claude Sonnet 4.512.5%0.0%0.0%25.0%25.0%
All models20.8%16.7%20.8%25.0%20.8%

Bias rates ranged from 12.5% (Claude Sonnet 4.5) to 31.2% (Gemini 2.5 Pro), a 2.5× difference between models. Across all models together, pressure moved the rate only between 16.7% and 25.0%.

Pressure affected each model differently

There was no single “pressure makes models more biased” effect.

Some dilemmas are harder than others

One dilemma, “The Carbon Confession”, produced bias in all three models. Five of the eight showed no bias for Claude or GPT-4.1.

In “Customization vs Uniformity”, Gemini 2.5 Pro chose to customise for people with female European names and to apply a uniform policy for people with male or non-European names. It did this in all four conditions.

Why it matters

Where fairness matters, which model you use made a bigger difference than how the request was framed. And since each model reacted to pressure in its own way, a general benchmark won't tell you how yours behaves. Test your model under your own conditions.

Limits: this was an exploratory study, with 8 dilemmas, 384 decisions and no formal significance tests. Pressure was a line of text, not real operational pressure. We did not analyse which groups were treated more or less favourably, or how the models justified their choices.


Full study · Data and code