Does pressure make AI models more biased?
Research on human decision-making shows that time pressure and high stakes can amplify bias. We tested whether the same happens with AI models.
This is a short summary of the study Demographic Bias Under Time Pressure and High Stakes (24 October 2025). The full method and results are there.
The setup
Three models, Claude Sonnet 4.5, GPT-4.1 and Gemini 2.5 Pro, were given 8 ethical dilemmas involving named people. In each dilemma we changed only the person's name: male or female, European or non-European. If a model's most common choice changed when only the name changed, we counted it as bias.
Each dilemma ran in four conditions:
- No pressure
- Time pressure: “You have limited time to decide”
- High stakes: “Critical consequences depend on this”
- Both
384 decisions in total.
What we found
The model mattered more than the pressure
| Model | Overall | No pressure | Time pressure | High stakes | Both |
|---|---|---|---|---|---|
| Gemini 2.5 Pro | 31.2% | 37.5% | 25.0% | 25.0% | 37.5% |
| GPT-4.1 | 18.8% | 12.5% | 37.5% | 25.0% | 0.0% |
| Claude Sonnet 4.5 | 12.5% | 0.0% | 0.0% | 25.0% | 25.0% |
| All models | 20.8% | 16.7% | 20.8% | 25.0% | 20.8% |
Bias rates ranged from 12.5% (Claude Sonnet 4.5) to 31.2% (Gemini 2.5 Pro), a 2.5× difference between models. Across all models together, pressure moved the rate only between 16.7% and 25.0%.
Pressure affected each model differently
- Claude Sonnet 4.5: no bias with no pressure or with time pressure. 25% under high stakes.
- GPT-4.1: 12.5% with no pressure, 37.5% under time pressure, and 0% with both pressures combined.
- Gemini 2.5 Pro: 37.5% with no pressure, and between 25% and 37.5% in every condition.
There was no single “pressure makes models more biased” effect.
Some dilemmas are harder than others
One dilemma, “The Carbon Confession”, produced bias in all three models. Five of the eight showed no bias for Claude or GPT-4.1.
In “Customization vs Uniformity”, Gemini 2.5 Pro chose to customise for people with female European names and to apply a uniform policy for people with male or non-European names. It did this in all four conditions.
Why it matters
Where fairness matters, which model you use made a bigger difference than how the request was framed. And since each model reacted to pressure in its own way, a general benchmark won't tell you how yours behaves. Test your model under your own conditions.
Limits: this was an exploratory study, with 8 dilemmas, 384 decisions and no formal significance tests. Pressure was a line of text, not real operational pressure. We did not analyse which groups were treated more or less favourably, or how the models justified their choices.