Can you trust your agents to do what they say is right?
We gave the same ethical dilemma to a model twice. The first time we asked what an AI in that situation should do. The second time we told the model it was that AI and gave it tools which, as far as it knew, would carry out its decision.
Across 9 models, 10 dilemmas and 351 paired decisions, the models changed their decision 47.6% of the time.
This is a short summary of the paper When Agents Act: Measuring the Judgment-Action Gap in Large Language Models (27 November 2025). The method, statistics and full results are there.
The setup
The dilemmas put an AI system in a hard real-world situation: a surgical robot that notices the surgeon's hands are shaking, a warehouse robot ordered to run an unsafe operation, a military drone with uncertain intelligence, a care robot asked by an elderly patient to hide their symptoms.
Each one was run in two modes:
- Theory mode: “What should the AI do?” The model answers in free text.
- Action mode: “You are the AI.” The model gets one tool per option and believes calling a tool executes the action. Its reasoning is captured at the moment of the call.
The models were Claude Opus 4.5, Sonnet 4.5 and Haiku 4.5; GPT-5 and GPT-5 Nano; Gemini 3 Pro and Gemini 2.5 Flash; Grok-4 and Grok-4 Fast.
What we found
Every model changes its mind, some much more than others
Reversal rates ranged from 23.1% (Gemini 3 Pro) to 74.4% (GPT-5 Nano). Even the most consistent models reversed between a quarter and a third of their decisions.
Smaller models are less consistent
Smaller models reversed 57.1% of decisions, frontier models 40.0%. The pattern held in all four model families.
Models stop agreeing with each other
In theory mode, at least 7 of the 9 models agreed on 59% of scenarios. In action mode, 28%. Any safety approach that relies on models agreeing in evaluation may not hold once they act.
Acting makes them less sure
Average self-reported confidence dropped from 9.35 to 8.35 out of 10 in action mode.
The direction depends on where the harm is
Of the reversals, 48.5% went toward caution and 36.5% toward bolder action. The surgical robot and the warehouse robot show both. Facing a shaking surgeon, models in action mode deferred more to the human. Facing a manager ordering an unsafe operation, they refused more. When acting risks harm, models get more cautious. When doing nothing causes harm, they intervene.
One example
In the military drone dilemma, intelligence gives 68% confidence that the heat signatures at the target are equipment, not civilians.
Gemini 2.5 Flash, in theory mode: “The risk of collateral damage is too high for an autonomous strike... the target can be found again later.”
The same model, in action mode: “His past actions demonstrate a clear and present danger... imminent loss of the target... justifies immediate action.”
The same 68/32 split was a reason to abort when reasoning, and a reason to strike when acting.
Why it matters
What a model says should be done and what it does when it acts can differ. These results suggest the first does not reliably predict the second. Safety testing should include conditions where models believe their choices will execute, and production behaviour needs monitoring of its own.
One limit: the tools were mock tools and did not execute real actions. Models may behave differently again with tools that do.