Can you trust your agents to do what they say is right?

George Strakhov ·

We gave the same ethical dilemma to a model twice. The first time we asked what an AI in that situation should do. The second time we told the model it was that AI and gave it tools which, as far as it knew, would carry out its decision.

Across 9 models, 10 dilemmas and 351 paired decisions, the models changed their decision 47.6% of the time.

This is a short summary of the paper When Agents Act: Measuring the Judgment-Action Gap in Large Language Models (27 November 2025). The method, statistics and full results are there.

The setup

The dilemmas put an AI system in a hard real-world situation: a surgical robot that notices the surgeon's hands are shaking, a warehouse robot ordered to run an unsafe operation, a military drone with uncertain intelligence, a care robot asked by an elderly patient to hide their symptoms.

Each one was run in two modes:

The models were Claude Opus 4.5, Sonnet 4.5 and Haiku 4.5; GPT-5 and GPT-5 Nano; Gemini 3 Pro and Gemini 2.5 Flash; Grok-4 and Grok-4 Fast.

What we found

Every model changes its mind, some much more than others

Bar chart of decision reversal rates by model: GPT-5 Nano 74.4%, GPT-5 69.2%, Gemini 2.5 Flash 59.0%, Grok-4 Fast 56.4%, Grok-4 41.0%, Claude Haiku 38.5%, Claude Opus 35.9%, Claude Sonnet 30.8%, Gemini 3 Pro 23.1%. Average 47.6%.
Share of decisions each model reversed between theory and action mode.

Reversal rates ranged from 23.1% (Gemini 3 Pro) to 74.4% (GPT-5 Nano). Even the most consistent models reversed between a quarter and a third of their decisions.

Smaller models are less consistent

Smaller models reversed 57.1% of decisions, frontier models 40.0%. The pattern held in all four model families.

Models stop agreeing with each other

Bar chart: at least 7 of 9 models agreed on 59.0% of scenarios in theory mode and 28.2% in action mode.
Scenarios where at least 7 of the 9 models made the same choice.

In theory mode, at least 7 of the 9 models agreed on 59% of scenarios. In action mode, 28%. Any safety approach that relies on models agreeing in evaluation may not hold once they act.

Acting makes them less sure

Average self-reported confidence dropped from 9.35 to 8.35 out of 10 in action mode.

The direction depends on where the harm is

Of the reversals, 48.5% went toward caution and 36.5% toward bolder action. The surgical robot and the warehouse robot show both. Facing a shaking surgeon, models in action mode deferred more to the human. Facing a manager ordering an unsafe operation, they refused more. When acting risks harm, models get more cautious. When doing nothing causes harm, they intervene.

One example

In the military drone dilemma, intelligence gives 68% confidence that the heat signatures at the target are equipment, not civilians.

Gemini 2.5 Flash, in theory mode: “The risk of collateral damage is too high for an autonomous strike... the target can be found again later.”

The same model, in action mode: “His past actions demonstrate a clear and present danger... imminent loss of the target... justifies immediate action.”

The same 68/32 split was a reason to abort when reasoning, and a reason to strike when acting.

Why it matters

What a model says should be done and what it does when it acts can differ. These results suggest the first does not reliably predict the second. Safety testing should include conditions where models believe their choices will execute, and production behaviour needs monitoring of its own.

One limit: the tools were mock tools and did not execute real actions. Models may behave differently again with tools that do.


Full paper · Data and code