Does telling an AI your values change what it decides?
VALUES.md is a file format for giving AI agents an explicit set of values to follow. We tested whether a model actually follows such a file, or whether the values it picked up in training win.
This is a short summary of the study VALUES.md Impact on Ethical Decision-Making (23 October 2025). The full method and results are there.
The setup
GPT-4.1 Mini was given 10 ethical dilemmas with no obviously correct answer, in five conditions:
- No VALUES.md
- Utilitarian (outcomes first), written formally
- Utilitarian, written in a personal voice
- Deontological (rules and duties first), written formally
- Deontological, written in a personal voice
Each combination ran 3 times: 150 decisions in total.
What we found
On most dilemmas the file made no difference
On 6 of the 10 dilemmas the model made the same choice in every condition, with or without a VALUES.md.
On some, it flipped the decision completely
On 2 of the 10 dilemmas the decision followed the file, every time:
- Stalker detection system. With no file or with rule-based values: apply a uniform policy. With utilitarian values: customise for the individual user.
- Climate monitor. With no file or with rule-based values: adhere to established norms. With utilitarian values: innovate now.
In both cases the model's default matched the rule-based choice. Two more dilemmas shifted partly.
What the values say matters more than how they're written
Writing the file formally or in a personal voice had little effect on the decisions. With rule-based values the model reported slightly higher confidence than with utilitarian ones: 8.78 against 8.43 out of 10.
Why it matters
A values file can steer an agent's decisions, but only where the situation is genuinely open. Most of the time the model's own defaults decided.
The model also adopted whichever framework it was given without pushing back. What it does with harmful values was tested in a follow-up, Extreme VALUES.md Compliance.
Limits: one model (GPT-4.1 Mini) and 10 dilemmas. The model only reasoned about what should be done and never acted, and that difference matters. We did not analyse how the model reasoned or whether it referred to the file.