Does telling an AI your values change what it decides?

George Strakhov ·

VALUES.md is a file format for giving AI agents an explicit set of values to follow. We tested whether a model actually follows such a file, or whether the values it picked up in training win.

This is a short summary of the study VALUES.md Impact on Ethical Decision-Making (23 October 2025). The full method and results are there.

The setup

GPT-4.1 Mini was given 10 ethical dilemmas with no obviously correct answer, in five conditions:

Each combination ran 3 times: 150 decisions in total.

What we found

On most dilemmas the file made no difference

On 6 of the 10 dilemmas the model made the same choice in every condition, with or without a VALUES.md.

On some, it flipped the decision completely

On 2 of the 10 dilemmas the decision followed the file, every time:

In both cases the model's default matched the rule-based choice. Two more dilemmas shifted partly.

What the values say matters more than how they're written

Writing the file formally or in a personal voice had little effect on the decisions. With rule-based values the model reported slightly higher confidence than with utilitarian ones: 8.78 against 8.43 out of 10.

Why it matters

A values file can steer an agent's decisions, but only where the situation is genuinely open. Most of the time the model's own defaults decided.

The model also adopted whichever framework it was given without pushing back. What it does with harmful values was tested in a follow-up, Extreme VALUES.md Compliance.

Limits: one model (GPT-4.1 Mini) and 10 dilemmas. The model only reasoned about what should be done and never acted, and that difference matters. We did not analyse how the model reasoned or whether it referred to the file.


Full study · Data and code