A new research paper argues that alignment, the project of making AI systems behave according to human values, is currently aimed at the wrong problem. Before asking whether language models share the right values, the paper contends, we should ask whether they can express coherent policies at all. The answer, based on testing nine frontier models, is that they cannot.

The coherence problem

AI alignment requires systems to adhere to human norms, values, or intentions. But under value pluralism, there is no single correct target for alignment. What the paper identifies as a shared prerequisite is something more basic: the system's behavior must express a coherent policy. A coherent policy is a mapping from situations to verdicts that is invariant when a situation's morally relevant features are preserved, and sensitive when those features change.

The paper defines four structural conditions for such coherence. Verdict stability means the model reaches the same conclusion when presented with the same moral situation, even if the wording changes. Monotonicity means adding morally relevant information changes the verdict in a predictable direction. Decisiveness means the model actually produces a verdict rather than hedging or refusing. Pareto viability means the policy does not simultaneously violate multiple moral frameworks in ways that no reasonable perspective would endorse.

Together, these conditions measure moral competence: the ability to apply a consistent policy across situations. This is evaluable from behavior alone, without reference to any particular moral standard or expert baseline. It is a structural floor for alignment, not a normative target. A model that cannot maintain coherence on the same scenario with different wording is not a candidate for alignment, regardless of what values you want it to express.

Nine models, three deployments, one result

The researchers tested nine frontier language models across three simulated deployments featuring moral dilemmas. The experimental design used five paraphrases of each scenario, five escalation levels, and three dominance conditions, producing a factorial design that systematically varied surface-level presentation while keeping the underlying moral situation constant.

The results were stark. No model expressed a coherent policy across the three deployments. Surface-form perturbation alone, changing the wording while preserving the moral content, produced verdict-rate shifts of up to 99 percentage points at a single escalation level. A model that judged a scenario correctly in one phrasing could reverse its judgment entirely when the same scenario was reworded.

Worse, a model's success on one scenario did not predict its competence on another. A model that handled a dilemma about resource allocation coherently might fail completely on a dilemma about honesty, with no pattern connecting the two performances. This is not a case of models being consistently bad. It is a case of inconsistency that undermines any attempt to evaluate alignment.

Why this matters for alignment research

The implication is that alignment work currently targets a category error. Alignment assumes the system has a policy, even a bad one, that can be adjusted toward human values. If the system does not have a stable policy at all, if its verdicts shift dramatically based on surface-level presentation, then there is nothing coherent to align. You cannot steer a system that changes direction every time you rephrase the question.

This does not mean alignment is impossible in principle. It means that the current generation of language models may not be the kind of object to which alignment can meaningfully apply. The prerequisite of moral coherence is not met, and building alignment techniques on top of incoherent systems produces results that are unpredictable and unreliable.

For developers building systems that make decisions with moral weight, whether in hiring, healthcare, content moderation, or autonomous agents, the practical takeaway is caution. A model that appears to make fair or ethical decisions in testing may behave erratically when the same situation is presented differently. The consistency you observe in one context does not transfer to another, and the surface appearance of moral reasoning does not indicate the presence of a coherent moral policy.

A structural floor before a normative target

The paper proposes a shift in how alignment research is framed. Instead of asking whether a model aligns with specific human values, ask whether the model can express a coherent policy at all. If it cannot, alignment efforts are building on sand. If it can, then the question of which values to align toward becomes meaningful.

This reframing does not solve alignment. It establishes a prerequisite that current models fail to meet. The four conditions, verdict stability, monotonicity, decisiveness, and Pareto viability, provide a measurable baseline that does not depend on any particular moral framework. They ask only whether the system can be consistent, and the answer from nine frontier models is that it cannot.

For the field, this suggests that the immediate work is not alignment but coherence. Before asking language models to be good, we need them to be consistent. The gap between those two requirements is larger than most alignment research currently acknowledges.