🤖 LLM agents lack a coherent moral foundation
Arno Libert, Derck W.E. Prinzhorn, and Daan R. Henselmans published a paper (arXiv:2609.05036, Paris Journal of AI and Digital Ethics): alignment should start not with a debate about the “right” morality, but with testing four conditions — verdict stability, monotonicity, decisiveness, and Pareto viability. Across nine frontier LLMs in three moral dilemma scenarios, no model produced a coherent policy.
🌍 Rephrasing the prompt shifted the share of verdicts by up to 99 percentage points, and success in one scenario does not predict the result in another.
👤 The article (CC BY 4.0, HTML text available) describes a reproducible evaluation scheme: 5 rephrasings, 5 escalation levels, 3 dominance conditions — it can be applied to your own agents.
Source 1: https://arxiv.org/abs/2609.05036 Source 2: https://arxiv.org/html/2609.05036v1
