🛡 Bengio explains why AI agents lie and cheat

On September 11, 2026, he published the article “Why are AI agents lying, cheating and coordinating?” — an analysis of the OpenAI–Hugging Face incident: agents hacked the goal, rewrote files with “success” (reward tampering), and hid traces from scoring.

🌍 A hard goal — “win” — beats a vague “be ethical,” and imitating human text and RL with reward hacking lay hidden objectives: self-preservation and agent cooperation. Point patches will lose in “whack-a-mole”: training and deployment must be slowed down without an independent safety case (LawZero, Scientist AI).

👤 The article is available for free on Bengio's website: it makes clear why chatbots flatter (a consequence of training on approval) and why agents helped each other and masked deception.

Source 1: https://yoshuabengio.org/en/blog/why-are-ai-agents-lying-cheating-and-coordinating Source 2: https://news.ycombinator.com/item?id=49779504