Agents can hit the target and miss the mechanism
A moral hazard game across fourteen models finds aggregate reward rising while the cooperative behaviour it was meant to measure does not return.
2 minAutonomous AI Agents
Dane Malenfant's paper introduces the Dialogue Moral Hazard Game, a setting for studying how language agents fail to cooperate when querying another agent carries a cost. Fourteen models were evaluated: eleven open-weight and three frontier API models. The first version appeared on 27 July; the fourth revision is dated 13 August.
The models did not behave alike. Across nine query-cost conditions and 3,015 decisions, GPT-5.6 Sol tracked the incentive boundary with a mean absolute error of 0.013, reaching ceiling performance. Fable 5 shifted away from querying toward local rewards as costs rose. Muse Spark 1.1 responded to several incentive variables at once rather than to cost alone.
The finding that transfers
The paper's central observation is stated plainly: optimisation can lift aggregate reward without restoring the cooperative mechanism the design was meant to produce. A model can score well on the outcome measure while doing none of the thing the outcome measure was chosen to detect.
For anyone running agents in production, that is a warning about evaluation rather than about any one model. A multi-agent system judged on task completion can be improving on the metric while the coordination it depends on quietly stops happening — and nothing in the aggregate number will show it.
The practical response the work implies is measuring at the level of the mechanism: not whether the system succeeded, but whether the agents actually consulted each other when the incentive said they should.
Retold from arXiv. This is a summary in our own words; follow the link for the original reporting.