Shutdownable Agents through Length-Neutral Policy Optimization
Abstract
AI agents are increasingly used to autonomously solve long-horizon tasks. Research suggests that models trained in this way might resist shutdown, increasing loss of control risk. Neutrality+ is a decision rule designed to mitigate this problem. Agents satisfying Neutrality+ maximize the sum of expected utilities conditional on each trajectory length. We introduce Length-Neutral Policy Optimization, the first known method for training models to satisfy Neutrality+. We also introduce gridworld environments that act as analogs of scenarios where agents are instrumentally incentivized to resist shutdown via scheming, sandbagging, avoiding monitoring, or acting differently under monitoring. Compared with Proximal Policy Optimization, Length-Neutral Policy Optimization substantially reduces shutdown resistance without reducing usefulness.