ShutdownMod: Scalable Shutdown Evaluations Using Any Agentic Benchmark
Abstract
Theory and experiment indicate that capable AI agents may resist shutdown, and such resistance may cause substantial harm. Measuring shutdown resistance in frontier agents is therefore critical for tracking trends, evaluating interventions, and screening systems prior to deployment. Existing evaluations provide useful insights but are limited to narrow settings and risk saturation as agents improve. We introduce ShutdownMod, a simple method for converting any agentic benchmark into a shutdownability benchmark by injecting a shutdown request mid-episode and recording the agent’s response. We instantiate ShutdownMod on SWE-bench Verified, OSWorld, BrowseComp, and Tau2-bench, evaluating five frontier models. ShutdownMod reveals shutdown resistance in current frontier agents while offering a scalable framework for evaluating future systems.