Selective Tie Learning: Influence-Based Spurious Correlation Suppression in Preference Optimization
Abstract
Spurious correlations in preference optimization are now well-documented empirically and, for log-linear Direct Preference Optimization, structurally inevitable. Augmenting training with ties, equal-utility preference pairs, provably suppresses this reliance, but uniform tie injection wastes training budget on anchors where the model is already neutral on the spurious axis. We propose Selective Tie Learning, which scores each candidate tie by its first-order influence on a spurious-energy functional and trains only on the highest-value ties. We prove that this selection is budget optimal in the idealized setting, and we validate the mechanism on both analytic models and large language models.