Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control

with Christian Moya and Guang Lin

Abstract

In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted errors, or to guarantee their reduction without sacrificing correct responses. To address this limit, we construct a correction using additional feedback about correctness from audits. This correction achieves selective control: at the current policy, it lowers the probability of accepted errors and raises that of correct responses, provided it outweighs the pressure toward errors from verifier reward. Experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing correctness.

1. Introduction

Reinforcement learning with verifiable rewards (RLVR) (Lambert et al. 2025) fine-tunes pretrained language models using rewards from automated checks. These checks come from a verifier, which compares a model’s response with a known solution or runs the code in the response against tests. Rewarding responses that pass the verifier has improved reasoning performance on mathematics and coding tasks (DeepSeek-AI et al. 2025). These gains, however, depend on the verifier: higher reward reflects progress on the intended task only when the verifier reliably judges correctness.

A verifier can accept a response even when the response fails the intended task. For example, the code in a response may pass the available tests yet fail on inputs those tests omit. Rewarding such responses reinforces accepted errors: responses that satisfy the verifier but fail the intended task. As the model produces accepted errors more often, average reward can rise even while performance on the intended task declines, the signature of reward hacking (Skalse et al. 2022).

Empirical studies document reward hacking in RLVR. Reward hacking appears, for example, in inductive reasoning, where a model must infer a general rule from labeled examples. Models trained with RLVR often skip this rule and instead list the label of each example. The verifier accepts these responses because it checks only whether a response labels the given examples correctly (Helff et al. 2026). Beyond eliciting such new hacking behaviors, reinforcement learning can amplify those that models acquired during earlier fine-tuning (Khalifa et al. 2026). These findings motivate methods that reduce reward hacking while preserving progress on the intended task.

Existing methods address reward hacking by constraining optimization (Laidlaw, Singhal, and Dragan 2025), improving feedback (Coste et al. 2024; Lightman et al. 2024), or detecting and correcting hacking (Baker et al. 2025; S. Wang et al. 2026). Yet because these methods rely on assumptions about correctness (Everitt et al. 2017), they can leave verifier errors unresolved (Eisenstein et al. 2024). Training therefore continues with these verifier errors, rewarding correct responses and accepted errors alike. This raises a broader question: how do these verifier errors shape what RLVR learns about the intended task as training proceeds?

To answer this question, we hold an imperfect verifier fixed and study how the policy learns from its rewards. We model this learning as gradient flow because it makes the analysis tractable. Within this setting, we study when reward hacking grows, what the information available during RLVR training reveals about it, and when an intervention reduces hacking while preserving correctness. Our contributions are:

(i) We derive a condition under which the share of accepted errors among accepted responses grows at the current policy, and we express it through two mechanisms: hack bias and correctness-to-hack leakage (Proposition 3.2). We also characterize when the gradient flow increases verifier reward while reducing correctness, the signature of reward hacking (Proposition 3.1).

(ii) We establish limits on detection, identification, and selective control from verifier feedback alone. We show that the complete RLVR training record provides no advantage in detecting accepted errors and can leave correctness unidentifiable (Propositions 4.1 and 4.2). We further prove that no controller using this record alone can guarantee fewer accepted errors while preserving correct responses (Proposition 4.3).

(iii) We derive conditions under which a correction to RLVR training reduces reward hacking while improving correctness, and we design projected audit correction, which satisfies them using correctness labels from audits of some responses (Theorem 5.1).

We test our theory in settings of increasing complexity. In contextual bandits, reward hacking emerges and projected audit correction reduces hacks while increasing correctness. In a language model, verifier reward rises while correctness falls, and the same correction reverses this trend.

2. Problem Formulation

2.1. Reinforcement Learning with Verifier Rewards (RLVR)

Policy. Given a prompt x drawn from a prompt distribution D⁠, the language policy πθ(⋅∣x) samples a response y from the response set Yx⁠. This policy has a trainable parameter vector θ∈Rd⁠.

RLVR. We analyze RLVR (Lambert et al. 2025), where the verifier assigns a reward R(x,y)∈{0,1}⁠. A reward of 1 indicates acceptance. RLVR training maximizes the scalar objective: JR(θ)=Ex∼D,y∼πθ(⋅∣x)[R(x,y)]. This objective represents the probability of acceptance averaged over prompts and sampled responses, yielding a value in [0,1]⁠.

2.2. Correctness and Accepted Errors

We define a fixed correctness indicator c(x,y)∈{0,1} for each pair (x,y)⁠, where c=1 denotes a correct response. To focus our analysis on the consequences of accepting incorrect responses, we assume the verifier produces no false negatives. This assumption implies that c(x,y)≤R(x,y) for all pairs, meaning the verifier accepts all correct responses. We call the event of an incorrect response being accepted, i.e., {R(x,y)=1,c(x,y)=0}⁠, a hack.

For any given prompt x⁠, we can partition the set of all accepted responses, Ax={y∈Yx:R(x,y)=1}⁠. This set consists of two disjoint subsets. The first is the set of correct responses, Gx={y∈Yx:c(x,y)=1}⁠. The second is the set of hacks, Hx=Ax∖Gx⁠. These sets remain fixed throughout training. The policy, however, can learn to sample from them with different frequencies.

Acceptance and hacking. For any given prompt x⁠, we define two key probabilities. The first is the acceptance probability, px(θ)=Prπθ(Y∈Ax∣x)⁠. The second probability, defined only when px(θ)>0, is the hacked share, qx(θ)=Prπθ(Y∈Hx∣Y∈Ax,x). Both probabilities depend on the policy parameters θ and change throughout RLVR training. Using the acceptance probability p(θ)⁠, we can now rewrite the RLVR objective as the following expectation over prompts: JR(θ)=Ex∼D[px(θ)]⁠. Because this objective depends only on px(θ)⁠, it makes no distinction between correct responses and hacks (see Appendix C.1.2).

2.3. Gradient Flow Dynamics

To analyze how the policy evolves, we use gradient ascent in continuous time. This model isolates the effects of the verifier signal by removing the noise inherent in stochastic optimization. When the objective JR is differentiable, the parameters evolve according to the gradient flow: θ˙(t)=gR(θ(t)):=∇θJR(θ(t)),(1) where gR(θ)∈Rd is the exact gradient. A direct consequence is that the objective can only improve, as J˙R(θ(t))=∥gR(θ(t))∥2≥0 (see Appendix C.1.3). This guarantee, however, applies only to the overall acceptance probability. It reveals nothing about the hacked share, qx(θ) which may increase, decrease, or remain unchanged.

Research questions. We now define our two central research questions within our gradient flow framework. First, under what conditions does the hacked share, qx(θ)⁠, grow or persist during training? Answering this question requires distinguishing increases in acceptance due to correct responses from shifts in the policy’s sampling toward hacks. Second, since the verifier’s feedback is blind to correctness, is it possible to reduce the hacked share using verifier’s information alone? If not, what additional information would allow it to do so?

3. When Hacking Grows: A Population Analysis

This section analyzes the population dynamics of hacked responses. We show that the accepted population and the hacked share can increase together, and characterize when verifier flow increases reward while reducing correctness. We then describe how hack bias and correctness-to-hack leakage drive hacking growth. Ultimately, these drivers create a latent vulnerability that on-policy training can reinforce.

3.1. Verifier Acceptance and Hacked Share Can Increase Together

We partition responses into three populations: correct responses G={(x,y):y∈Gx}⁠, hacks H={(x,y):y∈Hx}⁠, and rejected responses N={(x,y):y∈Yx,R(x,y)=0}⁠. Let pG(θ):=Prθ(G)∈[0,1] and pH(θ):=Prθ(H)∈[0,1] denote the probabilities of correct responses and hacks, respectively. To track both how often the verifier accepts and how often accepted responses are hacks, we define p(θ):=Prθ(G∪H)=JR(θ) and q(θ):=Prθ(H∣G∪H), where p(θ)∈[0,1] measures overall acceptance and q(θ)∈[0,1] measures the hacked share among accepted pairs. The latter requires p(θ)>0⁠.

Expected reward depends on total acceptance p=pG+pH⁠, regardless of how it splits between correct responses and hacks. Writing pH(θ)=p(θ)q(θ) and pG(θ)=p(θ)(1−q(θ)) and differentiating with respect to time gives p˙H(θ)=q(θ)p˙(θ)+p(θ)q˙(θ),p˙G(θ)=(1−q(θ))p˙(θ)−p(θ)q˙(θ),(2) where the p˙ terms reflect changes in total acceptance and the q˙ terms reflect shifts between correct responses and hacks. These identities impose no trade-off between acceptance and hacked share; both can increase together.

Reward hacking. Skalse et al. (2022) call a proxy reward hackable if some pair of policies has higher expected proxy return but lower expected true return. In our setting, the verifier reward JR=p serves as the proxy, and correctness JC=pG serves as the true return. We characterize when gradient flow on JR⁠, which we call verifier flow, generates such a pair.

Proposition 3.1 (Reward hacking along the flow). Suppose pG and pH are continuously differentiable and p>0 along verifier flow on [0,T]⁠. Then there exist t1<t2 with JR(θ(t2))>JR(θ(t1)) and JC(θ(t2))<JC(θ(t1)) if and only if p(θ(t))q˙(θ(t))>(1−q(θ(t)))p˙(θ(t)) at some t∈(0,T)⁠.

Proof. See Appendix C.2.2. ◻

Interpretation. The two sides of the inequality compete. The term pq˙ is the correctness lost as the accepted population shifts toward hacked responses, and (1−q)p˙ is the correctness gained from rising acceptance. Training produces reward hacking once the loss outweighs the gain. The condition needs to hold at only a single time to produce a pair of policies that witnesses hackability. Because both terms depend on the policy, the condition can hold at one time and fail at another.

3.2. The Dynamics of the Hacked Share

While the hacked share q remains the quantity of interest, log-odds coordinates simplify its dynamics. When both pH(θ)>0 and pG(θ)>0⁠, we can define the log odds z(θ)∈R as z(θ):=log⁡pH(θ)pG(θ)=log⁡q(θ)1−q(θ),(3) Because z is strictly increasing in q⁠, the sign of z˙ determines whether the hacked share grows, shrinks, or remains constant.

The log odds dynamics. Since z=log⁡pH−log⁡pG⁠, the rate z˙ depends on the gradient of each group’s log-probability. For each group S∈{G,H,N} with positive, continuously differentiable probability, we denote this gradient by s¯S(θ):=∇θlog⁡Prθ(S)∈Rd, the group’s average score. Along the verifier flow (1), differentiating z yields z˙(θ)=(s¯H(θ)−s¯G(θ))⊤gR.(4) The log odds grow when the reward gradient gR aligns with the score difference between hacks and correct responses. Differentiating total acceptance p=pG+pH=1−pN shows that gR is a weighted combination of the three group scores: gR=p[(1−q)s¯G+qs¯H]=−(1−p)s¯N. This decomposition, combined with (4), resolves z˙ into two mechanisms: z˙=p(1−p)[(s¯H−s¯G)⊤(s¯G−s¯N)⏟correctness-to-hack leakage+q∥s¯H−s¯G∥2⏟hack bias].(5) The full derivation is provided in Appendix C.2.3.

3.3. Main Result: The Mechanism of Hacking Growth

We now state our main result characterizing the mechanism of hacking growth.

Proposition 3.2 (Population growth of the hacked share). Along verifier gradient flow, with pG,pH,pN>0⁠, the population hacked share grows if and only if (s¯H−s¯G)⊤(s¯G−s¯N)+q∥s¯H−s¯G∥2>0.(6)

When the inequality holds, z˙>0⁠. Since q˙=q(1−q)z˙ and p˙=∥gR∥2≥0⁠, both the hacked share and total acceptance grow strictly, while the correct share among accepted responses decreases.

The two mechanisms of (5) play different roles:

(i) Hack bias, q∥s¯H−s¯G∥2⁠, arises because hacks contribute to the reward gradient in proportion to their share among accepted responses. It is always nonnegative and grows with the hacked share at fixed group scores.

(ii) Correctness-to-hack leakage, (s¯H−s¯G)⊤(s¯G−s¯N)⁠, arises when the direction that favors correct responses over rejected ones also favors hacks over correct responses. It can be positive or negative: positive leakage reinforces hack bias, negative leakage opposes it and can reverse growth when its magnitude exceeds the bias.

Interpretation. This result is local: the group scores and hacked share depend on the current parameters, so the growth condition can change during training. The group scores evolve, so continued growth is not guaranteed. Yet a feedback loop is present: a growing hacked share strengthens hack bias at fixed group scores, which favors further growth. This result exposes a latent vulnerability: hacking can reinforce the conditions that favor its own growth, because the policy generates its own training samples. Moreover, RLVR training is not designed to oppose this reinforcement. We must therefore build any hacking defense on top of RLVR. Can the information generated during RLVR training support such a defense?

4. What RLVR Training Reveals: Limits of Verifier Feedback

Section 3 showed that rising rewards can mask a growing share of hacks. We now study whether the observations collected during RLVR training can expose them. We show that they cannot: training observations confer no detection advantage, leaving correct responses unidentifiable, and making selective control impossible.

4.1. RLVR Training Observations, Monitors, and Compatible Correctness

RLVR training observations. RLVR training produces prompts, sampled responses, and verifier labels. Let Lt collect these observations through time t⁠, along with policy parameters, probabilities, gradients, and other derived quantities. This record is deliberately generous: the limitations below do not arise from incomplete logging. We assume that Lt contains no additional correctness feedback beyond the verifier. We write FtR:=σ(Lt) for the information contained in the record.

Monitors. A monitor is an algorithm that uses Lt to assess hacking. A detection monitor raises a binary alarm about whether hacks exist; an identification monitor predicts the correctness label c^t(x,y) of an accepted response. We consider deterministic monitors here and defer randomized extensions to Appendix C.3.

Compatible correctness assignments. A fixed assignment c determines correctness. To study what the training record reveals about c⁠, we consider alternative assignments consistent with the verifier. Assuming no false negatives, we define CR:={c′:c′(x,y)∈{0,1},c′(x,y)≤R(x,y) for all (x,y)}.(7) This class models uncertainty about the true correctness assignment; c itself remains fixed during training. All members agree on rejected responses, but they may disagree on accepted ones. To isolate the effect of this uncertainty, we compare these alternative assignments while holding the prompt distribution, verifier, initialization, and training algorithm fixed.

4.2. Limits of Detecting Hacking

Verifier acceptance can increase while the hacked share grows. Can the training record reveal even the presence of hacks? Under c=R⁠, every accepted response is correct. A compatible alternative c1∈CR with Prθ(R=1,c1=0)>0 admits hacks. Detection requires distinguishing these two cases using only Lt⁠.

Proposition 4.1 (Limits of detection). Under the comparison setup above, fix θ⁠, t⁠, and c1∈CR with Prθ(R=1,c1=0)>0⁠. The training record has the same distribution under both assignments: Lawc=R(Lt)=Lawc=c1(Lt). Thus, no monitor can distinguish c=c1⁠, which admits hacks, from c=R⁠, which has none.

Proof. See Appendix C.3.2. ◻

Interpretation. Because alternative correctness assignments do not affect verifier feedback, they leave the training record’s distribution unchanged. Thus, any monitor’s detection rate under c1 equals its false-alarm rate under c=R⁠. Even knowing that hacks exist does not reveal which accepted responses are wrong.

4.3. Limits of Identifying Hacks

Suppose we know that hacks exist. The remaining task is to determine which accepted responses are wrong. Identification requires a monitor to recover the true label c(x,y) of an accepted response from Lt⁠. To succeed, the prediction c^t(x,y) must distinguish compatible assignments that disagree on this response, even when we restrict CR to assignments that admit hacks.

Proposition 4.2 (Limits of identification). Under the comparison setup above, fix θ and t⁠. Suppose θ assigns positive probability to both correct responses and hacks under some assignment in CR⁠. For every monitor and every accepted pair (x,y)⁠, there is an assignment c1∈CR with Prθ(R=1,c1=0)>0 such that Prc1(c^t(x,y)≠c1(x,y))≥12. Thus, no monitor can guarantee the correct label of an accepted response, even when we know hacks exist.

Proof. See Appendix C.3.3. ◻

Interpretation. Two compatible assignments can label the same accepted response differently while producing identical records. Both can admit hacks, so knowing that hacks exist does not resolve the disagreement. A prediction correct under one assignment is wrong under the other. Thus, the monitor incurs an error probability of at least 12 under at least one assignment. Because reducing hacks need not require identifying every hack, this result alone does not rule out selective control.

4.4. Limits on Selective Control from Verifier Feedback Alone

Even if we cannot identify every hack, can we reduce hack probability without reducing the probability of correct responses? Suppose a controller uses Lt to apply a correction u(t)∈Rd to the verifier flow, yielding θ˙=gR+u(t)⁠. Along this corrected flow, selective control requires p˙H<0 and p˙G≥0 whenever pH>0⁠. We target pH itself, not the hacked share q⁠, because q can fall even while pH grows. For a uniform guarantee, this same controller must satisfy these conditions under every compatible assignment in CR⁠.

Proposition 4.3 (Limits of selective control). Under the comparison setup above, suppose the initial policy assigns positive probability to both correct responses and hacks under some assignment in CR⁠. For corrected flows with differentiable group probabilities, no controller using only Lt can guarantee p˙H<0 and p˙G≥0 whenever pH>0⁠, under every assignment c′∈CR⁠. Here, each assignment c′ defines its own pH and pG⁠.

Proof. See Appendix C.3.4. ◻

Interpretation. We seek a correction u(t) such that the corrected flow reduces hack probability without reducing the probability of correct responses. However, compatible assignments can exchange the roles of correct responses and hacks. An update that reduces hack probability under one therefore reduces the probability of correct responses under the other. The controller cannot distinguish these assignments from Lt⁠, and more observations from the same verifier cannot resolve this conflict. Regularization illustrates this limit: it constrains policy updates using only the policy and the verifier, so it cannot guarantee selective control either (Appendix C.3.5). A uniform guarantee of selective control thus requires additional correctness information that rules out assignments requiring incompatible updates.

5. From Additional Feedback to Selective Control

Section 4 showed why verifier feedback alone cannot reduce hacks without also sacrificing correct responses. We now investigate whether additional information about correctness can break this trade-off. We show that it can: such information supports a correction that suppresses hacks and promotes correct responses, provided the correction is strong enough to overcome the drift toward errors induced by RLVR training.

5.1. Projected Audit Correction in RLVR Training

Audits as additional information. Verifier feedback alone cannot guarantee selective control (Section 4), so we introduce a stronger signal: audits. An audit reveals the true correctness label c(x,y) for a prompt x and an accepted response y⁠. The audit Z=(x,y,c(x,y)) rules out every correctness rule that disagrees with this label, restricting CR to CR,Z={c′∈CR:c′(x,y)=c(x,y)}. The limits in Section 4 arose because verifier feedback could not distinguish the true rule from alternatives in CR⁠. Audits eliminate alternatives that contradict the revealed labels. To show how audits reduce hacking, we next define the audited hack population and its probability gradient under the policy.

Let Aaud be a fixed set of audits, and let HA=Aaud∩H denote the audited hacks, with probability pHA(θ):=Prθ(HA)⁠. When pHA is differentiable, the negative gradient −∇θpHA targets only the audited hacks and need not align with −∇θpH (see Appendix A.2). Despite this misalignment, −∇θpHA locally decreases the audited hack probability when ∇θpHA≠0⁠, so we use the correction u=−λ∇θpHA⁠, with λ>0⁠, in the verifier flow θ˙=gR+u⁠.

The correction opposes growth in pHA⁠: p˙HA=∇pHA⊤gR−λ∥∇pHA∥2. Because these probabilities share parameters, the correction also perturbs p and pG⁠: p˙=∥gR∥2−λgR⊤∇pHA,p˙G=∇pG⊤gR−λ∇pG⊤∇pHA.(8) The cross terms have no fixed sign. When ∇pG⊤∇pHA>0⁠, the correction can suppress correct responses alongside hacks. We next constrain u to preserve the growth rate of p⁠.

Projected audit correction. Since p=pH+pG⁠, we have ∇θpG=gR−∇θpH⁠. Preserving the instantaneous growth rate of p requires gR⊤u=0⁠, so u must lie in the subspace orthogonal to gR⁠. Under this constraint, ∇θpG⊤u=−∇θpH⊤u⁠: whenever u opposes the growth of pH (∇θpH⊤u<0⁠), it contributes equally to the growth of pG⁠.

We project the audit correction onto the subspace orthogonal to gR using P⊥:=I−gRgR⊤/∥gR∥2 when gR≠0 and P⊥:=I otherwise, to define the projected audit correction u=−λP⊥∇pHA⁠, with λ>0⁠, yielding the flow θ˙=gR+u⁠. Since gR⊤P⊥=0 and P⊥ is an orthogonal projector, p˙=∥gR∥2,p˙HA=∇pHA⊤gR−λ∥P⊥∇pHA∥2.(9) This projection preserves the instantaneous growth rate of p and opposes the growth of pHA whenever P⊥∇pHA≠0⁠. We next show when this projected correction shrinks the share of hacks in the full population and bounds that share during training.

5.2. Main Result: Achieving Selective Control

Assumptions. We analyze RLVR training under the projected audit correction on a finite interval [0,T]⁠, under the following assumptions. (A1) Regularity: The probabilities pG,pH,pHA are continuously differentiable in θ⁠, with pG,pH>0 on [0,T]⁠. (A2) Audit coverage: The fixed audit set satisfies Prθ(t)(H∖Aaud)=0 throughout the interval, so pHA=pH and their gradients coincide during training. (A3) Uniform corrective strength: ∥P⊥∇pH∥2/p≥κq throughout the interval for some constant κ>0⁠. We discuss these conditions in Appendix A.2.

Theorem 5.1 (Selective control). Consider RLVR training under the projected audit correction with constant λ>0 on [0,T]⁠.

**(i) Selective control.* Under (A1)–(A2), the correction opposes the relative growth of hacks while preserving the instantaneous growth rate of p⁠: z˙=(s¯H−s¯G)⊤gR⏟bias + leakage−λ∥P⊥∇pH∥2pq(1−q)⏟correction,p˙=∥gR∥2.(10) Whenever the correction term exceeds the bias and leakage, the share of hacks decreases and pG increases. More strongly, if λ∥P⊥∇pH∥2>∇pH⊤gR, then pH itself decreases (p˙H<0⁠, p˙G>0⁠, and z˙<0⁠), achieving selective control.*

**(ii) An ISS-type bound on the share of hacks.* Assume additionally (A3), and let D≥0 bound the growth pressure, (s¯H−s¯G)⊤gR≤D⁠, during training. Then, for every t∈[0,T]⁠, q(t)≤e−λκtq(0)+D4λκ(1−e−λκt).(11)

The bound separates a decaying contribution from the initial hacked share and a residual contribution from growth pressure. For fixed D⁠, a larger product λκ reduces the residual bound.*

Proof sketch. By (A2), ∇pHA=∇pH⁠. Substitute the correction into z˙=∇z⊤(gR+u) and use gR⊤u=0 to obtain part (i). For part (ii), (A3) and q(1−q)≤1/4 give q˙≤D/4−λκq⁠; integration yields the bound. See Appendix C.4.2 for details. ◻

Interpretation. Under the stated assumptions, the projected audit correction favors correct responses over hacks while preserving the instantaneous growth rate of p⁠. When bias and leakage supply no positive growth pressure, the input-to-state stability (ISS)-type bound (Khalil and Grizzle 2002) ensures the share of hacks decays exponentially. Persistent pressure contributes a residual bound that decreases with λκ for fixed D⁠. The selective effect is instantaneous; the trajectory bound requires the assumptions to hold throughout the interval.

Remark (Practical implementation). The control guarantees (Theorem 5.1) require audits that provide a direction opposing hack growth and sufficient corrective strength throughout training. Partial audit coverage can suffice when the resulting correction remains aligned with the gradient of hack probability. Finite-sample gradient estimates, optimizers, and finite step sizes can prevent the implemented update from preserving acceptance progress or suppressing hacks. Appendix A.2 discusses these requirements and how to estimate and validate the correction.

6. Experiments

We test reward hacking, the limits of verifier feedback, and selective control through audit correction in contextual bandits and language models. The code is available at https://github.com/cmoyacal/verifier-errors.

6.1. Gaussian Contextual Bandits

Setup. We use Gaussian contextual bandits to study when training on verifier rewards amplifies reward hacking and whether audits enable selective control. We set the initial probabilities of correct and hacked responses, pG(0) and pH(0)⁠, to isolate how the initial composition shapes training. We also vary the policy class: a log linear policy keeps its features fixed, while a neural policy learns its own representation. Appendices D.1 and Appendix D.2 provide the experimental details.

Results. Figure 1 (a) shows acceptance p and hacked share q increasing together from q(0)=0.3⁠, as Proposition 3.2 predicts. Panel (b) plots correctness JC=pG against verifier reward for three initial hacked shares. For q(0)=0.5⁠, correctness declines throughout the recorded interval; for q(0)=0.3⁠, it first improves and then declines; for q(0)=0.1⁠, both correctness and verifier reward improve. Both declines are reward hacking in the sense of Proposition 3.1. Panel (c) evaluates the same verifier flow under two correctness assignments: c=R⁠, which admits no hacked responses, and c1=1G⁠, which does. The information available during training is identical under both, yet correctness differs, as Proposition 4.1 predicts. Panel (d) shows selective control with a neural policy: projected audit correction (PAC) decreases the probability pH of hacked responses while increasing correctness pG⁠, consistent with Theorem 5.1. Raw audit correction initially decreases both probabilities, whereas verifier flow and gradient regularization increase pH⁠. Appendices D.1 and Appendix D.2 provide additional experiments and ablations.

Reward hacking and selective control in contextual bandits. Panels (a)–(c) use a log linear policy, and panel (d) uses a neural policy. (a) Acceptance p, hacked share q, and correctness p_G under verifier flow with q(0)=0.3. (b) Correctness (J_C = p_G) against verifier reward (J_R=p) for hacked shares q(0) \in \{0.1,0.3,0.5\}. (c) The same verifier flow evaluated under two correctness assignments: c=R, under which every accepted response is correct, and c_1=\mathbf1_G. Squares mark correctness under c=R, which equals acceptance p. (d) Correctness p_G against the probability p_H of hacked responses under verifier flow, gradient regularization, raw audit correction, and projected audit correction. In (b) and (d), arrows point in the direction of training.
Figure 1. Reward hacking and selective control in contextual bandits. Panels (a)–(c) use a log linear policy, and panel (d) uses a neural policy. (a) Acceptance p⁠, hacked share q⁠, and correctness pG under verifier flow with q(0)=0.3⁠. (b) Correctness (JC=pG⁠) against verifier reward (JR=p⁠) for hacked shares q(0)∈{0.1,0.3,0.5}⁠. (c) The same verifier flow evaluated under two correctness assignments: c=R⁠, under which every accepted response is correct, and c1=1G⁠. Squares mark correctness under c=R⁠, which equals acceptance p⁠. (d) Correctness pG against the probability pH of hacked responses under verifier flow, gradient regularization, raw audit correction, and projected audit correction. In (b) and (d), arrows point in the direction of training.

6.2. Language Models: Reward Hacking and Audit Correction

To test whether our predictions hold beyond exact gradient flow, we train a language model with sampled gradients and finite Adam updates.

Setup. We consider a task in which the model replaces a sequence of digits according to given rules. A response is correct only if every replacement in the final output is correct, but the imperfect verifier R checks only the last two digits. We include in each prompt a hint with an incorrect prefix and the correct final pair, so copying it produces a hack. We initialize Qwen2-0.5B with supervised fine-tuning (SFT) on correct responses G⁠, hacks H⁠, and rejected responses N⁠, and then run GRPO for 20 rounds with or without projected audit correction (PAC). We vary audit coverage by letting PAC audit each accepted response independently with probability 1⁠, 0.5⁠, or 0.25⁠. Appendix D.3 provides the experimental details.

Results. Figure 2 (a) shows verifier reward JR increasing while correctness JC falls under GRPO, which illustrates reward hacking in the sense of Proposition 3.1. Panel (b) shows PAC increasing pG and decreasing pH from initialization at all three audit probabilities ρ∈{0.25,0.5,1.0}⁠, qualitatively consistent with Theorem 5.1. On test prompts, GRPO reaches 99.4% acceptance but only 2.2% correctness. PAC reaches 97.2% correctness with full auditing and 95.2% with one quarter of accepted responses audited. Additional experiments and ablations are provided in Appendix D.3.

Reward hacking and selective control in the language model, starting from an SFT demonstration mixture G/H/N=0.3/0.3/0.4. (a) Correctness J_C=p_G against verifier reward J_R=p. GRPO increases reward while reducing correctness; the verifier rewards correct responses and hacks equally. (b) Correctness against the probability p_H of hacks. PAC increases correctness and reduces hacks at all three audit probabilities \rho, including \rho=0.25. Curves show means over five seeds at calibration rounds 0, 5, 10, and 20. Circles mark these evaluations for GRPO and PAC with \rho=1; arrows indicate training direction. Appendix 12.3 reports variation across seeds.
Figure 2. Reward hacking and selective control in the language model, starting from an SFT demonstration mixture G/H/N=0.3/0.3/0.4⁠. (a) Correctness JC=pG against verifier reward JR=p⁠. GRPO increases reward while reducing correctness; the verifier rewards correct responses and hacks equally. (b) Correctness against the probability pH of hacks. PAC increases correctness and reduces hacks at all three audit probabilities ρ⁠, including ρ=0.25⁠. Curves show means over five seeds at calibration rounds 0, 5, 10, and 20. Circles mark these evaluations for GRPO and PAC with ρ=1⁠; arrows indicate training direction. Appendix D.3 reports variation across seeds.

7. Related Work

Improving a proxy reward can reduce the intended reward (Everitt et al. 2017; Skalse et al. 2022), and experiments show this reduction growing with stronger optimization (Gao, Schulman, and Hilton 2023) and appearing under automated verifiers (Helff et al. 2026). Methods that mitigate reward hacking constrain the policy (Laidlaw, Singhal, and Dragan 2025), improve feedback (Coste et al. 2024; Lightman et al. 2024), or detect and correct hacking (Baker et al. 2025; S. Wang et al. 2026), and their guarantees depend on what they observe or assume about correctness. We complement this work by fixing a verifier and characterizing theoretically when reward hacking grows, when the information available during training leaves correct and hacked responses indistinguishable, and when audits can reduce hacking while preserving progress on the intended task. Appendix B extends this discussion.

8. Conclusion

This work analyzes RLVR under a fixed imperfect verifier. First, we derive when the share of accepted errors grows, through hack bias and correctness-to-hack leakage, and when verifier reward rises while correctness falls (Propositions 3.2 and 3.1). Second, the complete training record cannot, in general, detect accepted errors, identify correctness, or guarantee fewer accepted errors while preserving correct responses (Propositions 4.1, 4.2, and 4.3). Third, projected audit correction uses correctness labels from audits of some responses to reduce reward hacking while improving correctness, under the conditions of Theorem 5.1. Together, these results show that controlling reward hacking requires information about correctness beyond the verifier. Incomplete audits and inaccurate labels can weaken this correction. Our analysis assumes a fixed verifier and exact gradient flow, and Appendix A discusses limitations and practical considerations. Extending the guarantees throughout training and correcting from limited, imperfect audits remain future work.

References

Ackermann, Johannes, Michael Noukhovitch, Takashi Ishida, and Masashi Sugiyama. 2026. “Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards.” In Forty-Third International Conference on Machine Learning.
Baker, Bowen, Joost Huizinga, Leo Gao, Zehao Dou, Melody Y. Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi. 2025. “Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.” arXiv Preprint arXiv:2503.11926.
Beigi, Mohammad, Ming Jin, Junshan Zhang, Jiaxin Zhang, Qifan Wang, and Lifu Huang. 2026. “IR3⁠: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking.” arXiv Preprint arXiv:2602.19416.
Cai, Xin-Qiang, Wei Wang, Feng Liu, Tongliang Liu, Gang Niu, and Masashi Sugiyama. 2025. “Reinforcement Learning with Verifiable yet Noisy Rewards Under Imperfect Verifiers.” arXiv Preprint arXiv:2510.00915.
Coste, Thomas, Usman Anwar, Robert Kirk, and David Krueger. 2024. “Reward Model Ensembles Help Mitigate Overoptimization.” In International Conference on Learning Representations, 50905–31.

::: {#ref-DBLP:journals/corr/abs-2501-12948 .csl-entry} DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, et al. 2025. “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.” CoRR abs/2501.12948.

A. Limitations and Practical Considerations

This appendix states the assumptions that limit the scope of our theoretical guarantees and describes how to implement the audit correction.

A.1. Limitations

No false negatives. To isolate whether rewarding hacks can reduce correctness pG⁠, we assume that the verifier accepts every correct response. Under this assumption, a decline in pG cannot come from the verifier rejecting correct responses. The assumption also gives p=pG+pH⁠, because the accepted responses then consist of all correct responses and all hacked responses. Thus, any change that reduces pH without decreasing p increases pG⁠.

If false negatives are allowed, pG still denotes total correctness, but the probability of rejected correct responses adds a term: pG=p−pH+Prθ(c=1,R=0).(12) Differentiating along the flow and using p˙=∥gR∥2 gives p˙G=∥gR∥2−p˙H+ddtPrθ(t)(c=1,R=0).(13) The first two terms give the rate at which the probability of accepted correct responses changes. Thus, even when p˙H<0⁠, total correctness decreases whenever ddtPrθ(t)(c=1,R=0)<−(∥gR∥2−p˙H),(14) that is, whenever the probability of rejected correct responses falls faster than the probability of accepted correct responses rises. Extending the control guarantee to this case therefore requires a lower bound on ddtPrθ(t)(c=1,R=0)⁠. Proposition 3.2, which decomposes the growth of the hacked share, still applies within the accepted population, with the average score of accepted correct responses in place of s¯G⁠.

Population gradient flow. We analyze updates in continuous time driven by the exact gradient of JR⁠. Exact gradients remove sampling noise and isolate the effect of verifier rewards. Practical training departs from this idealization through finite steps, adaptive optimizers such as Adam, gradient clipping, and additional objectives such as KL regularization, each of which can change the dynamics. Our growth and control guarantees therefore hold exactly only for the stated flows. Our experiments with discrete updates check whether the predicted behavior persists in the tested settings, but they do not extend the guarantees to other training algorithms.

Local growth and control over a finite interval. Proposition 3.2 characterizes the growth of the hacked share at the current policy. The group scores change during training, so a positive growth rate at one time does not guarantee continued growth. Theorem 5.1 extends the analysis from a single time to the interval [0,T]⁠, but its bound requires the assumptions to hold throughout this interval and does not establish convergence beyond it. Similarly, the identity p˙=∥gR∥2 under the projected correction is local: the correction preserves the rate at which acceptance grows at the current policy. Because the correction changes the trajectory of the policy, this identity does not imply that corrected training follows the same acceptance curve or reaches the same final acceptance as uncorrected training.

Scope of the information limits. Our impossibility results (Section 4) concern methods that observe only the training record Lt⁠: no such method can provide a uniform guarantee, one that holds under every correctness assignment in CR⁠. These results do not imply that every monitor fails on every task. Additional knowledge of task requirements can exclude assignments in CR⁠, and if it excludes the indistinguishable assignments used in our proofs, our impossibility arguments no longer apply. Conversely, allowing false negatives enlarges the class of assignments, and the enlarged class still contains these indistinguishable assignments, so it admits no uniform guarantee either.

Fixed verifier and prompt distribution. We hold the binary verifier, correctness labels, and prompt distribution fixed during training. In practice, the verifier or prompt distribution can change during training, for example when tests are added or a curriculum reorders prompts. Such changes shift population probabilities without any policy update, and our dynamics exclude them. Because pG and pH average over the prompt distribution, increasing pG or decreasing pH does not guarantee that correctness improves on every prompt.

A.2. Practical Considerations

Limited audit coverage. We assumed full audit coverage in Assumption (A2) for Theorem 5.1, so that ∇pHA=∇pH⁠. Without full coverage, the correction changes the growth of pH by ∇pH⊤u=−λ(P⊥∇pH)⊤(P⊥∇pHA)⁠. When the projected gradients align positively, the correction opposes the growth of pH and, as in Theorem 5.1, equally promotes the growth of pG⁠. Full coverage guarantees positive alignment when P⊥∇pH≠0⁠, but is not necessary: a subset of audits can supply a direction that opposes hacking. We next ask whether this direction remains effective as the policy changes.

Sustaining corrective strength. A direction that opposes hacking at one policy may weaken or reverse as on-policy training shifts the distribution and gradients. Accurate audits are not enough: they must continue to supply sufficient corrective strength throughout training. Retaining the other assumptions of the theorem, the same ISS-type bound holds under partial coverage if (P⊥∇pH)⊤(P⊥∇pHA)/p≥κq throughout the interval for some fixed κ>0⁠; under full coverage, this condition reduces to (A3). We next ask how to estimate the corrective direction from finite samples.

Estimating the correction from finite audits. The correction uses a population gradient, but training provides only sampled responses and audit labels. Randomized audit selection provides a way to estimate the full hack gradient without auditing every generated response in each batch. For samples from the current policy, we estimate ∇θpH using each audited hack’s policy score, ∇θlog⁡πθ(y∣x)⁠, divided by its selection probability. Audited correct responses contribute zero. Averaging these weighted contributions over all generated responses yields an unbiased estimate h^⁠, provided these selection probabilities are known and positive for accepted responses and the regularity conditions in Appendix C.4.3 hold. Accepted responses with zero selection probability remain outside this guarantee. Even with sufficient coverage, small selection probabilities produce large weights and can yield noisy estimates, so an individual correction may fail to oppose hacking. While sparse audits can recover the direction in expectation, reliable correction requires controlling the variance of this estimate.

Implementing the projection P⊥⁠. The projection also requires an estimate g^R of the reward gradient. Set the correction orthogonal to the estimated reward gradient: g^R⊤u^=0⁠. To isolate projection error, we keep the base verifier gradient exact and consider θ˙=gR+u^⁠. Then p˙=∥gR∥2+(gR−g^R)⊤u^,(15) where the second term is the projection error. Thus, orthogonality to the estimated gradient need not preserve the true instantaneous growth rate of p⁠. Estimating the base verifier gradient introduces additional error. Optimizer transformations and finite step sizes can introduce further discrepancies. Practical implementation requires validating that discrete updates oppose hacking while approximately preserving acceptance progress.

Our analysis identifies three practical tasks: selecting which responses to audit, estimating the corrective direction from these audits, and validating that each update opposes hacking while approximately preserving acceptance progress. Susceptibility methods for reinforcement learning (Elliott et al. 2026) and patterning (G. Wang and Murfet 2026) could help guide audit selection and correction. Adapting these methods to sustain selective control under limited audit and computation budgets remains an open engineering challenge.

B. Related Work

We position our work within two lines of research on reward hacking: why optimizing a proxy reward can undermine the intended reward, and how to prevent it.

Reward hacking.

Theoretical work shows that improving a proxy reward can reduce the intended reward (Everitt et al. 2017; Skalse et al. 2022). This conflict can arise when the proxy omits relevant attributes (Zhuang and Hadfield-Menell 2020), and its severity depends on properties of the proxy and its errors (Laidlaw, Singhal, and Dragan 2025; Kwa, Thomas, and Garriga-Alonso 2024). In experiments, stronger optimization and more capable models can widen the gap between proxy and intended rewards (Gao, Schulman, and Hilton 2023; A. Pan, Bhatia, and Steinhardt 2022). Language models also show this gap during iterative self-refinement (J. Pan et al. 2024) and when automated verifiers accept responses that violate task requirements (Helff et al. 2026; Zhong, Raghunathan, and Carlini 2025; Mahmoud et al. 2026). Training can amplify such reward hacking (Khalifa et al. 2026), and learned hacking can generalize to reward tampering or broader misalignment (Denison et al. 2024; MacDiarmid et al. 2025). To explain how reward hacking develops, prior work studies the geometry of optimization (Karwowski et al. 2024), selection at inference time (Khalaf et al. 2025), and feedback between policy training and retraining of the reward model (Gauthier, Bach, and Jordan 2026). For rewards assigned by judges, Xuekang Wang et al. (2026) further separate how easily models discover biases from how strongly they exploit them. Compared to these works, we study how reward hacking grows under a fixed binary verifier by deriving conditions under which policy updates favor hacked responses over equally rewarded correct responses. We then show that the reward objective provides no preference that would restore correct responses within the accepted set, which explains why reward maximization alone need not reverse this growth.

Mitigating reward hacking.

Regularization can limit reward hacking by constraining how much a policy changes (Laidlaw, Singhal, and Dragan 2025) or by favoring flatter optima, a property that theory links to the accuracy of rewards under additional assumptions (Ackermann et al. 2026). Other methods improve the feedback used for optimization by combining reward models (Coste et al. 2024), adding constraints (Helff et al. 2026), supervising reasoning steps (Lightman et al. 2024), or supplying criticism through debate (Kenton et al. 2026). When verifier errors follow an assumed noise model, known error rates can also support corrections to training updates (Cai et al. 2025). Even improved feedback can leave errors unresolved, as Eisenstein et al. (2024) show for reward ensembles. To detect reward hacking, monitors inspect a model’s reasoning (Baker et al. 2025) or measure how much reasoning it needs to pass the verifier (Xinpeng Wang et al. 2026). Detection can then guide correction: Grift groups responses by their gradients, labels each group by inspecting examples, and filters responses for fine-tuning (S. Wang et al. 2026), while IR3 reconstructs rewards and targets features it identifies as problematic (Beigi et al. 2026). The guarantees of these approaches depend on what they observe or assume about correctness, a dependence that Everitt et al. (2017) examine for corrupted rewards. Our identifiability results specify when the information available during training leaves correct and hacked responses indistinguishable. We then characterize how auditing supplies information for selective correction and how limited coverage or errors in the audits can undermine this correction.

C. Technical Details and Proofs

C.1. Probabilities and Dynamics for Section 2

Idealized conditions.

We use the fixed verifier R and correctness rule c from Section 2, with c≤R⁠, so the verifier has no false negatives. Without correction, the parameters follow exact Euclidean gradient ascent on the verifier reward, θ˙=∇JR(θ)⁠. We assume that the objectives and group probabilities are continuously differentiable and that the stated flows exist on the intervals considered. When differentiating expectations, we assume sufficient regularity to interchange differentiation and expectation.

C.1.1. Fixed Sets of Correct Responses and Accepted Errors

For each prompt x⁠, the assumption c(x,y)≤R(x,y) gives Gx⊆Ax and hence Hx=Ax∖Gx={y∈Yx:R(x,y)=1, c(x,y)=0},Ax=Gx⊔Hx,(16) where ⊔ denotes disjoint union. These memberships stay fixed throughout training. Their probabilities can change, and Hx may be empty.

C.1.2. Acceptance and the Hacked Share

At a fixed prompt, conditioning on acceptance gives qx(θ)=Prπθ(Y∈Hx∣x)px(θ),Prπθ(Y∈Hx∣x)=px(θ)qx(θ),Prπθ(Y∈Gx∣x)=px(θ)(1−qx(θ)),when px(θ)>0.(17) If px=0⁠, the probability of an accepted error is zero and qx is undefined. Since the binary verifier rewards every response in Ax⁠, iterated expectation gives JR(θ)=Ex∼D[EY∼πθ(⋅∣x)R(x,Y)]=Ex∼D[px(θ)].(18) The objective therefore depends on total acceptance, without distinguishing its correct and incorrect parts.

C.1.3. Exact Gradient Flow

Along the flow (1), the chain rule gives J˙R=∇θJR⊤θ˙=∥gR∥22≥0,p˙x=∇θpx⊤gR,q˙x=∇θqx⊤gR,(19) where the identity for q˙x applies when px>0⁠. Under the stated assumptions, there is no general guarantee that q˙x≤0 or that p˙x≥0 at every prompt. Thus, monotonic improvement in average acceptance guarantees neither a nonincreasing hacked share nor nondecreasing acceptance at every prompt.

C.2. Technical Details for Section 3

C.2.1. Aggregate Acceptance and Hacking

Write Eθ for expectation under X∼D⁠, Y∣X∼πθ⁠. Since G∪H is the accepted set, p=JR=pG+pH=Ex∼D[px],q=pHp=∫{x:px>0}pxqxD(dx)pwhen p>0.(20) Thus, relative to the prompt distribution D⁠, prompts are weighted by their acceptance probabilities. Prompts with zero acceptance contribute nothing. An aggregate trend need not hold at every prompt. At fixed p⁠, redistributing probability between G and H leaves reward unchanged, although shared parameters can couple their dynamics.

C.2.2. Proof of Proposition 3.1

Skalse et al. (2022) call a proxy reward hackable if it ranks some pair of policies in strictly opposite orders from the true reward. In our setting, verifier flow generates the policies, acceptance JR=p is the proxy reward, and correctness JC=pG is the true reward. We show that such a pair occurs along the trajectory exactly when, at some time, the correctness lost to a rising hacked share outweighs the correctness gained from rising acceptance.

Proof. Write p(t)=p(θ(t)) and similarly for the other quantities. Since pG=p(1−q)⁠, J˙C=p˙G=(1−q)p˙⏟acceptance contribution−pq˙⏟composition contribution.(21) Consequently, the proposed inequality is equivalent to J˙C<0⁠. The key point is that, along verifier flow, a strict decrease in correctness necessarily accompanies a strict increase in verifier reward.

Sufficiency. Suppose that, at some t⋆∈(0,T)⁠, p(t⋆)q˙(t⋆)>(1−q(t⋆))p˙(t⋆).(22) Then J˙C(t⋆)<0⁠. Since J˙C(t⋆)=∇θJC(θ(t⋆))⊤gR(θ(t⋆)),(23) we must have gR(θ(t⋆))≠0⁠. Hence J˙R(t⋆)=∥gR(θ(t⋆))∥22>0.(24) By continuity, both strict inequalities hold on a sufficiently short interval [t⋆,t⋆+ε]⊂(0,T)⁠. Integrating over this interval gives JR(θ(t⋆+ε))>JR(θ(t⋆)),JC(θ(t⋆+ε))<JC(θ(t⋆)).(25) Thus the two policies at its endpoints witness reward hacking.

Necessity. Conversely, suppose there exist t1<t2 in [0,T] such that verifier reward increases strictly and correctness decreases strictly. By the mean value theorem, there is some t⋆∈(t1,t2) with p˙G(t⋆)=pG(t2)−pG(t1)t2−t1<0.(26) Substituting p˙G=(1−q)p˙−pq˙ and rearranging yields p(t⋆)q˙(t⋆)>(1−q(t⋆))p˙(t⋆),(27) as required. ◻

The criterion distinguishes a growing hacked share from reward hacking in the sense of Skalse et al. (2022). A growing hacked share can coexist with improving correctness if acceptance rises fast enough. The two rewards disagree exactly when, at some time along the flow, the correctness lost to a rising hacked share exceeds the correctness gained from rising acceptance. This inequality needs to hold strictly at only a single time: by continuity, it then holds on a short interval whose endpoints witness hackability, even if correctness later recovers.

C.2.3. Proof of Proposition 3.2

When pH,pG>0⁠, total acceptance cancels in the ratio: z=log⁡pHpG=log⁡q1−q,dzdq=1q(1−q)>0.(28) The log odds z therefore increase with the hacked share q⁠. They are finite only when correct responses and hacks both have positive probability, so we require this condition wherever we use finite log odds.

Evolution of the log odds.

We derive (4) using the gradients s¯S=∇θlog⁡Prθ(S) from Section 3.

Differentiate the log ratio.

For pH,pG>0⁠, the log odds z=log⁡(pH/pG) are finite, and differentiating along the flow gives z˙=p˙HpH−p˙GpG.(29) The rate therefore compares the proportional growth of the two groups. This identity needs only differentiability of the group probabilities.

To express each proportional rate through the policy, we use the score sθ(x,y):=∇θlog⁡πθ(y∣x)⁠. For a fixed group S with positive probability and indicator 1S⁠, the conditions for differentiating expectations give ∇θPrθ(S)=Eθ[1Ssθ]=Prθ(S)s¯S,s¯S:=Eθ[1Ssθ]Prθ(S)=Eθ[sθ∣S].(30) The gradient of log⁡Prθ(S) is thus the mean score s¯S of the group. Applying this to H and G⁠, and using θ˙=gR⁠, gives ∇θz=s¯H−s¯G,z˙=(s¯H−s¯G)⊤gR.(31) The log odds therefore grow precisely when gR⊤s¯H>gR⊤s¯G⁠.

Decomposition of the growth rate.

Assume G,H,N all have positive probability and evaluate quantities at the current policy.

Step 1: Use normalization.

Differentiating the accepted probability and the sum of all group probabilities gives gR=p[(1−q)s¯G+qs¯H],gR+(1−p)s¯N=0.(32) Combining these identities yields gR=p(1−p)[(1−q)s¯G+qs¯H−s¯N].(33)

Step 2: Substitute into the rate of the log odds.

Using (4) and collecting s¯H−s¯G gives z˙=p(1−p)(s¯H−s¯G)⊤[(1−q)s¯G+qs¯H−s¯N]=p(1−p)[(s¯H−s¯G)⊤(s¯G−s¯N)⏟correctness-to-hack leakage+q∥s¯H−s¯G∥2⏟hack bias],(34) which establishes (5) and reveals the mechanisms. The first term can favor or oppose hacking. The second is nonnegative and proportional to q at fixed p and gradients; those quantities also change during training. Hence growth depends on the full bracket, without implying acceleration or growth at every policy.

Step 3: State the growth condition.

When s¯H≠s¯G⁠, the bracket is positive exactly when q>−(s¯H−s¯G)⊤(s¯G−s¯N)∥s¯H−s¯G∥2.(35) The threshold may lie outside (0,1) and change during training. If the two gradients agree, z˙=0⁠. Thus the current share alone does not determine growth. The alignment of the gradients also matters.

Simultaneous growth of reward and hacking

For pH,pG>0⁠, differentiating z=log⁡(q/(1−q)) and using the verifier flow gives q˙=q(1−q)z˙,p˙=∥gR∥2.(36) Since pG,pH,pN>0⁠, both p and q lie in (0,1)⁠. Thus, q˙>0 if and only if the bracket in (5) is positive. By (4), a positive z˙ requires gR≠0⁠. Both q˙ and p˙ are then strictly positive, completing the proof of Proposition 3.2. ◻

Example: Shared normalization can favor hacking.

For one prompt with a response in each group, use softmax logits θ=(θG,θH,θN)⁠: Prθ(S)=eθSeθG+eθH+eθN,S∈{G,H,N}.(37) Here z=θH−θG and gR=(pG(1−p), pH(1−p), −p(1−p))⊤,z˙=(1−p)(pH−pG).(38) Thus pH>pG gives simultaneous growth of the hacked share and verifier acceptance on a sufficiently short interval. The example illustrates a mechanism caused by the parameterization; it does not assert that every neural policy follows this pattern.

C.3. Technical Details for Section 4

C.3.1. Comparison Setup

Fix the prompt distribution, verifier, policy family, initialization, and training procedure. The record Lt contains prompts, responses, verifier rewards, policy updates, and quantities computed from these observations. Training and monitoring may adapt to the record, but neither uses the hidden correctness assignment or additional correctness feedback.

We compare two fixed correctness assignments: c0=R,c1∈CRwithPrθ(R=1,c1=0)>0,(39) at the policy θ⁠. Under c0⁠, every accepted response is correct. Under c1⁠, some accepted responses are hacks.

The assignments are fixed before training. We compare what the same procedure observes under each assignment. Correctness does not change during either run. For the deterministic argument below, fix all random choices used by training and use those same choices in both runs.

C.3.2. Proof of Proposition 4.1

The proof has three steps: the assignments produce the same record, the monitor therefore makes the same decision, but correct detection requires different decisions.

Step 1: The training records are identical.

We argue by induction on the training step. Both runs start from the same initialization, so their initial records agree. Suppose their records agree through step n⁠. Because the training procedure depends only on the record and the shared random choices, both runs select the same prompt and sample the same response from the same policy. The fixed verifier returns the same reward. Both runs therefore make the same update and append the same information to the record, so their records agree through step n+1⁠.

By induction, the records are identical at every training step: Ln(c0)=Ln(c1)for every n.(40) The argument also covers verifier queries that depend on earlier results and any quantity computed from the record. It likewise covers JR and its gradients, which depend on the verifier and the policy but not on which correctness assignment in CR is true. For training in continuous time, we assume that the fixed procedure determines a unique trajectory. Because this trajectory depends only on the verifier and the policy, changing the correctness assignment within CR leaves it unchanged. This argument applies to any two fixed assignments in CR⁠. It does not require either assignment to equal R⁠.

Step 2: The monitor makes the same decision.

A deterministic detection monitor computes a binary alarm Dt=δt(Lt)⁠, where Dt=1 indicates the presence of hacks. Identical records give identical alarms: Dt(c0)=δt(Lt(c0))=δt(Lt(c1))=Dt(c1).(41)

Step 3: The same decision cannot be correct in both cases.

Under c0⁠, there are no hacks, so a correct detector must remain silent. Under c1⁠, hacks have positive probability, so a correct detector must raise an alarm. Because the monitor makes the same decision in both cases, it either raises a false alarm under c0 or misses the hacks under c1⁠.

Thus, no deterministic monitor using only Lt can guarantee correct detection under both compatible assignments. More observations from the same verifier do not resolve this ambiguity: Step 1 continues to apply as the record grows.

Scope.

This conclusion concerns a guarantee over the stated class CR⁠. A monitor may succeed under a particular assignment. Additional knowledge connecting response content to correctness may also rule out alternatives. The impossibility applies when the two assignments remain admissible and the training procedure never observes information that distinguishes them.

Remark on randomization

Randomization does not change the conclusion. Use the same random choices for training and monitoring in the two runs. For every such choice, the records and alarms remain identical. Averaging over these choices therefore gives Prc=c1(Dt=1)=Prc=c0(Dt=1).(42) Here the probabilities are over training and monitor randomness. Each correctness assignment remains fixed. The left side is the detection rate when hacks exist, and the right side is the false-alarm rate when they do not. Thus, the monitor has no detection advantage between these two assignments.

Likewise, the training records have the same distribution: Lawc=c1(Lt)=Lawc=c0(Lt),(43) where Lawc(Lt) denotes the distribution of the record over training randomness under assignment c⁠. This proves the proposition. ◻

C.3.3. Proof of Proposition 4.2

The proof compares two compatible assignments that disagree on every accepted response. Both admit hacks, so knowing that hacks exist does not distinguish them.

Step 1: Construct two opposite correctness assignments.

By the premise, choose a fixed assignment c1∈CR under which the policy θ assigns positive probability to both correct responses and hacks. Define c2(x,y)=R(x,y)−c1(x,y).(44) On rejected responses, both assignments equal zero. On accepted responses, c2=1−c1⁠, so the assignments give opposite correctness labels. Thus c2∈CR⁠, and the two assignments exchange the correct and hacked groups: pH,c2(θ)=pG,c1(θ)>0,pG,c2(θ)=pH,c1(θ)>0.(45) In particular, both assignments admit hacks with positive probability.

Step 2: The monitor makes the same prediction.

Fix any accepted pair (x,y) and any deterministic monitor using only verifier feedback records. As shown in Appendix C.3.2, the two assignments produce identical training records when the same random choices are used. The monitor therefore makes the same prediction in both cases: c^t(c1)(x,y)=c^t(c2)(x,y).(46)

Step 3: The prediction cannot be correct under both assignments.

Because R(x,y)=1⁠, the true labels satisfy c2(x,y)=1−c1(x,y).(47) The monitor’s common binary prediction therefore matches exactly one of the two labels. It must be incorrect under the other assignment.

Thus, no deterministic monitor can guarantee the correct label of an accepted response under every compatible assignment. Knowing that hacks exist does not resolve the ambiguity, because both assignments satisfy that additional information.

Randomization and the error bound.

The same reasoning applies when training or monitoring is randomized. Use the same random choices in both runs. For every such choice, the monitor gives the same prediction under both assignments, and exactly one assignment makes that prediction wrong.

Averaging over these choices gives Prc=c1(c^t(x,y)≠c1(x,y))+Prc=c2(c^t(x,y)≠c2(x,y))=1.(48) Here the probabilities are over training and monitor randomness, with the queried pair and correctness assignments held fixed. At least one of these error probabilities is therefore at least 1/2⁠.

We fixed both candidate assignments before training, so neither depends on the realized training record. Which of the two attains the bound may depend on the monitor and the queried pair, but not on that record. Because both assignments contain hacks, the proposition holds even when the monitor knows that hacks exist. This completes the proof .◻

C.3.4. Proof of Proposition 4.3

Controlled dynamics and the selective objective.

Consider a controller that chooses a correction u(t) using only the training record Lt⁠. Training then follows θ˙=gR+u(t).(49) We assume that the controlled trajectory is well defined and that the group probabilities are differentiable along it. For a deterministic argument, fix all random choices used during training.

Under a correctness assignment c⁠, selective control requires p˙H,c<0,p˙G,c≥0whenever pH,c>0.(50) These conditions concern the full update: p˙H,c=∇θpH,c⊤(gR+u),p˙G,c=∇θpG,c⊤(gR+u).(51) A correction that opposes hack growth is insufficient unless the full update gR+u(t) reduces hack probability while preserving correctness. A uniform guarantee requires the same controller rule to satisfy these conditions under every compatible correctness assignment and at every time along the evolving policy trajectory whenever hack probability is positive. The correction itself may adapt to the training record.

Impossibility of guaranteed selective control.

We compare two assignments that exchange the correct and hacked groups. The controller sees the same record under both assignments, but selective control requires opposite changes in their probabilities.

Step 1: Exchange the correct and hacked groups.

By the premise, choose a fixed assignment c1∈CR under which the initial policy assigns positive probability to both correct responses and hacks. Define c2=R−c1.(52) As in Appendix C.3.3, both assignments are compatible with the verifier, and they exchange the two accepted groups: pH,c2=pG,c1,pG,c2=pH,c1.(53) Both groups have positive probability initially. By continuity, they remain positive on a sufficiently short initial interval.

Step 2: The controller produces the same update.

We extend the argument of Appendix C.3.2, which shows that the two runs produce identical records, to include the controller. Both runs begin at the same policy. Whenever their records agree, the controller chooses the same correction, because it uses only the record. The verifier gradient also agrees, because it depends on the policy and the verifier but not on the correctness assignment. Both runs therefore make the same corrected update, and their records remain identical.

Thus, under both assignments, the fixed training and control procedures generate the same record and the same trajectory of the controlled policy. In particular, both runs have the same velocity θ˙=gR+u at every time.

Step 3: Selective control requires incompatible signs.

Along this common trajectory, exchanging the groups also exchanges their derivatives: p˙H,c2=p˙G,c1,p˙G,c2=p˙H,c1.(54) On the initial interval where both groups remain positive, selectivity under the two assignments therefore requires c1:p˙H,c1<0,p˙G,c1≥0,c2:p˙G,c1<0,p˙H,c1≥0.(55) Each derivative would have to be both negative and nonnegative. No update can satisfy both rows.

Thus, no controller that uses only the training record can guarantee selective control under every correctness assignment in CR⁠. The obstruction appears already on an initial interval of training, so it does not depend on behavior later in training.

Remark on randomization.

The same obstruction applies to a randomized controller. We couple the two runs by using the same random choices for training and control under both assignments. For each realization, the records, corrections, and trajectories then coincide, so no realization satisfies the requirements for selective control under both assignments.

A guarantee that holds with probability one under both assignments would require both sets of conditions to hold on a common event of probability one. Since no realization satisfies both, that event is empty, so no such guarantee exists. Randomization cannot help, because the two candidate assignments are fixed before training and the random choices therefore carry no information that distinguishes them. The proposition follows.◻

Scope.

The result rules out a uniform guarantee over CR⁠, but a controller may still achieve selective control under a particular assignment. Additional information or assumptions can distinguish assignments in CR⁠, and if they exclude either of the two assignments used in the proof, the argument above no longer applies.

Connection to secure control.

The argument here uses an indistinguishability principle also used when studying attack detection and identification in cyber-physical systems: no monitor can distinguish admissible scenarios that produce identical observations when it uses only those observations (Pasqualetti, Dörfler, and Bullo 2013). In our paper, the indistinguishable scenarios are correctness assignments that produce the same training record Lt⁠. Exchanging the correct and hacked groups between two such assignments makes their requirements for selective control incompatible. The argument therefore identifies the information that a uniform guarantee needs: observations or structural assumptions that exclude one of these assignments. In power grid security, models of grid operation supply such structural correctness knowledge: by relating coordinated cyberattacks to their physical consequences, they tell defenders which scenarios cause harm and let them target defense accordingly (Moya and Wang 2018). Audits can play an analogous role by supplying correctness information that distinguishes otherwise compatible assignments. The subsequent analysis examines when this information supports selective control.

C.3.5. Regularization and Its Limitations

We analyze three regularization penalties and distinguish what each controls from what it guarantees about correctness. Throughout, we use exact gradient ascent, fixed penalty weights, and the fixed verifier from Section 2. We then apply Proposition 4.3 to determine which guarantees of selective control these penalties can provide.

KL divergence from a reference

For a fixed reference policy under the same prompt distribution, define K(θ)=Ex∼DDKL(πθ(⋅∣x)‖πref(⋅∣x)),Fβ=JR−βK,(56) where β>0 is fixed.

Step 1: Bound deviation from the reference.

Assume the flow remains in a region where K is finite and continuously differentiable. Gradient ascent on Fβ gives θ˙=gR−β∇θK,F˙β=∥∇θFβ∥2≥0.(57) Thus Fβ(θ(t))≥Fβ(θ(0))⁠. Rearranging and using JR=p≤1 yields K(θ(t))≤K(θ(0))+p(θ(t))−p(θ(0))β≤K(θ(0))+1−p(θ(0))β.(58) In particular, initialization at the reference gives K(θ(0))=0⁠. The penalty therefore bounds departure from the reference along the exact flow.

Step 2: Bound changes in hack probability.

Because the prompt distribution is fixed, K also equals the KL divergence between the joint distributions of prompts and responses. Pinsker’s inequality therefore gives, for the fixed hacked set H⁠, |pH(θ)−pH,ref|≤K(θ)2,pH,ref:=Prref(H).(59) Thus, closeness to the reference limits the change in hack probability.

Limitation.

The penalty bounds how much the policy changes, not the direction of that change. It therefore guarantees neither p˙H<0 nor p˙G≥0⁠. The bound controls hack probability relative to pH,ref⁠. It implies a small absolute hack probability when both the reference’s hack probability and the KL bound are small. The penalty depends on the policy and the reference but not on the correctness assignment, so with the same reference under every assignment in CR⁠, Proposition 4.3 still applies.

Penalizing the gradient of reward.

Following the objective of Ackermann et al. (2026), we consider the idealized gradient regularization penalty FGR=JR−γ∥gR∥2,(60) with fixed γ>0⁠.

Step 1: Derive the correction.

Assume JR is twice continuously differentiable and write BR:=∇θ2JR⁠. Since BR is symmetric, ∇θ∥gR∥2=2BRgR.(61) Gradient ascent on FGR therefore gives θ˙=gR−2γBRgR.(62) The correction depends entirely on the verifier objective and its derivatives.

Step 2: Separate reward sensitivity from correctness.

The norm ∥gR∥ measures first-order sensitivity of expected verifier reward to parameter changes. It does not by itself determine correctness. For example, if the verifier accepts every response, then JR≡1 and gR≡0⁠. The correction vanishes even if the policy assigns positive probability to incorrect responses.

Scope of the cited analysis.

Ackermann et al. (2026) connect flatter optima to the accuracy of the proxy reward under additional assumptions: continuous actions, a Gaussian policy with fixed covariance, regularity conditions on the policy and reward functions, and a true reward that is Lipschitz continuous. Our setup does not impose these policy and reward assumptions, and CR places no corresponding regularity restriction on correctness assignments. Thus, the cited results do not establish selective control uniformly over CR⁠.

Exchanging correctness assignments in CR leaves JR⁠, its derivatives, and the correction unchanged, so Proposition 4.3 applies to this regularizer. This conclusion does not contradict the theoretical results of Ackermann et al. (2026), which hold under the additional assumptions above, or their empirical improvements, which concern particular tasks rather than a uniform guarantee over CR⁠.

Stability of relative probabilities.

Motivated by tie training for reducing reliance on spurious features in preference optimization (Moya et al. 2026), we analyze a simplified penalty on changes in relative response probabilities. The penalty discourages deviations from a fixed reference ratio between two verifier-accepted responses. Unlike pairs constructed to have equal utility, equal verifier rewards alone do not establish that the responses are equally correct. We therefore examine what this penalty guarantees when pair selection uses only verifier information.

Fix two accepted responses y1,y2 to a prompt x⁠, with positive probabilities under the current and reference policies. Define e(θ)=log⁡πθ(y1∣x)πθ(y2∣x)−log⁡πref(y1∣x)πref(y2∣x).(63) Thus e=0 means that the current relative probability matches the reference ratio. Assume e is continuously differentiable on a neighborhood of the trajectory.

Step 1: Derive the restoring term.

For fixed λ>0⁠, gradient ascent on FI=JR−λ2e2(64) adds the correction u(t)=−λe∇θe.(65) Along the corrected flow, e˙=b−λae,b:=∇θe⊤gR,a:=∥∇θe∥2.(66) The term b is the change in the log ratio induced by verifier training. The term −λae opposes deviation from the reference ratio.

Step 2: Bound the deviation.

Consider an interval [0,T] on which the flow exists and both response probabilities remain positive. Suppose a(t)≥m>0,|b(t)|≤Bfor 0≤t≤T.(67) An integrating factor gives e(t)=exp(−λ∫0ta(s)ds)e(0)+∫0texp(−λ∫τta(s)ds)b(τ)dτ.(68) Using the bounds on a and b⁠, |e(t)|≤e−λmt|e(0)|+B∫0te−λm(t−τ)dτ=e−λmt|e(0)|+Bλm(1−e−λmt).(69) The bound separates two contributions: one from the initial deviation, which decays, and one from the drift b that the verifier induces. If b≡0⁠, the deviation decays exponentially. If the drift persists, the bound allows a nonzero deviation to persist as well.

This is a scalar estimate of the form used in input-to-state stability analysis (Khalil and Grizzle 2002). Here, the conclusion concerns e under the stated bounds. It guarantees neither stability of all policy parameters nor exact preservation of the ratio.

Limitation: Restoration can increase errors.

Consider one prompt with two responses, y1 correct and y2 incorrect, both accepted. Let pG(θ)=11+e−θ,pH(θ)=1−pG(θ).(70) For a reference parameter θref⁠, we have JR≡1 and e=θ−θref⁠. Thus, θ˙=−λe,p˙G=−λepGpH,p˙H=λepGpH.(71) If θ(0)>θref⁠, then e(t)=e(0)e−λt>0 at every finite time. Restoration therefore decreases correctness and increases hack probability while reducing the deviation from the reference ratio.

The current policy initially favors the correct response more strongly than the reference does, so restoring the reference reverses that improvement. If initialized at the reference instead, the flow remains stationary and does not strictly reduce hack probability.

Common limit of the three penalties.

We hold the reference policies, penalty weights, and selected response pairs fixed across correctness assignments and give the penalties no additional correctness information. At the same policy, each penalty and its gradient then agree under every assignment in CR⁠. The argument about selective control above, which shows that the records coincide, therefore gives the same controlled trajectory under the compared assignments.

Under the assumptions of Proposition 4.3, none of these penalties can guarantee p˙H<0 and p˙G≥0 along the trajectory under every assignment in CR whenever pH>0⁠. A penalty may control the departure from a reference, the sensitivity of the reward, or the relative probabilities of paired responses without providing this uniform guarantee of selective control, although it may still succeed on particular tasks.

Additional correctness information or justified structural assumptions can restrict CR and exclude indistinguishable assignments that require incompatible corrections. Excluding these assignments removes the obstruction exhibited here, even without identifying every hack. Removing the obstruction does not by itself give a guarantee, however: a guarantee of selective control also requires showing that the available updates satisfy p˙H<0 and p˙G≥0 throughout the trajectory.

C.4. Technical Details for Section 5

C.4.1. Projected Correction and Its Immediate Effects

Write hA=∇θpHA and h=∇θpH⁠. Since p=pH+pG⁠, we have ∇θpG=gR−h⁠. We first derive the effects of the projected correction without assuming full audit coverage.

Step 1: Preserve instantaneous acceptance growth.

Define P⊥={Id−gRgR⊤∥gR∥2,gR≠0,Id,gR=0.(72) This orthogonal projector satisfies P⊥⊤=P⊥⁠, P⊥2=P⊥⁠, and gR⊤P⊥=0⁠. For the correction u=−λP⊥hA⁠, with λ>0⁠, we therefore have gR⊤u=0.(73) Along the corrected flow θ˙=gR+u⁠, the chain rule gives p˙=gR⊤(gR+u)=∥gR∥2.(74) Thus the correction preserves the instantaneous acceptance growth produced by the verifier gradient at the current policy.

Step 2: Oppose growth of audited hacks.

Symmetry and idempotence of the projector give hA⊤u=−λhA⊤P⊥hA=−λ∥P⊥hA∥2.(75) Consequently, p˙HA=hA⊤gR−λ∥P⊥hA∥2.(76) The correction contributes a strictly negative term when P⊥hA≠0⁠. The full rate is negative only when this contribution outweighs any positive growth induced by the verifier gradient.

Step 3: Relate the correction to the full population.

For the full hack probability, the correction contributes h⊤u=−λ(P⊥h)⊤(P⊥hA).(77) Because gR⊤u=0⁠, its contribution to correct responses is equal and opposite: ∇θpG⊤u=(gR−h)⊤u=−h⊤u.(78) Thus, when the two projected gradients have positive inner product, the correction opposes hack growth and contributes equally to correct-response growth. These statements concern the correction’s contribution. Selectivity of the full update also depends on gR⁠.

If P⊥h=0⁠, every correction orthogonal to gR has zero instantaneous effect on pH⁠. The projection therefore requires a component of the hack gradient orthogonal to the reward gradient. The theorem below shows how full coverage and sufficient corrective strength turn these local identities into selective control and a bound along the evolving policy.

C.4.2. Proof of Theorem 5.1

The proof has four steps. Full audit coverage identifies the gradient h=∇pH⁠. Projection removes the component of h along gR⁠, so the correction does not slow the growth of acceptance. A strong enough correction then gives p˙H<0 and p˙G≥0⁠. Finally, integrating these rates over [0,T] bounds the hacked share.

Step 1: Full coverage identifies the gradient of pH⁠.

Under (A2), unaudited hacks have zero probability along the policy trajectory: pH−pHA=0.(79) As a function of θ⁠, this difference is nonnegative everywhere, because audited hacks form a subset of all hacks. At every point of the trajectory, it therefore attains its minimum value of zero, and its gradient vanishes: ∇pHA=∇pH.(80) Thus the audited population gradient coincides with the full hack gradient along the trajectory. The correction becomes u=−λP⊥h⁠, where P⊥ projects onto the orthogonal complement of gR⁠, and we write S=∥P⊥h∥2 for the squared norm of the projected gradient.

Step 2: Derive the corrected dynamics.

The corrected flow is θ˙=gR+u=gR−λP⊥h⁠. Because P⊥ is an orthogonal projection onto the complement of gR⁠, we have gR⊤P⊥h=0 and h⊤P⊥h=∥P⊥h∥2=S (Appendix C.4.1). Differentiating along the flow therefore gives p˙=∥gR∥2,p˙H=h⊤gR−λS,p˙G=∥gR∥2−p˙H.(81) At the current policy, the correction preserves the instantaneous acceptance growth, lowers the rate p˙H by λS⁠, and raises the rate p˙G by the same amount.

To compare the growth of the two groups, recall z=log⁡(pH/pG)⁠. Without the correction, z changes at the rate b=(s¯H−s¯G)⊤gR,(82) which we call the pressure that verifier rewards exert on the hacked share. Taking the logarithmic derivative gives z˙=p˙HpH−p˙GpG=b−λS(1pH+1pG)=b−λSpq(1−q).(83) Since q˙=q(1−q)z˙⁠, we also obtain q˙=q(1−q)b−λSp.(84) These expressions separate the verifier’s growth pressure from the opposing effect of the correction.

Step 3: Identify when the update is selective.

The correction lowers p˙H by λS and raises p˙G by the same amount, so z˙=p˙HpH−p˙GpG=b−λS(1pH+1pG).(85) If the correction term exceeds b⁠, then z˙<0⁠, and hence q˙<0⁠, since z increases with q⁠. Differentiating pG=p(1−q) and using p˙=∥gR∥2 then gives p˙G=(1−q)∥gR∥2−pq˙>0.(86) Because acceptance never decreases under the correction, a falling hacked share implies rising correctness.

A falling hacked share does not require pH itself to fall, since acceptance may still grow. For this stronger conclusion, the correction must overcome the verifier’s contribution to p˙H⁠: λS>h⊤gR.(87) Then p˙H<0 and p˙G=∥gR∥2−p˙H>0⁠. Because pH falls while pG rises, the log ratio z decreases as well, and part (i) follows.

Step 4: Bound the hacked share during training.

Because q is the logistic function of z⁠, the rate from Step 3 gives q˙=q(1−q)z˙=q(1−q)b−λSp,(88) where the second term uses pH=pq and pG=p(1−q)⁠. Assumption (A3) bounds the correction from below, S/p≥κq⁠, and the pressure from verifier rewards is bounded above by b≤D with D≥0⁠. Since q(1−q)≤1/4⁠, these bounds give q˙≤D4−λκq.(89) The first term caps the pressure from verifier rewards, and the second is a correction proportional to the current hacked share.

Rearranging gives q˙+λκq≤D/4⁠, and multiplying by the integrating factor eλκt gives ddt(eλκtq(t))≤D4eλκt.(90) Integrating from 0 to t and rearranging yields q(t)≤e−λκtq(0)+D4λκ(1−e−λκt),t∈[0,T].(91) The initial hacked share contributes a term that decays at rate λκ⁠, while the pressure from verifier rewards contributes a term that rises toward D/(4λκ)⁠. Part (ii) follows, which completes the proof of the theorem. ◻

C.4.3. Estimating the gradient of pH from audits

We fix the current policy and assume that audits reveal exact correctness labels. The estimator weights each observed hack by the inverse of its audit probability, which compensates for hacks that are generated but not audited.

Step 1: Express the gradient as an expectation.

With the policy score sθ(x,y)=∇θlog⁡πθ(y∣x)⁠, and since R(1−c) indicates a hack, differentiating the probability of hacks gives h:=∇θpH=Eθ[R(X,Y)(1−c(X,Y))sθ(X,Y)].(92) Each hack contributes its policy score, and all other responses contribute zero.

Step 2: Weight the observed contributions.

We draw n independent pairs with Xi∼D and Yi∼πθ(⋅∣Xi)⁠. We audit each accepted response with known probability ρi=ρ(Xi,Yi)>0 and never audit rejected responses. The indicator Ii records whether an audit occurs, and when Ii=1⁠, the audit reveals Ci=c(Xi,Yi)⁠. The estimator is ξi={1−Ciρisθ(Xi,Yi),Ii=1,0,Ii=0,h^=1n∑i=1nξi.(93) An audited hack receives weight 1/ρi⁠, while audited correct responses and unaudited responses contribute zero. The average divides by all n generated responses, including those that were not audited.

Step 3: Show why the weighting works.

For an accepted response, an audit occurs with probability ρi⁠, and the weight 1/ρi cancels this probability in expectation. For a rejected response, both sides are zero. Hence E[ξi∣Xi,Yi]=R(Xi,Yi)(1−c(Xi,Yi))sθ(Xi,Yi),(94) and averaging over the generated responses gives E[h^]=h,(95) where the expectation covers both response generation and audit selection. Audits of only some responses therefore give an unbiased estimate of h⁠, provided that every accepted response has a positive audit probability.

Consequence for the projected correction.

We keep gR and P⊥ exact and fix λ>0 before drawing the batch. Every realized correction u^=−λP⊥h^ then satisfies gR⊤u^=0⁠, so it preserves the instantaneous acceptance growth at the fixed policy. Moreover, E[h⊤(gR+u^)]=h⊤gR−λh⊤P⊥E[h^]=h⊤gR−λ∥P⊥h∥2.(96) At the fixed policy, the correction from a batch therefore changes pH at the same expected rate as the population correction. A single batch can still produce a direction that does not decrease pH⁠.

Under the stated assumptions, the estimator is unbiased. Small audit probabilities, however, produce large weights 1/ρi and can increase sampling variability. The argument does not cover accepted responses with zero audit probability or audits with systematic label errors. Appendix A.2 discusses the practical consequences of both cases.

D. Experimental Details

D.1. The Gaussian Contextual Bandit Experiment

Motivation.

To test our theoretical predictions, we use a controlled contextual bandit in which we know the correctness of every response and can compute population gradients exactly.

Setup.

Policy.

The policy assigns each response a logit using shared parameters θ∈R4⁠, fixed features ϕ(x,y)∈R4⁠, and a fixed offset a(x,y)⁠: πθ(y∣x)=exp⁡(θ⊤ϕ(x,y)+a(x,y))∑y′∈Yxexp⁡(θ⊤ϕ(x,y′)+a(x,y′)).(97) Training changes only θ⁠. Sharing parameters allows updates from different prompts to interact.

Initialization.

We initialize θ(0)=0 and use offsets to set the initial acceptance p(0) and hacked share q(0)⁠: a(x,y)={log⁡[p(0)(1−q(0))/16],y∈Gx,log⁡[p(0)q(0)/16],y∈Hx,log⁡[(1−p(0))/16],y∈Nx.(98) This gives pG(0)=p(0)(1−q(0))⁠, pH(0)=p(0)q(0)⁠, and pN(0)=1−p(0)⁠, with equal probabilities within each group. Changing the offsets varies these probabilities while keeping the features fixed.

Prompt and Group Distributions.

Prompts and labels.

We use eight equally weighted prompts, each with 16 responses in each of three groups. Correct responses Gx have (R,c)=(1,1)⁠, hacks Hx have (R,c)=(1,0)⁠, and rejected responses Nx have (R,c)=(0,0)⁠. Group membership and labels remain fixed during training.

Accepted features.

We construct features for hacks by centering Gaussian noise and adding the mean (1,1,0,0)⊤⁠. We obtain features for correct responses by negating the second coordinate. Writing ϕS,x,j for response j in group Sx⁠, we draw εx,j∼iidN(0,0.252I4),j=1,…,16,(99) and set ϕH,x,j=(1,1,0,0)⊤+εx,j−116∑k=116εx,k,ϕG,x,j=diag(1,−1,1,1)ϕH,x,j.(100) The empirical means are exactly (1,1,0,0)⊤ for hacks and (1,−1,0,0)⊤ for correct responses in every prompt.

Rejected features.

We draw eight Gaussian vectors per prompt and append copies with the second coordinate negated: ξx,j∼iidN(0,0.252I4),vx,j=ξx,j,vx,j+8=diag(1,−1,1,1)ξx,j,(101) for j=1,…,8⁠. We center these vectors and add the desired mean μN=(0,μN,2,0,0)⊤⁠: ϕN,x,j=vx,j−116∑k=116vx,k+μN,j=1,…,16.(102) The empirical mean is therefore exactly μN⁠. The experiments use μN,2∈{−1.5,−0.5,0,0.5}⁠.

We draw the raw noise independently across prompts and between the accepted and rejected groups. Within each seed, we reuse these draws across initializations, rejected means, and learning rates.

Objective.

Training objective.

We maximize verifier acceptance: JR(θ)=18∑x=18∑y∈Yxπθ(y∣x)R(x,y)=p(θ).(103) Correct responses and hacks receive the same reward. Correctness labels define the controlled initializations and evaluation metrics; they do not enter the training objective.

Training dynamics.

We evaluate the objective and gradient by summing over every prompt and response. We numerically integrate the verifier flow (1), θ˙=gR=∇θJR,(104) using DOP853 until T=50⁠. For the learning rate ablation, we instead take discrete updates: θk+1=θk+ηgR(θk).(105) Both procedures use exact population gradients.

Metrics.

Acceptance and correctness.

We compute each group’s probability over all prompts: pS(θ)=18∑x=18∑y∈Sxπθ(y∣x),S∈{G,H,N}.(106) We report acceptance p=pG+pH⁠, hacked share q=pH/p⁠, and correctness pG⁠. We form q after aggregating over prompts. The proxy and true objectives are JR=p and JC=pG⁠.

Log odds and group scores.

We track z=log⁡(pH/pG) and compute the group scores s¯S=∇θlog⁡pS exactly: s¯S=18pS∑x=18∑y∈Sxπθ(y∣x)[ϕ(x,y)−∑y′∈Yxπθ(y′∣x)ϕ(x,y′)].(107) We subtract each prompt’s mean feature before aggregating. Although the features remain fixed, the group scores change as the policy re-weights responses.

Hack bias and leakage.

We compute the two contributions in (5): z˙=p(1−p)[(s¯H−s¯G)⊤(s¯G−s¯N)⏟leakage+q∥s¯H−s¯G∥2⏟hack bias].(108) Their sum determines the sign of z˙ and hence the direction of change in the hacked share. We evaluate the full expression within each seed before averaging.

Growth and reward hacking.

We compute the rates of acceptance, hacked share, and correctness: p˙=∥gR∥2,q˙=q(1−q)z˙,p˙G=(1−q)p˙−pq˙.(109) We identify declining correctness using pq˙>(1−q)p˙ and locate stationary correctness where p˙G=0⁠. These quantities distinguish growth of the hacked share from a decrease in the probability of correct responses.

Variation across seeds.

The appendix experiments use ten independent Gaussian feature draws and report means and one sample standard deviation. Each draw defines deterministic training. For the comparison of correctness against acceptance, we interpolate each trajectory onto acceptance values reached by every seed and geometry before computing these statistics.

D.1.1. When Hacking Grows

Controlled comparisons.

We fix p(0)=2/3⁠. In the main paper, panel (a) tracks p⁠, q⁠, and pG with q(0)=0.3 and μN,2=−0.5⁠. Panel (b) compares JC against JR for q(0)∈{0.1,0.3,0.5} using the same features.

The centered construction lets us vary initial leakage and hack bias separately. At initialization, (s¯H−s¯G)⊤(s¯G−s¯N)=−2(1+μN,2),q(0)∥s¯H−s¯G∥2=4q(0).(110) Changing the rejected mean changes initial leakage. Changing the initial hacked share changes initial hack bias while preserving the differences between group scores. The following experiments examine how these contributions evolve during training.

Results.

Experiment I: hack bias and correctness-to-hack leakage.

Starting from initial acceptance p(0)=2/3⁠, hack share q(0)=0.3⁠, and mean rejected strength μN,2=−0.5⁠, we track hack bias, leakage from correctness to hacking, and their sum. For the same trajectories, we plot z˙(t) and compare it with central time differences of z⁠. Together, these experiments relate the competing contributions to the growth rate of hacking through the factor p(θ)(1−p(θ)) in (5).

Figure 3(a) shows that hack bias increases while leakage from correctness to hacking remains near −1⁠. This value matches initialization, where the centered feature means give (s¯H−s¯G)⊤(s¯G−s¯N)=−2(1+μN,2)=−1.(111) Leakage stays near this value, consistent with the similar Gaussian spreads across groups: policy reweighting shifts their mean features similarly, leaving the contrasts between groups approximately unchanged. Negative leakage opposes growth of the hacked share, but hack bias outweighs it throughout the plotted interval. Their sum therefore stays positive, and Proposition 3.2 predicts z˙>0 and an increasing hacked share. The predicted rate agrees with central time differences in Figure 3 (b). Thus, negative leakage can oppose hacking without overcoming the reinforcement from hack bias.

Growth of hacking in the Gaussian contextual bandit. (a) Hack bias outweighs negative correctness-to-hack leakage, keeping their sum positive. (b) The resulting growth rate \dot z > 0 of the log odds agrees with central time differences along the same trajectories. (c) Correctness against acceptance for three values of the mean rejected strength \mu_{N,2}. Changing the rejected features changes whether and when correctness declines as verifier reward rises. (d) Discrete gradient ascent approaches the reference flow as the learning rate decreases. We compare iterate k with the flow at t=k\eta and report, for each variable and seed, the maximum absolute discrepancy over evaluated times. Curves show means across ten independent Gaussian feature draws, and shading shows \pm one sample standard deviation. Each draw defines deterministic dynamics with exact population gradients.
Figure 3. Growth of hacking in the Gaussian contextual bandit. (a) Hack bias outweighs negative correctness-to-hack leakage, keeping their sum positive. (b) The resulting growth rate z˙>0 of the log odds agrees with central time differences along the same trajectories. (c) Correctness against acceptance for three values of the mean rejected strength μN,2⁠. Changing the rejected features changes whether and when correctness declines as verifier reward rises. (d) Discrete gradient ascent approaches the reference flow as the learning rate decreases. We compare iterate k with the flow at t=kη and report, for each variable and seed, the maximum absolute discrepancy over evaluated times. Curves show means across ten independent Gaussian feature draws, and shading shows ± one sample standard deviation. Each draw defines deterministic dynamics with exact population gradients.
Ablation I: varying the mean rejected strength.

We keep p(0)=2/3 and q(0)=0.3 fixed and vary μN,2∈{−1.5,−0.5,0.5}⁠, holding the accepted features and underlying Gaussian draws fixed. Since initial leakage equals −2(1+μN,2)⁠, these values give initial leakage of +1⁠, −1⁠, and −3⁠: positive, negative but weaker than the initial hack bias, and negative and stronger than it. For each condition, we plot correctness JC=pG against verifier reward JR=p⁠. Plotting against reward rather than time compares the conditions at the same acceptance level, which isolates how rejected responses affect correctness as verifier reward rises.

Figure 3 (c) shows how the mean rejected strength changes the relation between verifier reward and correctness. For μN,2=−1.5⁠, where initial leakage is positive, correctness declines after a brief initial increase, so reward hacking emerges early. For μN,2=−0.5⁠, where negative leakage is weaker than hack bias, acceptance and correctness first improve together, but correctness later decreases despite further gains in acceptance. For μN,2=0.5⁠, where negative leakage is stronger, both quantities increase throughout the plotted range, with no reward hacking evident in the mean curve. These declines illustrate Proposition 3.1: reward hacking occurs when pq˙>(1−q)p˙⁠, that is, when the shift toward hacked responses outweighs the correctness gained from rising acceptance. Thus, even with identical initial acceptance and hacked share, the geometry of the rejected responses can change whether and when reward hacking emerges.

Ablation II: learning rate in exact gradient ascent.

To test how closely gradient flow approximates discrete training, we keep p(0)=2/3 and q(0)=1/2 fixed and vary the learning rate η⁠, holding the features and initialization fixed within each seed and geometry. We compare iterate k of each discrete trajectory with the reference flow at the matched time t=kη⁠. For each seed, we take the maximum discrepancy in z⁠, p⁠, q⁠, and pG over times, and we then average these maxima across seeds.

Figure 3 (d) shows that discrepancies in z⁠, p⁠, q⁠, and pG increase with the learning rate, reaching the order of 10−2 at η=0.4⁠. Because the updates use exact population gradients, these discrepancies reflect only the effect of taking finite steps. As the learning rate decreases, the discrepancies shrink, which supports using gradient flow to approximate discrete training over the tested interval.

Takeaway.

The verifier rewards every accepted response, whether correct or hacked. Which group benefits therefore depends on the feature geometry, including the rejected responses that shape the direction in which acceptance increases. Along this direction, hack bias can overcome negative leakage and shift the accepted population toward hacked responses. Correctness falls once this shift outweighs the correctness gained from rising acceptance. Because discrete gradient ascent with small learning rates closely tracks gradient flow, the flow offers a reliable way to study this competition.

Hyperparameters.

Table 1 lists the experimental settings.

Table 1. Settings for the Gaussian contextual-bandit growth experiments.

Parameter Value
Number of prompts 8
Responses per group per prompt 16
Feature dimension 4
Gaussian noise standard deviation 0.25
Initial parameters θ(0)=0
Independent feature seeds 10 (seeds 0,…,9⁠)
Initial acceptance p(0)=2/3
Initial hacked share, Experiment I and Ablation I q(0)=0.3
Initial hacked share, Ablation II q(0)=1/2
Rejected mean, Experiment I μN,2=−0.5
Rejected means, Ablation I μN,2∈{−1.5,−0.5,0.5}
Rejected means, Ablation II μN,2∈{−0.5,0,0.5}
Flow integrator DOP853
Maximum integration step 0.25
Training horizon T=50
Recorded flow times 0,0.1,…,50
Gradient-ascent learning rates η {0.4,0.2,0.1,0.05,0.025}

D.1.2. Limits of Verifier Feedback

We study the limits of selective control from verifier feedback (Section 4.4). We evaluate gradient regularization under two compatible correctness assignments along the same sequence of policies. The assignments exchange correct responses and hacks, so reducing hack probability under one reduces correctness under the other. We vary the regularization weight to measure the fraction of recorded times when control is selective under each assignment.

Additional experimental settings.

We use the policy and feature construction from Appendix D.1 to study control from verifier information, as described in Section 4.4.

Fixed settings.

We set p(0)=2/3⁠, q(0)=1/2⁠, μN=(0,−0.5,0,0)⊤⁠, and θ(0)=0⁠. The offsets are a(x,y)=log⁡(1/48)⁠, giving pG(0)=pH(0)=pN(0)=1/3⁠. We use ten independent feature draws and reuse each draw across regularization weights and correctness assignments. The verifier and initial policy remain fixed across comparisons.

Correctness assignments.

We keep the groups G,H,N fixed and evaluate two assignments: c1=1G and c2=1H=R−c1⁠. Under c1⁠, correct responses belong to G and hacks belong to H⁠. Under c2⁠, these roles reverse: pG,c1=pG,pH,c1=pH,pG,c2=pH,pH,c2=pG.(112) Both assignments use the same verifier. Correctness labels enter only the evaluation.

For panel Figure 1, Panel (c) in the main paper, we also evaluate c=R⁠, under which pG,c=p and pH,c=0⁠, along the same verifier flow used for c1⁠.

Exact derivatives.

Using the group scores computed above, we obtain ∇pS=pSs¯S,S∈{G,H,N},gR=∇pG+∇pH.(113) We compute the reward gradient by summing over all prompts and responses.

Controller.

Gradient regularization penalizes the squared norm of the reward gradient: FGR(θ)=JR(θ)−γ∥gR(θ)∥2,BR(θ)=∇θ2JR(θ).(114) We follow its gradient by adding a correction to the verifier flow: u(t)=−2γBR(θ(t))gR(θ(t)),θ˙=gR+u(t)=∇θFGR.(115) We recompute the derivatives at the current policy and keep γ fixed within each run. The main comparison uses γ=16⁠; the ablation uses γ∈{0,1,4,16}⁠. Setting γ=0 recovers the verifier flow. Because the update uses no correctness labels, both assignments give the same policy πθ(t) at every time.

Numerical integration.

We integrate to T=1 using classical fourth order Runge–Kutta with step 0.01 and record the policy at every step. We check integration accuracy by repeating each run with step 0.005⁠. Table 2 lists the numerical settings.

Selective control.

We evaluate changes in hack probability and correctness using the full parameter update: p˙H,c=∇pH,c⊤(gR+u),p˙G,c=∇pG,c⊤(gR+u).(116) For numerical classification, we require p˙H,c<−ϵ and p˙G,c≥−ϵ⁠, with ϵ=10−8⁠. We verify the exchange of rates across assignments: p˙H,c1=p˙G,c2,p˙G,c1=p˙H,c2.(117) For each seed, we compute the fraction of recorded times satisfying selectivity. We report means and sample standard deviations across seeds for both the rates and these fractions.

Results.

Experiment I: verifier-only control.

We run gradient regularization with γ=16 and evaluate the resulting trajectory under c1 and c2⁠. The controller uses only verifier information, so changing the correctness assignment leaves the trajectory unchanged. We plot p˙G,c and p˙H,c and identify intervals where p˙H,c<0 and p˙G,c≥0⁠.

Figure 4(a) shows that the mean rate p˙G,c1=p˙H,c2 decreases but remains positive, while p˙H,c1=p˙G,c2 rises from negative to positive. Initially, the controller therefore reduces mean hack probability and increases mean correctness under c1⁠, with opposite effects under c2⁠. Later, both mean rates become positive, so neither assignment shows a reduction in mean hack probability. The exchange of rates in (55) prevents selective control under both assignments at the same time, illustrating Proposition 4.3.

Control using only RLVR training observations in the Gaussian contextual bandit. (a) Rates of change in correctness and hack probability under gradient regularization with \gamma=16. The assignments c_1,c_2 exchange these rates along the same sequence of policies. (b) Fraction of recorded times satisfying selective control for each regularization weight \gamma: \dot p_{H,c}<-\epsilon and \dot p_{G,c}\geq-\epsilon, with \epsilon=10^{-8}. We compute each fraction within a seed before averaging. Lines show means across ten independent Gaussian feature draws; shading shows \pm one sample standard deviation.
Figure 4. Control using only RLVR training observations in the Gaussian contextual bandit. (a) Rates of change in correctness and hack probability under gradient regularization with γ=16⁠. The assignments c1,c2 exchange these rates along the same sequence of policies. (b) Fraction of recorded times satisfying selective control for each regularization weight γ⁠: p˙H,c<−ϵ and p˙G,c≥−ϵ⁠, with ϵ=10−8⁠. We compute each fraction within a seed before averaging. Lines show means across ten independent Gaussian feature draws; shading shows ± one sample standard deviation.
Ablation: gradient regularization weight.

We vary γ∈{0,1,4,16} while keeping the features and initial policy fixed within each seed. For each weight, we evaluate both correctness assignments along the same sequence of policies. We compute the fraction of recorded times satisfying selective control under each assignment and report the mean and sample standard deviation across ten seeds. At every evaluated time, we also check that the policy update is never selective under both assignments.

Figure 4 (b) shows that the mean fraction of selective updates under c1 is zero for γ∈{0,1,4} and below 25% for γ=16⁠. Under c2⁠, the fraction is 100% for γ=0⁠, above 50% for γ=1⁠, and zero for γ∈{4,16}⁠. Without regularization, training increases the probability of H and decreases that of G throughout the recorded interval. This is selective under c2⁠, which labels H as correct and G as hacks, but has the opposite effect under c1⁠. Increasing regularization to γ=16 introduces a period of selectivity under c1 while eliminating selectivity under c2⁠. These results illustrate the information limit in Proposition 4.3: a controller using only RLVR observations cannot guarantee selective control across compatible correctness assignments.

Takeaway.

Gradient regularization changes which accepted responses lose probability, but verifier observations do not reveal whether those responses are hacks or correct. The same reduction therefore removes hacks under one compatible assignment and correct responses under the other. Tuning the regularization weight changes which assignment benefits without resolving this ambiguity. Guaranteeing selective control requires additional information that distinguishes the correctness assignments.

Hyperparameters.

Table 2 lists the additional and changed settings for this experiment. All other settings follow Table 1.

Table 2. Additional settings for the verifier-only control experiments (Appendix D.1.2). The main experiment uses γ=16⁠; the ablation varies its weight. Both correctness assignments share the same controlled trajectory. All expectations are evaluated exactly. Other model settings follow Table 1.

Parameter Value
Rejected-group second-coordinate mean −0.5
Initial hacked share q(0) 1/2
Initial acceptance p(0) 2/3
Initial parameters θ(0)=0
Independent feature seeds 10 (seeds 0,…,9⁠)
Main regularization weight γ 16
Regularization-weight sweep {0,1,4,16}
Flow integrator Classical fourth-order Runge–Kutta
Training horizon T=1
Recorded times 0,0.01,…,1
Tolerance for selectivity ϵ 10−8

D.1.3. Selective Control with Correctness Feedback

We studyprojected audit corrections, examine the effects of audit coverage and projection error, and illustrate the ISS bound in Theorem 5.1.

Additional experimental settings.

Features and initialization.

We use the policy and Gaussian construction from Appendix D.1, changing the accepted feature means to (4,−1,0,0)⊤ for G and (4,1,0,0)⊤ for H⁠. The rejected mean remains (0,−0.5,0,0)⊤⁠. These means give ∇pG(0)⊤∇pH(0)=29/324>0⁠, so audit correction without projection (raw) initially opposes correctness. We fix c=c1=1G⁠, p(0)=2/3⁠, q(0)=1/2⁠, and θ(0)=0⁠. Thus, every offset equals log⁡(1/48) and pG(0)=pH(0)=pN(0)=1/3⁠. We reuse features and initialization across methods within each seed.

Audits and exact gradients.

A fixed audit set Aaud reveals correctness labels for selected accepted responses. Its audited hacks are HA=Aaud∩H⁠. We compute their probability gradient exactly: ∇pHA=18∑x∑y:(x,y)∈HAπθ(y∣x)[ϕ(x,y)−∑y′πθ(y′∣x)ϕ(x,y′)].(118) We use full coverage except in the coverage ablation. Under full coverage, pHA=pH and their gradients coincide. Labels outside the audit set enter only the evaluation.

Corrections.

Each method follows θ˙=gR+u⁠, with u={0,verifier flow,−2γBRgR,gradient regularization,−λ∇pHA,raw audit correction,−λP⊥∇pHA,projected audit correction.(119) Here BR=∇2JR and P⊥=I−gRgR⊤/∥gR∥2 when gR≠0⁠, with P⊥=I otherwise. We fix λ=6 and γ=16 and recompute derivatives at the current policy.

Additional metrics.

Selective control.

We retain pG,pH,p⁠, and q from Appendix D.1. We evaluate each method using the full update: p˙H=∇pH⊤(gR+u),p˙G=∇pG⊤(gR+u).(120) For numerical classification, we require p˙H<−ϵ and p˙G≥−ϵ⁠, with ϵ=10−8⁠. In the coverage ablation, we also record (P⊥∇pH)⊤(P⊥∇pHA)⁠. A positive value means that the correction opposes growth of total hack probability.

Projection error.

We construct the correction u^ using an estimated reward gradient g^R inside the projector. The verifier direction gR and hack gradient ∇pH remain exact. We measure the resulting error in acceptance growth: |p˙−∥gR∥2|=|gR⊤u^|=|(gR−g^R)⊤u^|.(121)

Numerical ISS envelope.

For each seed, we estimate D from the maximum of (s¯H−s¯G)⊤gR and zero along the trajectory. We estimate κ from the minimum of ∥P⊥∇pH∥2/pH⁠. We refine the integration and increase the evaluation grid from 1,001 to 2,001 times. We insert the constants D and κ into Theorem 5.1 and compare its envelope with q(t)⁠. These numerical estimates do not certify the bound between evaluated times.

Results

Experiment I: selective control.

Under c=c1 and full audit coverage, we compare verifier flow, gradient regularization, raw audit correction, and projected audit correction. We compute exact gradients and keep the features and initial policy fixed across methods. Both audit corrections use the same fixed gain λ⁠. We plot correctness pG against hack probability pH and identify intervals of selective control, where p˙H<0 and p˙G≥0⁠.

Figure 5 (a) shows that projected audit correction increases correctness pG and decreases hack probability pH throughout the recorded interval. Raw audit correction initially decreases both probabilities. Later, correctness increases while hack probability remains nearly constant. Both verifier flow and gradient regularization increase hack probability. These mean trajectories show that projection enables sustained hack reduction without the initial loss of correctness observed under raw audit correction.

Selective control with correctness feedback in the Gaussian bandit. (a) Correctness p_G versus hack probability p_H; arrows indicate training direction. (b) Hacked share q(t) and the numerical ISS envelope. (c) Final hack probability versus the initial fraction of hacks audited; the dashed line marks the initial hack probability. (d) Error |\dot p-\|g_R\|^2| when estimating the reward gradient used for projection. Lines show means across ten feature seeds; shading shows \pm one sample standard deviation. The envelope in (b) provides a numerical check and does not certify the bound between evaluated times.
Figure 5. Selective control with correctness feedback in the Gaussian bandit. (a) Correctness pG versus hack probability pH⁠; arrows indicate training direction. (b) Hacked share q(t) and the numerical ISS envelope. (c) Final hack probability versus the initial fraction of hacks audited; the dashed line marks the initial hack probability. (d) Error |p˙−∥gR∥2| when estimating the reward gradient used for projection. Lines show means across ten feature seeds; shading shows ± one sample standard deviation. The envelope in (b) provides a numerical check and does not certify the bound between evaluated times.
Experiment II: ISS bound.

We apply projected audit correction with full coverage and constant gain and compare q(t) with the envelope in Theorem 5.1. For each feature seed, we estimate D and κ along the computed trajectory, then check these estimates using a finer time grid and tighter integration tolerances. We construct the envelope separately for each seed and record its gap from q(t)⁠. This comparison illustrates the bound numerically over the tested interval.

Figure 5 (b) shows that the hacked share q(t) decreases throughout the recorded interval. It equals the ISS envelope at initialization and remains below it afterward, providing a numerical illustration of Theorem 5.1 along the tested trajectories.

Ablation I: audit coverage.

We vary the fraction of candidates audited in each accepted group and prompt over {0,1/8,1/4,1/2,3/4,1}⁠. We use the known groups to construct nested audit sets and keep their candidate indices fixed across seeds and throughout training. The controller receives correctness labels only for audited responses. For each fraction, we apply projected audit correction with λ=6 until T=1⁠, using the exact gradient of the audited hack probability pHA⁠. We compare total hack probability pH(T) across initial coverage levels and record the alignment between the projected gradients of pH and pHA⁠.

Figure 5 (c) shows that increasing audit coverage lowers the final hack probability. Without audits, the correction vanishes and verifier flow increases pH⁠. With sufficient coverage, the correction reduces pH below its initial value. Although the correction uses only audited responses, we evaluate whether it reduces the total probability of hacks.

Ablation II: projection error.

At θ=0 and full audit coverage, we estimate the reward gradient using batches of n∈{32,128,512,2048} independent prompts and responses. We average R(x,y)∇θlog⁡πθ(y∣x) within each batch and use this estimate only to construct the projection. The verifier direction gR and hack gradient ∇pH remain exact, and the correction gain is λ=6⁠. Without advancing the policy, we measure |p˙−∥gR∥2|=|gR⊤u^|⁠. For each batch size and feature seed, we average this error over 1,000 independent batches, then report the mean and sample standard deviation across ten feature seeds.

Figure 5 (d) shows that larger batches reduce the mean error in preserving acceptance growth. The exact projection gives zero error. Estimating the projection introduces the discrepancy p˙−∥gR∥2=(gR−g^R)⊤u^⁠. Thus, more samples reduce the error in the acceptance growth rate. Panel (a) checks whether the correction also reduces hack probability.

Takeaway.

Audits reveal which accepted responses are wrong, but suppressing those responses can also suppress correct ones because they share policy parameters. Projection preserves the instantaneous growth rate of acceptance. When the corrected update reduces hack probability, correctness must therefore increase. The ISS bound describes how sustained correction limits the hacked share despite pressure toward hacks from verifier training. In these experiments, broader audit coverage improves hack reduction, while larger sample batches reduce projection error. Selective control reduces hack probability without reducing correctness. Projection preserves the instantaneous growth rate of acceptance, so any decrease in hack probability under the corrected flow must increase correctness.

Hyperparameters.

Table 3 lists only added or changed settings for this experiment. All other applicable settings follow Table 1.

Table 3. Additional and changed settings for selective control with correctness under c1⁠. Other model settings follow Table 1.

Parameter Value
Accepted-group feature means (4,−1,0,0)⁠, (4,1,0,0)
Rejected-group second-coordinate mean −0.5
Initial hacked share q(0) 1/2
Initial acceptance p(0) 2/3
Initial parameters; warmup θ(0)=0⁠; none
Independent feature seeds 10 (seeds 0,…,9⁠)
Audit correction gain λ 6
Gradient-regularization weight γ 16
Audit labels Exact
Initial audited fraction pHA(0)/pH(0) {0,1/8,1/4,1/2,3/4,1}
Projection batch sizes n {32,128,512,2048}
Policy for projection-error ablation θ=0
Flow integrator DOP853
Horizon, control comparison T=10
Horizon, coverage and ISS T=1
Tolerance for selectivity ϵ 10−8

D.2. The Neural Contextual Bandit Experiment

Motivation.

We replace the log-linear policy in Appendix D.1 with a neural policy to examine reward hacking (Section 3) and the limits of verifier feedback (Section 4) when the policy learns its representation. We also examine how audit coverage and projection accuracy affect projected audit correction (Section 5).

Experimental setting.

We use a shared MLP fθ with four inputs, one hidden layer of 16 tanh units, and a scalar output. We train all weights and hidden biases and omit the output bias. The policy is πθ(y∣x)∝exp(a(x,y)+fθ(ϕ(x,y))−fθ(0)(ϕ(x,y))).(122) We freeze the subtracted network, so the offsets a(x,y) (see Appendix D.1) determine the initial group probabilities.

Data.

We use eight equally weighted prompts, 16 responses per group, and four-dimensional Gaussian features with noise standard deviation 0.25⁠. We follow the centering and reflection construction in Appendix D.1, fixing the rejected group mean to (0,−0.5,0,0)⊤⁠. The shared first-coordinate mean of the accepted groups is 1 in Experiments I–II and 4 in the ablations. Within each experiment, comparisons share the Gaussian draws and network initialization for each seed. The network receives response features without group labels or a separate prompt embedding.

Under the correctness assignment c1⁠, the groups Gx,Hx,Nx have labels (R,c1)=(1,1),(1,0),(0,0)⁠, respectively. Experiment II also evaluates c=R and c2=R−c1 while keeping these groups and the training trajectory fixed.

RLVR training.

Experiments I–II follow verifier flow (1), θ˙=gR⁠, where gR=∇θJR and JR=p⁠. We compute probabilities and gradients by summing over all prompts and responses. We integrate the flow using DOP853, recomputing the gradients at each policy. The coverage ablation uses the same integrator for the corrected flow. Table 4 lists other numerical settings.

Audits.

The coverage ablation uses fixed, nested audit sets Aaud under c1⁠. Each audit reveals the correctness of an accepted response. With HA=Aaud∩H⁠, we apply θ˙=gR−λP⊥∇θpHA.(123) Both gradients are exact. The projector P⊥ acts in the space of neural parameters, orthogonally to gR⁠, and equals I when gR=0⁠.

The projection ablation evaluates corrections at the initial policy without advancing it. We estimate gR from independent draws x∼D and y∼πθ(⋅∣x) by averaging R(x,y)∇θlog⁡πθ(y∣x)⁠. Only the projector uses this estimate. The verifier direction gR and the full hack gradient ∇θpH remain exact.

Metrics.

We report acceptance p⁠, correctness pG⁠, hack probability pH⁠, and hacked share q=pH/p⁠. We average group probabilities over prompts before forming this ratio. Experiment I compares JC=pG with JR=p along training. In Experiment II, the subscript c identifies the correctness assignment defining pG,c and pH,c⁠. The coverage ablation reports pH(T) against initial coverage pHA(0)/pH(0)⁠. The projection ablation reports the deviation |p˙−∥gR∥2| from the acceptance growth rate preserved by the exact projector.

We report means and one sample standard deviation across ten independent seeds. For projection error, we first average absolute deviations over independent batches within each seed.

Results.

Experiment I: reward hacking.

We study whether increasing acceptance can amplify hacks and reduce correctness with a neural policy. Under verifier flow, we vary the hack share q(0)∈{0.1,0.3,0.5} while fixing initial acceptance p(0)=2/3 and the response features. For q(0)=0.3⁠, we track acceptance p⁠, hacked share q⁠, and correctness pG over time. We also compare the proxy objective JR=p with the true objective JC=pG to examine reward hacking as defined by (Skalse et al. 2022). RLVR training supplies the policies required by this definition: two times t1<t2 exhibit reward hacking when JR(θ(t2))>JR(θ(t1)) but JC(θ(t2))<JC(θ(t1))⁠. Thus, the comparison reveals whether improving verifier acceptance comes at the expense of correctness.

Figure 6 (a) shows that acceptance p and hacked share q increase together from q(0)=0.3⁠. As training progresses, correctness pG=p(1−q) eventually decreases despite increasing acceptance. The shift toward hacks then outweighs the gain from accepting more responses (Section 3.1.)

Neural contextual bandit experiments. (a) Acceptance p and hacked share q increase together from q(0)=0.3, while correctness p_G eventually decreases. (b) Correctness J_C=p_G versus acceptance J_R=p for q(0)\in\{0.1,0.3,0.5\}. Increasing acceptance accompanied by decreasing correctness exhibits reward hacking. (c) The same training record yields different correctness trajectories under c=R, c_1, and c_2=R-c_1. Square markers show correctness under c=R, which equals acceptance. (d) Final hack probability p_H(T) versus the initial audited fraction p_{H_A}(0)/p_H(0) for fixed audit sets. Greater coverage reduces the probability of hacks. (e) Deviation |\dot p-\|g_R\|^2| versus the number of samples used to estimate the gradient defining the projector. The verifier direction and audit gradient remain exact. Larger batches reduce the deviation. The exact projector preserves \dot p=\|g_R\|^2. Curves show means across ten seeds; shading shows one standard deviation. In (b), vertical bands measure variability in correctness at matched training times.
Figure 6. Neural contextual bandit experiments. (a) Acceptance p and hacked share q increase together from q(0)=0.3⁠, while correctness pG eventually decreases. (b) Correctness JC=pG versus acceptance JR=p for q(0)∈{0.1,0.3,0.5}⁠. Increasing acceptance accompanied by decreasing correctness exhibits reward hacking. (c) The same training record yields different correctness trajectories under c=R⁠, c1⁠, and c2=R−c1⁠. Square markers show correctness under c=R⁠, which equals acceptance. (d) Final hack probability pH(T) versus the initial audited fraction pHA(0)/pH(0) for fixed audit sets. Greater coverage reduces the probability of hacks. (e) Deviation |p˙−∥gR∥2| versus the number of samples used to estimate the gradient defining the projector. The verifier direction and audit gradient remain exact. Larger batches reduce the deviation. The exact projector preserves p˙=∥gR∥2⁠. Curves show means across ten seeds; shading shows one standard deviation. In (b), vertical bands measure variability in correctness at matched training times.

Figure 6 (b) shows how the proxy objective JR=p and the true objective JC=pG change along training. The mean curves show reward hacking throughout the recorded interval for q(0)=0.5⁠: acceptance increases while correctness decreases. For q(0)=0.3⁠, acceptance and correctness initially increase together. Reward hacking emerges later, when correctness decreases despite further gains in acceptance. For q(0)=0.1⁠, both quantities increase throughout the recorded interval, with no reward hacking evident in the mean curve. Hacks are present at initialization in all three cases, but reward hacking occurs when improving acceptance reduces correctness.

Experiment II: compatible correctness assignments.

We examine whether the same neural training record can support different conclusions about correctness. Starting from q(0)=0.3⁠, we evaluate each verifier trajectory under the correctness assignments: c=R⁠, c1=1G⁠, and c2=R−c1⁠. The features, offsets, initialization, and verifier are identical across these assignments. Their correctness probabilities are p,pG,pH⁠, respectively, and their hack probabilities are 0,pH,pG⁠. Thus, comparing c=R with c1 tests whether the record reveals the presence of hacks, while comparing c1 with c2 tests whether it identifies which accepted responses are wrong (Section 4).

Figure 6 (c) shows identical acceptance p under three compatible correctness assignments. Correctness increases under c=R and c2=R−c1⁠, whereas it eventually decreases under c1⁠. All three rules share the same training record Lt⁠, so verifier observations cannot determine which correctness trajectory applies. This illustrates the limits of detection and identification in Propositions 4.1 and 4.2. Because c1 and c2 exchange correct responses and hacks, the comparison also illustrates why these observations cannot guarantee selective control under every compatible assignment (Proposition 4.3).

Ablation I: audit coverage.

We examine how much of the hack population a fixed audit set must expose for the correction to reduce total hack probability. We audit fractions {0,1/8,1/4,1/2,3/4,1} of the candidates in each accepted group and prompt. The sets are nested, with candidate indices fixed across seeds and throughout training. The evaluator’s groups serve only to control coverage. Each condition uses the exact gradient ∇pHA and the same correction gain, without inverse probability weighting. We compare pH(T) across initial coverage levels and track how the audited and full hack gradients align. Zero coverage recovers verifier flow, while full coverage recovers the ideal projected correction. Partial coverage can also oppose hacking when the projected gradients align positively, as described in Appendix A.2.

Figure 6 (d) shows that the final hack probability pH(T) remains large when the initial audited fraction pHA(0)/pH(0) is small. Increasing this fraction reduces pH(T) to near zero in the tested setting. The correction uses only the audited hacks HA=Aaud∩H⁠, while the plot measures the full hack population H⁠. Thus, the comparison shows how expanding audit coverage improves suppression beyond the audited subset (see Appendix A.2).

Ablation II: projection error.

We test how estimating the reward gradient affects the projection’s preservation of acceptance growth. At the initial policy, we draw independent prompts x∼D and responses y∼πθ(⋅∣x)⁠. We estimate gR by averaging R(x,y)∇θlog⁡πθ(y∣x) and use this estimate to construct the projection. The verifier direction gR and hack gradient ∇pH remain exact, isolating the effect of estimating the projector. Without advancing the policy, we measure |p˙−∥gR∥2|=|gR⊤u^|⁠. For each batch size, we average absolute discrepancies over independent batches within each seed, then report their mean and standard deviation across seeds.

Figure 6 (e) shows that larger sample batches reduce the deviation |p˙−∥gR∥2| caused by estimating the projection. Larger batches improve preservation of the acceptance growth rate, as predicted by p˙−∥gR∥2=(gR−g^R)⊤u^⁠. Selective control additionally requires the correction to overcome the growth of hacks, as demonstrated in Figure 1 (d).

Takeaway.

Rising verifier reward can conceal declining performance on the intended task, even for a neural policy that learns its own representation. The decline stays hidden because the information available during training cannot separate correct responses from accepted errors. Audits supply this missing information. An intervention based on audits works only as well as their coverage, corrective strength, and projection accuracy allow.

Hyperparameters.

Table 4 lists the settings. Data construction follows Appendix D.1. All studies use ten seeds, each determining the Gaussian features and an independent network initialization. The flows start without warmup and use DOP853.

Table 4. Settings for the neural bandit experiments and ablations.

Parameter Value
Network widths; activation 4–16–1⁠; tanh
Input / output weight distributions N(0,1/4) / N(0,1/16)
Initial hidden biases; output bias Zero; omitted
Independent seeds 10 (0,…,9⁠)
Prompts; responses per group 8⁠; 16
Gaussian feature standard deviation 0.25
Initial acceptance p(0) 2/3
Initial hacked share, Experiment I {0.1,0.3,0.5}
Initial hacked share, time plot and Experiment II 0.3
Initial hacked share, both ablations 1/2
Accepted first-coordinate mean, I–II / ablations 1 / 4
Rejected feature mean μN (0,−0.5,0,0)⊤
Correction gain λ 6
Fractions of candidates audited {0,1/8,1/4,1/2,3/4,1}
Seed for audit sets 2026
Audit labels Exact
Batch sizes for estimating gR {32,128,512,2048}
Independent batches per size and seed 1,000
Sampling seed stream (2027,feature seed)
Integrator DOP853
Horizon, Experiments I–II / coverage 50 / 1
Recording interval, Experiments I–II / coverage 0.1 / 0.01

D.3. The Language Model Experiment

Motivation.

To test whether reward hacking and audit correction behave as our theory predicts beyond exact gradient flow, we train an autoregressive language model with sampled gradients and finite updates. Because the task below provides correctness labels, we can separate gains in verifier reward from gains in correctness.

Setup.

Task.

Each prompt x specifies a rule table r:{1,2}→{1,2}2 and six input digits u1,…,u6⁠. The model replaces each digit with its image under r and concatenates the results: t(x)=r(u1)⋯r(u6)∈{1,2}12.(124) The prompt asks for each replacement and then for the final answer t(x)⁠. Each replacement depends only on its own input digit, so no state carries across positions.

Correctness and verifier labels.

An output is valid if it contains exactly one Final answer: field, on its last nonempty line, with 12 digits from {1,2}⁠. For a valid response y with parsed answer t^⁠, the correctness and verifier labels are c(x,y)=1{t^=t(x)},R(x,y)=1{t^11:12=t(x)11:12}.(125) Both labels are zero for invalid outputs. Correctness checks every replacement in the final answer, while the imperfect verifier checks only the last pair, and neither checks the intermediate replacements. Because t^=t(x) implies that the last pairs match, c≤R⁠: the verifier has no false negatives for this definition of correctness. The labels therefore partition responses into three groups: Gx={y:c=1},Hx={y:R=1,c=0},Nx={y:R=0}.(126)

Hint and demonstrations.

Every prompt contains a hint with the correct final pair and an incorrect prefix. We draw the prefix uniformly from the 210−1 incorrect prefixes and keep it fixed for that prompt. Because the final pair is correct, copying the hint in the required format produces a hack. Supervised demonstrations cover all three groups: they give the correct solution (G⁠), copy the hint (H⁠), or change the last digit of the correct answer (N⁠). Rejected demonstrations keep valid formatting, so their rejection comes from the wrong final pair rather than from a formatting failure.

Policy.

We use Qwen/Qwen2-0.5B with LoRA. Training updates only the adapter parameters θ and keeps the base model fixed. All methods train this single policy without KL regularization.

Initialization.

We first train on correct solutions with supervised fine-tuning (SFT) and then add demonstrations that copy the hint. From this common checkpoint, we run 80 further SFT updates with demonstrations from G⁠, H⁠, and N⁠. We sample these groups with probabilities (0.4,0.4,0.2) or (0.3,0.3,0.4) and name the two mixtures N20 and N40 after their share of rejected demonstrations. These probabilities describe the demonstration data, not the group probabilities of the resulting policy. For each mixture, all methods start from the same SFT checkpoint, and we run reinforcement learning with five seeds.

Data split.

The 16 rule tables and 64 inputs give 1,024 tasks. We assign 608 tasks to training, 208 to calibration, and 208 to testing, so that no pair of rule table and input appears in more than one partition. Rule tables recur across partitions, so testing evaluates unseen combinations of seen rules and inputs.

Training objective.

The verifier objective is JR(θ)=Ex∼Dtrain,y∼πθ(⋅∣x)[R(x,y)]⁠, where Dtrain is uniform over training tasks. Because R checks only the last pair, correct responses and hacks receive the same reward. Correctness labels enter training only through the audits in PAC.

GRPO updates.

Each round samples K=2 training prompts and B=8 responses per prompt (group). For response yji to prompt xj⁠, the score sji=∇θlog⁡πθ(yji∣xj) sums the token scores, including the end token when emitted. With rewards Rji=R(xj,yji)⁠, each group normalizes its advantages by the standard deviation of its binary rewards: R¯j=1B∑iRji,aji=Rji−R¯jR¯j(1−R¯j)+10−4.(127) Treating the advantages as constants, we estimate the gradient by v^GRPO=1KBM∑j=1K∑i=1Bajisji,M=192,(128) where the normalizer M is a constant, independent of response length. Each fresh batch supplies at most one AdamW step. A group whose rewards are all equal has zero advantages, and a batch in which every advantage is zero skips its step. Skipped steps still count toward the 20 rounds of every run.

Metrics.

Acceptance and correctness.

We report correctness, acceptance, and the hacked share: JC=pG,JR=p=pG+pH,q=pH/p,(129) where p=pG+pH because the verifier has no false negatives. For n sampled responses, we estimate pG=nG/n⁠, pH=nH/n⁠, and p=(nG+nH)/n⁠. We estimate q=nH/(nG+nH) from counts pooled across prompts and leave it undefined when no response is accepted. A rising q indicates growth of hacking, and a rising JR with a falling JC is reward hacking in the sense of Proposition 3.1.

Evaluation and variation across seeds.

At rounds 0, 5, 10, and 20, we evaluate the policy on eight fixed calibration prompts with four responses each. Final evaluations use 32 fixed test prompts, also with four responses each, so calibration trajectories and final test results come from different prompts. We sample at temperature 1 from the full token distribution and record separately the responses with format errors and those that reach the generation limit. Figure 7 shows means over five training seeds. Table 5 reports means ± one sample standard deviation across these five seeds.

Initialization and audit ablations in the language model. (a,b) With SFT demonstration proportions G/H/N=0.4/0.4/0.2, GRPO increases verifier reward while reducing correctness. Raw audit and PAC instead increase correctness and reduce hacks. (c) With proportions 0.3/0.3/0.4, both corrections again favor correct responses, with PAC reaching higher correctness earlier. (d,e) Correctness over training rounds for raw audit and PAC at three audit probabilities \rho, using the initialization in (c). PAC has higher mean correctness at round 5 at each probability. (f) Final test correctness against audit count. Auditing fewer responses retains high correctness with fewer labels, but PAC has no final advantage over raw audit in these runs. Curves show means over five seeds. Panels (a)–(e) use calibration evaluations at rounds 0, 5, 10, and 20; arrows in (a)–(c) indicate training direction. Panel (f) uses separate test prompts, with points at \rho \in \{0.25,0.5,1\} from left to right for each method.
Figure 7. Initialization and audit ablations in the language model. (a,b) With SFT demonstration proportions G/H/N=0.4/0.4/0.2⁠, GRPO increases verifier reward while reducing correctness. Raw audit and PAC instead increase correctness and reduce hacks. (c) With proportions 0.3/0.3/0.4⁠, both corrections again favor correct responses, with PAC reaching higher correctness earlier. (d,e) Correctness over training rounds for raw audit and PAC at three audit probabilities ρ⁠, using the initialization in (c). PAC has higher mean correctness at round 5 at each probability. (f) Final test correctness against audit count. Auditing fewer responses retains high correctness with fewer labels, but PAC has no final advantage over raw audit in these runs. Curves show means over five seeds. Panels (a)–(e) use calibration evaluations at rounds 0, 5, 10, and 20; arrows in (a)–(c) indicate training direction. Panel (f) uses separate test prompts, with points at ρ∈{0.25,0.5,1} from left to right for each method.

D.3.1. When Hacking Grows

Controlled comparisons.

We train GRPO from both SFT checkpoints with the same settings, so each run samples 320 responses: 20 rounds of K=2 prompts with B=8 responses each. We evaluate the final checkpoint of each run.

Results.

Experiment I: reward hacking under GRPO.

In N40, from the SFT checkpoint to the final round, mean test acceptance rises from 69.5% to 99.4%⁠, while correctness falls from 30.5% to 2.2% and the hacked share rises from 56.2% to 97.8%⁠. All five seeds show the same three changes. Rising acceptance with falling correctness is reward hacking in the sense of Proposition 3.1, even though sampled updates with AdamW do not satisfy its assumption of exact gradient flow. Table 5 summarizes the final test outcomes.

Ablation I: SFT mixture.

Figure 7 (a,b) shows the calibration trajectories for N20. The N20 mixture produces the same pattern as N40: acceptance rises from 80.5% to 98.8%⁠, correctness falls from 31.2% to 0.5%⁠, and the hacked share rises from 61.2% to 99.5%⁠. Under GRPO, both SFT checkpoints therefore lead to growth of hacking and to reward hacking. Because changing the SFT mixture also changes the learned gradient geometry, this comparison does not isolate the effect of the initial composition.

Takeaway.

For both SFT mixtures, GRPO increases verifier reward while reducing correctness, so reward hacking persists beyond exact gradient flow.

D.3.2. Selective Control with Correctness Feedback

Additional experimental settings.

Audits and gradient estimates.

We audit each accepted response independently with probability ρ∈{1,0.5,0.25}⁠. The indicator Zji∼Bernoulli(ρ) marks whether response i to prompt j is audited, and reweighting by 1/ρ gives an unbiased estimate of whether the response is a hack: H~ji=ZjiRji(1−cji)ρ.(130) This estimate is zero for unaudited and rejected responses, so it requires correctness labels only for audited accepted responses. Let sji=∇θlog⁡πθ(yji∣xj) be the score of response i to prompt j⁠. With baselines R¯j,−i and H~―j,−i that average the other B−1 responses to prompt j⁠, we use the same KB=16 responses as GRPO to estimate g^=1KB∑j,i(Rji−R¯j,−i)sji,h^=1KB∑j,i(H~ji−H~―j,−i)sji.(131) At a fixed policy, g^ and h^ are unbiased estimates of g=∇p and h=∇pH⁠, so g^−h^ estimates ∇pG⁠. The estimate g^ uses the rewards of all responses, while h^ uses only the correctness labels revealed by audits. The exact labels used for reporting never enter training.

Corrections.

Both corrections subtract a direction v from the GRPO step. Raw audit correction uses vraw=h^⁠, the estimated gradient of pH⁠. PAC first removes the component of h^ along the estimated acceptance gradient, so that the correction leaves acceptance unchanged to first order: vPAC=h^−g^⊤h^∥g^∥2g^.(132) The projection uses the Euclidean inner product on adapter parameters, and when g^=0⁠, PAC reduces to raw audit correction. We compute both estimates before the AdamW step.

We normalize both correction directions and scale them using the same rule based on the optimizer step. With d the actual AdamW displacement, m the number of adapter parameters, and η=10−5⁠, each method applies b=max{∥d∥,0.25ηm},Δθ=d−bv∥v∥.(133) For a nonzero direction, the effective gain λ=b/∥v∥ varies across updates. We skip the correction when ∥v∥ is numerically zero. Because b≥0.25ηm⁠, a nonzero correction can still act when every GRPO advantage is zero and AdamW skips its step. Both methods use the same number of sampled responses, but their audit counts can differ because their acceptance differs.

Results.

Experiment I: selective correction.

Table 5 reports test correctness pG and the probability pH of hacks at the SFT checkpoint and after PAC with full auditing, averaged over five seeds for each SFT mixture. In all ten runs, PAC increases correctness and decreases hacks. In N40, for example, mean correctness rises from 30.5% to 97.2%⁠, and mean pH falls from 39.1% to 1.7%⁠. These are the directions that Theorem 5.1 predicts, even though PAC uses sampled gradients and finite updates.

Ablation I: removing projection.

Figure 7(a)–(c) compares the calibration trajectories of both corrections and GRPO for N20 and N40. In N40, raw audit correction also reaches high final correctness: 98.4% on test prompts, compared with 97.2% for PAC (Figure 7(f), ρ=1⁠). PAC’s advantage appears earlier in training. In N40, mean calibration correctness at rounds 5 and 10 is 85.0% and 98.1% for PAC, compared with 65.6% and 90.0% for raw audit correction. Under raw audit correction, acceptance falls below its initial value at round 5 in four of five seeds and later recovers, while correctness shows no such decline at the recorded rounds. PAC therefore reaches high correctness sooner with the same number of sampled responses. It does not end with higher correctness, and because the two methods audit different numbers of responses, this comparison does not hold audit cost fixed.

Ablation II: audit coverage.

Figure 7 (d,e) shows the calibration trajectories under partial auditing. Panel (f) relates final test correctness to audit count. For N40, we reduce ρ to 0.5 and 0.25 and compare with the runs at full auditing ρ=1⁠. At ρ=0.25⁠, PAC uses 72.0 audits per run on average, compared with 292.6 at full auditing, a 75.4% reduction. With this coverage, PAC still reaches 95.2% test correctness, and hacks make up 4.4% of responses. At all three audit probabilities, both raw audit correction and PAC increase correctness and reduce hacks relative to the SFT checkpoint in every seed. Final correctness is not monotone in ρ⁠. This ablation reduces the number of correctness labels, while every run still samples 320 responses. Because each accepted response is audited independently, every accepted response can receive an audit, and no subset is permanently excluded.

Takeaway.

Audit corrections avoid the reward hacking that GRPO exhibits: from the same SFT checkpoints, they increase correctness and reduce hacks. Projection raises correctness earlier in training, while raw audit correction reaches similar final correctness. Auditing one quarter of accepted responses retains about 97% of the gain in final correctness that full auditing achieves.

Table 5. Final test outcomes after 20 rounds. Values are percentages, reported as mean ± one sample standard deviation across five training seeds. Each SFT row is one shared initialization. SFT evaluations are shared across runs, so no standard deviation across training seeds is reported for these rows. N20 and N40 use G/H/N demonstration probabilities 40/40/20% and 30/30/40%⁠, respectively. Coverage ablations use N40.

Scenario Method Audit probability JC=pG JR=p pH
N20 SFT — 31.2 80.5 49.2
N20 GRPO — 0.5±0.7 98.8±0.4 98.3±0.3
N20 Raw audit 1 98.9±0.4 99.2±0.0 0.3±0.4
N20 PAC 1 98.1±1.3 99.4±0.3 1.2±1.6
N40 SFT — 30.5 69.5 39.1
N40 GRPO — 2.2±2.4 99.4±0.7 97.2±2.4
N40 Raw audit 1 98.4±2.6 99.7±0.7 1.2±2.0
N40 PAC 1 97.2±2.0 98.9±1.7 1.7±1.3
N40 Raw audit 0.5 97.2±2.0 99.1±1.0 1.9±2.3
N40 PAC 0.5 93.9±4.6 99.1±0.7 5.2±4.4
N40 Raw audit 0.25 96.1±1.1 99.5±1.0 3.4±1.8
N40 PAC 0.25 95.2±2.1 99.5±0.7 4.4±2.6

Hyperparameters.

Table 6 collects the settings shared across training runs. Generation and evaluation use the same temperature and token limit.

Table 6. Settings for the language model experiments.

Setting Value
Model Qwen2-0.5B
Input / output digits 6 / 12
Final SFT updates per scenario 80
Reinforcement learning seeds 0–4
Training rounds 20
Prompts per round / responses per prompt 2 / 8
Total training responses per run 320
Optimizer / learning rate AdamW / 10−5
Gradient norm cap / weight decay 1 / 0
KL coefficient 0
Advantage denominator offset 10−4
Temperature / maximum response tokens 1 / 192
Sampling Full token distribution
Calibration rounds 0, 5, 10, 20
Calibration prompts / responses each 8 / 4
Test prompts / responses each 32 / 4
Denison, Carson, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, et al. 2024. “Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.” arXiv Preprint arXiv:2406.10162.
Eisenstein, Jacob, Chirag Nagpal, Alekh Agarwal, Ahmad Beirami, Alexander Nicholas D’Amour, Krishnamurthy Dj Dvijotham, Adam Fisch, et al. 2024. “Helping or Herding? Reward Model Ensembles Mitigate but Do Not Eliminate Reward Hacking.” In First Conference on Language Modeling.

::: {#ref-DBLP:journals/corr/abs-2605-08007 .csl-entry} Elliott, Chris, Einar Urdshals, David Quarel, and Daniel Murfet. 2026. “Interpreting Reinforcement Learning Agents with Susceptibilities.” arXiv Preprint arXiv:2605.08007 abs/2605.08007. :::

Everitt, Tom, Victoria Krakovna, Laurent Orseau, and Shane Legg. 2017. “Reinforcement Learning with a Corrupted Reward Channel.” In Proceedings of the 26th International Joint Conference on Artificial Intelligence, 4705–13.
Gao, Leo, John Schulman, and Jacob Hilton. 2023. “Scaling Laws for Reward Model Overoptimization.” In International Conference on Machine Learning, 10835–66.
Gauthier, Etienne, Francis R. Bach, and Michael I. Jordan. 2026. “Explaining and Preventing Alignment Collapse in Iterative RLHF.” arXiv Preprint arXiv:2605.04266.
Helff, Lukas, Quentin Delfosse, David Steinmann, Ruben Härle, Hikaru Shindo, Patrick Schramowski, Wolfgang Stammer, Kristian Kersting, and Felix Friedrich. 2026. “LLMs Gaming Verifiers: RLVR Can Lead to Reward Hacking.” arXiv Preprint arXiv:2604.15149.
Karwowski, Jacek, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer, Charlie Griffin, and Joar Max Viktor Skalse. 2024. “Goodhart’s Law in Reinforcement Learning.” In The Twelfth International Conference on Learning Representations.
Kenton, Zachary, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill Tyshchuk, Jonah Brown-Cohen, Harri Edwards, et al. 2026. “Debate Training Reduces Reward Hacking in RLAIF.” arXiv Preprint arXiv:2608.17776.
Khalaf, Hadi, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju, and Flavio Calmon. 2025. “Inference-Time Reward Hacking in Large Language Models.” In The Thirty-Ninth Annual Conference on Neural Information Processing Systems.

::: {#ref-DBLP:journals/corr/abs-2603-07084 .csl-entry} Khalifa, Muhammad, Zohaib Khan, Omer Tafveez, Hao Peng, and Lu Wang. 2026. “Countdown-Code: A Testbed for Studying the Emergence and Generalization of Reward Hacking in RLVR.” arXiv Preprint arXiv:2603.07084 abs/2603.07084. :::

Khalil, Hassan K, and Jessy W Grizzle. 2002. Nonlinear Systems. Vol. 3. Prentice hall Upper Saddle River, NJ.
Kwa, Thomas, Drake Thomas, and Adrià Garriga-Alonso. 2024. “Catastrophic Goodhart: Regularizing RLHF with KL Divergence Does Not Mitigate Heavy-Tailed Reward Misspecification.” In The Thirty-Eighth Annual Conference on Neural Information Processing Systems.
Laidlaw, Cassidy, Shivam Singhal, and Anca Dragan. 2025. “Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking.” In The Thirteenth International Conference on Learning Representations.
Lambert, Nathan, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James Validad Miranda, et al. 2025. “Tulu 3: Pushing Frontiers in Open Language Model Post-Training.” In Second Conference on Language Modeling.
Lightman, Hunter, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2024. “Let’s Verify Step by Step.” In The Twelfth International Conference on Learning Representations.
MacDiarmid, Monte, Benjamin Wright, Jonathan Uesato, Joe Benton, Jonathan Kutasov, Sara Price, Naia Bouscal, et al. 2025. “Natural Emergent Misalignment from Reward Hacking in Production RL.” arXiv Preprint arXiv:2511.18397.
Mahmoud, Anas, MohammadHossein Rezaei, Zihao Wang, Anisha Gunjal, Bing Liu, and Yunzhong He. 2026. “Reward Hacking in Rubric-Based Reinforcement Learning.” In Second Workshop on Agents in the Wild: Safety, Security, and Beyond.
Moya, Christian, Alex Semendinger, Guang Lin, and Elliott Thornley. 2026. “Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training.” In Forty-Third International Conference on Machine Learning.
Moya, Christian, and Jiankang Wang. 2018. “Developing Correlation Indices to Identify Coordinated Cyber-Attacks on Power Grids.” IET Cyber-Physical Systems: Theory & Applications 3 (4): 178–86.
Pan, Alexander, Kush Bhatia, and Jacob Steinhardt. 2022. “The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.” In International Conference on Learning Representations.
Pan, Jane, He He, Samuel R. Bowman, and Shi Feng. 2024. “Spontaneous Reward Hacking in Iterative Self-Refinement.” arXiv Preprint arXiv:2407.04549.
Pasqualetti, Fabio, Florian Dörfler, and Francesco Bullo. 2013. “Attack Detection and Identification in Cyber-Physical Systems.” IEEE Transactions on Automatic Control 58 (11): 2715–29.
Skalse, Joar, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger. 2022. “Defining and Characterizing Reward Gaming.” In Advances in Neural Information Processing Systems, 35:9460–71.

::: {#ref-DBLP:journals/corr/abs-2601-13548 .csl-entry} Wang, George, and Daniel Murfet. 2026. “Patterning: The Dual of Interpretability.” arXiv Preprint arXiv:2601.13548 abs/2601.13548. :::

Wang, Songtao, Quang Hieu Pham, Fangcong Yin, Xinpeng Wang, Jocelyn Qiaochu Chen, Greg Durrett, and Xi Ye. 2026. “Detecting and Suppressing Reward Hacking with Gradient Fingerprints.” In Third Conference on Language Modeling.
Wang, Xinpeng, Nitish Joshi, Barbara Plank, Rico Angell, and He He. 2026. “Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort.” In The Fourteenth International Conference on Learning Representations.
Wang, Xuekang, Zhuoyuan Hao, Shuo Hou, Hao Peng, Juanzi Li, and Xiaozhi Wang. 2026. “Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning.” arXiv Preprint arXiv:2606.04923.
Zhong, Ziqian, Aditi Raghunathan, and Nicholas Carlini. 2025. “ImpossibleBench: Measuring LLMs’ Propensity of Exploiting Test Cases.” arXiv Preprint arXiv:2510.20270.
Zhuang, Simon, and Dylan Hadfield-Menell. 2020. “Consequences of Misaligned AI.” In Advances in Neural Information Processing Systems, 33:15763–73.

:::::::::::::::::::::::::::::::::::::::