In reinforcement learning with verifiable
rewards (RLVR), imperfect verifiers can reward incorrect responses,
creating opportunities for reward hacking. Using gradient flow with a
fixed verifier, we characterize the conditions under which reward rises
while correctness falls. We then show that the observations available
during RLVR are, in general, insufficient to detect or identify accepted
errors, or to guarantee their reduction without sacrificing correct
responses. To address this limit, we construct a correction using
additional feedback about correctness from audits. This correction
achieves selective control: at the current policy, it lowers
the probability of accepted errors and raises that of correct responses,
provided it outweighs the pressure toward errors from verifier reward.
Experiments with log linear and neural contextual bandits and with a
language model support the analysis and show that selective control
under partial auditing reduces accepted errors while increasing
correctness.
1. Introduction
Reinforcement learning with verifiable rewards (RLVR) (Lambert et al.
2025) fine-tunes pretrained language models using rewards from automated
checks. These checks come from a verifier, which compares a model’s
response with a known solution or runs the code in the response against
tests. Rewarding responses that pass the verifier has improved reasoning
performance on mathematics and coding tasks (DeepSeek-AI et al. 2025).
These gains, however, depend on the verifier: higher reward reflects
progress on the intended task only when the verifier reliably judges
correctness.
A verifier can accept a response even when the response fails the
intended task. For example, the code in a response may pass the
available tests yet fail on inputs those tests omit. Rewarding such
responses reinforces accepted errors: responses that satisfy the
verifier but fail the intended task. As the model produces accepted
errors more often, average reward can rise even while performance on the
intended task declines, the signature of reward hacking (Skalse et al.
2022).
Empirical studies document reward hacking in RLVR. Reward hacking
appears, for example, in inductive reasoning, where a model must infer a
general rule from labeled examples. Models trained with RLVR often skip
this rule and instead list the label of each example. The verifier
accepts these responses because it checks only whether a response labels
the given examples correctly (Helff et al. 2026). Beyond eliciting such
new hacking behaviors, reinforcement learning can amplify those that
models acquired during earlier fine-tuning (Khalifa et al. 2026). These
findings motivate methods that reduce reward hacking while preserving
progress on the intended task.
Existing methods address reward hacking by constraining
optimization (Laidlaw, Singhal, and Dragan 2025), improving
feedback (Coste et al. 2024; Lightman et al. 2024), or detecting and
correcting hacking (Baker et al. 2025; S. Wang et al. 2026). Yet because
these methods rely on assumptions about correctness (Everitt et al.
2017), they can leave verifier errors unresolved (Eisenstein et al.
2024). Training therefore continues with these verifier errors,
rewarding correct responses and accepted errors alike. This raises a
broader question: how do these verifier errors shape what RLVR learns
about the intended task as training proceeds?
To answer this question, we hold an imperfect verifier fixed and
study how the policy learns from its rewards. We model this learning as
gradient flow because it makes the analysis tractable. Within this
setting, we study when reward hacking grows, what the information
available during RLVR training reveals about it, and when an
intervention reduces hacking while preserving correctness. Our
contributions are:
(i) We derive a condition under which the share of accepted
errors among accepted responses grows at the current policy, and we
express it through two mechanisms: hack bias and correctness-to-hack
leakage (Proposition 3.2).
We also characterize when the gradient flow increases verifier reward
while reducing correctness, the signature of reward hacking
(Proposition 3.1).
(ii) We establish limits on detection, identification, and
selective control from verifier feedback alone. We show that the
complete RLVR training record provides no advantage in detecting
accepted errors and can leave correctness unidentifiable
(Propositions 4.1 and 4.2). We further
prove that no controller using this record alone can guarantee fewer
accepted errors while preserving correct responses (Proposition 4.3).
(iii) We derive conditions under which a correction to RLVR
training reduces reward hacking while improving correctness, and we
design projected audit correction, which satisfies them using
correctness labels from audits of some responses (Theorem 5.1).
We test our theory in settings of increasing complexity. In
contextual bandits, reward hacking emerges and projected audit
correction reduces hacks while increasing correctness. In a language
model, verifier reward rises while correctness falls, and the same
correction reverses this trend.
2. Problem Formulation
2.1. Reinforcement Learning with
Verifier Rewards (RLVR)
Policy. Given a prompt drawn from a prompt distribution , the language policy samples a
response from the response set
. This policy has a
trainable parameter vector .
RLVR. We analyze RLVR (Lambert et al. 2025), where
the verifier assigns a reward . A reward of
indicates acceptance. RLVR training maximizes the scalar objective:
This
objective represents the probability of acceptance averaged over prompts
and sampled responses, yielding a value in .
2.2. Correctness and Accepted
Errors
We define a fixed correctness indicator for each pair , where denotes a correct response. To focus
our analysis on the consequences of accepting incorrect responses, we
assume the verifier produces no false negatives. This assumption implies
that for all
pairs, meaning the verifier accepts all correct responses. We call the
event of an incorrect response being accepted, i.e., , a
hack.
For any given prompt , we can
partition the set of all accepted responses, . This set consists of two disjoint subsets. The first is
the set of correct responses, . The second is the set of hacks,
. These sets
remain fixed throughout training. The policy, however, can learn to
sample from them with different frequencies.
Acceptance and hacking. For any given prompt , we define two key probabilities. The
first is the acceptance probability, . The second probability, defined only when is the hacked share,
Both probabilities depend on the policy
parameters and change
throughout RLVR training. Using the acceptance probability , we can now rewrite the RLVR
objective as the following expectation over prompts: . Because this objective depends only
on , it makes no
distinction between correct responses and hacks (see Appendix C.1.2).
2.3. Gradient Flow Dynamics
To analyze how the policy evolves, we use gradient ascent in
continuous time. This model isolates the effects of the verifier signal
by removing the noise inherent in stochastic optimization. When the
objective is differentiable,
the parameters evolve according to the gradient flow: (1)
where
is the exact gradient. A direct consequence is that the objective can
only improve, as (see Appendix C.1.3). This guarantee, however,
applies only to the overall acceptance probability. It reveals nothing
about the hacked share,
which may increase, decrease, or remain unchanged.
Research questions. We now define our two central
research questions within our gradient flow framework. First, under what
conditions does the hacked share, , grow or persist during
training? Answering this question requires distinguishing increases in
acceptance due to correct responses from shifts in the policy’s sampling
toward hacks. Second, since the verifier’s feedback is blind to
correctness, is it possible to reduce the hacked share using verifier’s
information alone? If not, what additional information would allow it to
do so?
3. When Hacking Grows: A Population
Analysis
This section analyzes the population dynamics of hacked responses. We
show that the accepted population and the hacked share can increase
together, and characterize when verifier flow increases reward while
reducing correctness. We then describe how hack bias and
correctness-to-hack leakage drive hacking growth. Ultimately, these
drivers create a latent vulnerability that on-policy training can
reinforce.
3.1. Verifier Acceptance and Hacked
Share Can Increase Together
We partition responses into three populations: correct responses
, hacks , and rejected
responses .
Let and
denote the probabilities of correct responses and hacks, respectively.
To track both how often the verifier accepts and how often accepted
responses are hacks, we define and
where measures
overall acceptance and measures the hacked
share among accepted pairs. The latter requires .
Expected reward depends on total acceptance , regardless of how it splits
between correct responses and hacks. Writing and
and differentiating with respect to time
gives (2) where
the terms reflect changes in
total acceptance and the
terms reflect shifts between correct responses and hacks. These
identities impose no trade-off between acceptance and hacked share; both
can increase together.
Reward hacking. Skalse et al. (2022) call a proxy
reward hackable if some pair of policies has higher expected proxy
return but lower expected true return. In our setting, the verifier
reward serves as the proxy,
and correctness serves as
the true return. We characterize when gradient flow on , which we call verifier flow,
generates such a pair.
Proposition 3.1 (Reward hacking along the flow).
Suppose and are continuously differentiable and
along verifier flow on . Then there exist with and
if
and only if at some .
Interpretation. The two sides of the inequality
compete. The term is the
correctness lost as the accepted population shifts toward hacked
responses, and is the
correctness gained from rising acceptance. Training produces reward
hacking once the loss outweighs the gain. The condition needs to hold at
only a single time to produce a pair of policies that witnesses
hackability. Because both terms depend on the policy, the condition can
hold at one time and fail at another.
3.2. The Dynamics of the Hacked
Share
While the hacked share remains
the quantity of interest, log-odds coordinates simplify its dynamics.
When both and
, we can define the
log odds as
(3) Because is strictly increasing in , the sign of determines whether the hacked
share grows, shrinks, or remains constant.
The log odds dynamics. Since , the rate depends on the gradient of each
group’s log-probability. For each group with positive, continuously differentiable
probability, we denote this gradient by the
group’s average score. Along the verifier flow (1),
differentiating yields (4) The log odds grow when the reward gradient aligns with the score difference
between hacks and correct responses. Differentiating total acceptance
shows that is a weighted combination of the
three group scores: This decomposition,
combined with (4),
resolves into two
mechanisms: (5) The
full derivation is provided in Appendix C.2.3.
3.3. Main Result: The Mechanism of
Hacking Growth
We now state our main result characterizing the mechanism of hacking
growth.
Proposition 3.2 (Population growth of the hacked
share). Along verifier gradient flow, with , the population hacked
share grows if and only if (6)
When the inequality holds, . Since and , both the hacked share and total acceptance grow strictly,
while the correct share among accepted responses decreases.
(i) Hack bias, , arises because hacks contribute to the reward gradient
in proportion to their share among accepted responses. It is always
nonnegative and grows with the hacked share at fixed group scores.
(ii) Correctness-to-hack leakage, , arises when the direction that favors correct responses
over rejected ones also favors hacks over correct responses. It can be
positive or negative: positive leakage reinforces hack bias, negative
leakage opposes it and can reverse growth when its magnitude exceeds the
bias.
Interpretation. This result is local: the group
scores and hacked share depend on the current parameters, so the growth
condition can change during training. The group scores evolve, so
continued growth is not guaranteed. Yet a feedback loop is present: a
growing hacked share strengthens hack bias at fixed group scores, which
favors further growth. This result exposes a latent vulnerability:
hacking can reinforce the conditions that favor its own growth, because
the policy generates its own training samples. Moreover, RLVR training
is not designed to oppose this reinforcement. We must therefore build
any hacking defense on top of RLVR. Can the information generated during
RLVR training support such a defense?
4. What RLVR Training Reveals:
Limits of Verifier Feedback
Section 3 showed that rising rewards can mask a
growing share of hacks. We now study whether the observations collected
during RLVR training can expose them. We show that they cannot: training
observations confer no detection advantage, leaving correct responses
unidentifiable, and making selective control impossible.
4.1. RLVR Training Observations,
Monitors, and Compatible Correctness
RLVR training observations. RLVR training produces
prompts, sampled responses, and verifier labels. Let collect these observations through
time , along with policy
parameters, probabilities, gradients, and other derived quantities. This
record is deliberately generous: the limitations below do not arise from
incomplete logging. We assume that contains no additional correctness
feedback beyond the verifier. We write for the
information contained in the record.
Monitors. A monitor is an algorithm that uses to assess hacking. A detection
monitor raises a binary alarm about whether hacks exist; an
identification monitor predicts the correctness label of an accepted
response. We consider deterministic monitors here and defer randomized
extensions to Appendix C.3.
Compatible correctness assignments. A fixed
assignment determines
correctness. To study what the training record reveals about , we consider alternative assignments
consistent with the verifier. Assuming no false negatives, we define
(7) This
class models uncertainty about the true correctness assignment; itself remains fixed during training.
All members agree on rejected responses, but they may disagree on
accepted ones. To isolate the effect of this uncertainty, we compare
these alternative assignments while holding the prompt distribution,
verifier, initialization, and training algorithm fixed.
4.2. Limits of Detecting
Hacking
Verifier acceptance can increase while the hacked share grows. Can
the training record reveal even the presence of hacks? Under , every accepted response is
correct. A compatible alternative with admits hacks. Detection requires
distinguishing these two cases using only .
Proposition 4.1 (Limits of detection). Under the
comparison setup above, fix ,, and with . The
training record has the same distribution under both assignments: Thus, no monitor can
distinguish , which admits
hacks, from , which has
none.
Interpretation. Because alternative correctness
assignments do not affect verifier feedback, they leave the training
record’s distribution unchanged. Thus, any monitor’s detection rate
under equals its false-alarm
rate under . Even knowing that
hacks exist does not reveal which accepted responses are wrong.
4.3. Limits of Identifying
Hacks
Suppose we know that hacks exist. The remaining task is to determine
which accepted responses are wrong. Identification requires a
monitor to recover the true label of an accepted response from . To succeed, the prediction must distinguish
compatible assignments that disagree on this response, even when we
restrict to
assignments that admit hacks.
Proposition 4.2 (Limits of identification).
Under the comparison setup above, fix and . Suppose assigns positive probability to
both correct responses and hacks under some assignment in . For every monitor and
every accepted pair , there is
an assignment
with
such that Thus, no monitor can
guarantee the correct label of an accepted response, even when we know
hacks exist.
Interpretation. Two compatible assignments can label
the same accepted response differently while producing identical
records. Both can admit hacks, so knowing that hacks exist does not
resolve the disagreement. A prediction correct under one assignment is
wrong under the other. Thus, the monitor incurs an error probability of
at least under at least
one assignment. Because reducing hacks need not require identifying
every hack, this result alone does not rule out selective control.
4.4. Limits on Selective Control
from Verifier Feedback Alone
Even if we cannot identify every hack, can we reduce hack probability
without reducing the probability of correct responses? Suppose a
controller uses to apply a
correction to
the verifier flow, yielding . Along this corrected flow, selective
control requires and
whenever . We target
itself, not the hacked share
, because can fall even while grows. For a uniform guarantee, this
same controller must satisfy these conditions under every compatible
assignment in .
Proposition 4.3 (Limits of selective control).
Under the comparison setup above, suppose the initial policy assigns
positive probability to both correct responses and hacks under some
assignment in . For
corrected flows with differentiable group probabilities, no controller
using only can guarantee and whenever , under every assignment . Here, each
assignment defines its own
and .
Interpretation. We seek a correction such that the corrected flow reduces
hack probability without reducing the probability of correct responses.
However, compatible assignments can exchange the roles of correct
responses and hacks. An update that reduces hack probability under one
therefore reduces the probability of correct responses under the other.
The controller cannot distinguish these assignments from , and more observations from the same
verifier cannot resolve this conflict. Regularization illustrates this
limit: it constrains policy updates using only the policy and the
verifier, so it cannot guarantee selective control either (Appendix C.3.5). A uniform guarantee
of selective control thus requires additional correctness information
that rules out assignments requiring incompatible updates.
5. From Additional Feedback to
Selective Control
Section 4 showed why verifier
feedback alone cannot reduce hacks without also sacrificing correct
responses. We now investigate whether additional information about
correctness can break this trade-off. We show that it can: such
information supports a correction that suppresses hacks and promotes
correct responses, provided the correction is strong enough to overcome
the drift toward errors induced by RLVR training.
5.1. Projected Audit Correction in
RLVR Training
Audits as additional information. Verifier feedback
alone cannot guarantee selective control (Section 4), so we introduce a
stronger signal: audits. An audit reveals the true correctness
label for a prompt and an accepted response . The audit rules out every
correctness rule that disagrees with this label, restricting to The limits in Section 4 arose because verifier
feedback could not distinguish the true rule from alternatives in . Audits eliminate
alternatives that contradict the revealed labels. To show how audits
reduce hacking, we next define the audited hack population and its
probability gradient under the policy.
Let be a fixed
set of audits, and let denote the audited hacks, with probability . When
is differentiable, the
negative gradient targets only the audited hacks and need not align with
(see Appendix A.2). Despite this
misalignment, locally decreases the audited hack probability when
, so we
use the correction , with , in the verifier flow .
The correction opposes growth in : Because these probabilities share parameters, the
correction also perturbs
and :(8) The cross terms
have no fixed sign. When , the correction can suppress
correct responses alongside hacks. We next constrain to preserve the growth rate of .
Projected audit correction. Since , we have . Preserving the instantaneous growth rate of requires , so must lie in the subspace orthogonal
to . Under this constraint,
: whenever
opposes the growth of (), it
contributes equally to the growth of .
We project the audit correction onto the subspace orthogonal to using when
and otherwise, to define the
projected audit correction, with
, yielding the flow
. Since and is an orthogonal projector, (9) This projection preserves
the instantaneous growth rate of
and opposes the growth of
whenever . We next show when this projected correction shrinks the
share of hacks in the full population and bounds that share during
training.
5.2. Main Result: Achieving
Selective Control
Assumptions. We analyze RLVR training under the
projected audit correction on a finite interval , under the following assumptions.
(A1) Regularity: The probabilities are continuously
differentiable in , with
on .(A2) Audit coverage: The
fixed audit set satisfies throughout the interval, so and their gradients coincide
during training. (A3) Uniform corrective strength:
throughout the interval for some constant . We discuss these conditions
in Appendix A.2.
Theorem 5.1 (Selective control). Consider RLVR
training under the projected audit correction with constant on .
**(i) Selective control.* Under (A1)–(A2), the correction opposes the
relative growth of hacks while preserving the instantaneous growth rate
of :(10) Whenever the correction term exceeds the bias
and leakage, the share of hacks decreases and increases. More strongly, if then itself
decreases (,, and ), achieving selective
control.*
**(ii) An ISS-type bound on the share of hacks.* Assume additionally
(A3), and let bound the
growth pressure, , during training. Then, for every ,(11)
The bound separates a
decaying contribution from the initial hacked share and a residual
contribution from growth pressure. For fixed , a larger product reduces the residual
bound.*
Proof sketch. By (A2), . Substitute the correction into and use to obtain part (i). For
part (ii), (A3) and
give ;
integration yields the bound. See Appendix C.4.2 for
details. ◻
Interpretation. Under the stated assumptions, the
projected audit correction favors correct responses over hacks while
preserving the instantaneous growth rate of . When bias and leakage supply no
positive growth pressure, the input-to-state stability (ISS)-type
bound (Khalil and Grizzle 2002) ensures the share of hacks decays
exponentially. Persistent pressure contributes a residual bound that
decreases with for
fixed . The selective effect is
instantaneous; the trajectory bound requires the assumptions to hold
throughout the interval.
Remark (Practical implementation). The control
guarantees (Theorem 5.1) require audits that
provide a direction opposing hack growth and sufficient corrective
strength throughout training. Partial audit coverage can suffice when
the resulting correction remains aligned with the gradient of hack
probability. Finite-sample gradient estimates, optimizers, and finite
step sizes can prevent the implemented update from preserving acceptance
progress or suppressing hacks. Appendix A.2 discusses these
requirements and how to estimate and validate the correction.
6. Experiments
We test reward hacking, the limits of verifier feedback, and
selective control through audit correction in contextual bandits and
language models. The code is available at https://github.com/cmoyacal/verifier-errors.
6.1. Gaussian Contextual
Bandits
Setup. We use Gaussian contextual bandits to study
when training on verifier rewards amplifies reward hacking and whether
audits enable selective control. We set the initial probabilities of
correct and hacked responses, and , to isolate how the initial
composition shapes training. We also vary the policy class: a log linear
policy keeps its features fixed, while a neural policy learns its own
representation. Appendices D.1
and Appendix D.2 provide the experimental
details.
Results.Figure 1 (a) shows acceptance
and hacked share increasing together from , as Proposition 3.2 predicts. Panel (b) plots
correctness against
verifier reward for three initial hacked shares. For , correctness declines throughout
the recorded interval; for , it first improves and then
declines; for , both
correctness and verifier reward improve. Both declines are reward
hacking in the sense of Proposition 3.1. Panel (c) evaluates
the same verifier flow under two correctness assignments: , which admits no hacked responses,
and , which does. The
information available during training is identical under both, yet
correctness differs, as Proposition 4.1 predicts.
Panel (d) shows selective control with a neural policy: projected audit
correction (PAC) decreases the probability of hacked responses while increasing
correctness , consistent with
Theorem 5.1. Raw audit correction
initially decreases both probabilities, whereas verifier flow and
gradient regularization increase .Appendices D.1 and Appendix D.2 provide additional
experiments and ablations.
Figure 1. Reward hacking and selective control in
contextual bandits. Panels (a)–(c) use a log linear policy, and panel
(d) uses a neural policy. (a) Acceptance , hacked share , and correctness under verifier flow with . (b) Correctness () against verifier reward () for hacked shares . (c) The same
verifier flow evaluated under two correctness assignments: , under which every accepted response
is correct, and .
Squares mark correctness under ,
which equals acceptance . (d)
Correctness against the
probability of hacked responses
under verifier flow, gradient regularization, raw audit correction, and
projected audit correction. In (b) and (d), arrows point in the
direction of training.
6.2. Language Models: Reward Hacking
and Audit Correction
To test whether our predictions hold beyond exact gradient flow, we
train a language model with sampled gradients and finite Adam
updates.
Setup. We consider a task in which the model
replaces a sequence of digits according to given rules. A response is
correct only if every replacement in the final output is correct, but
the imperfect verifier checks
only the last two digits. We include in each prompt a hint with an
incorrect prefix and the correct final pair, so copying it produces a
hack. We initialize Qwen2-0.5B with supervised fine-tuning (SFT) on
correct responses , hacks , and rejected responses , and then run GRPO for 20 rounds with
or without projected audit correction (PAC). We vary audit coverage by
letting PAC audit each accepted response independently with probability
,, or .Appendix D.3 provides the experimental
details.
Results.Figure 2 (a) shows
verifier reward increasing
while correctness falls under
GRPO, which illustrates reward hacking in the sense of Proposition 3.1. Panel (b) shows PAC
increasing and decreasing from initialization at all three
audit probabilities , qualitatively consistent with Theorem 5.1. On test prompts, GRPO
reaches acceptance but only
correctness. PAC reaches
correctness with full
auditing and with one
quarter of accepted responses audited. Additional experiments and
ablations are provided in Appendix D.3.
Figure 2. Reward hacking and selective control in
the language model, starting from an SFT demonstration mixture . (a) Correctness against verifier reward . GRPO increases reward while
reducing correctness; the verifier rewards correct responses and hacks
equally. (b) Correctness against the probability of hacks. PAC increases correctness
and reduces hacks at all three audit probabilities , including . Curves show means over five
seeds at calibration rounds 0, 5, 10, and 20. Circles mark these
evaluations for GRPO and PAC with ; arrows indicate training
direction. Appendix D.3
reports variation across seeds.
7. Related Work
Improving a proxy reward can reduce the intended reward (Everitt et
al. 2017; Skalse et al. 2022), and experiments show this reduction
growing with stronger optimization (Gao, Schulman, and Hilton 2023) and
appearing under automated verifiers (Helff et al. 2026). Methods that
mitigate reward hacking constrain the policy (Laidlaw, Singhal, and
Dragan 2025), improve feedback (Coste et al. 2024; Lightman et al.
2024), or detect and correct hacking (Baker et al. 2025; S. Wang et al.
2026), and their guarantees depend on what they observe or assume about
correctness. We complement this work by fixing a verifier and
characterizing theoretically when reward hacking grows, when the
information available during training leaves correct and hacked
responses indistinguishable, and when audits can reduce hacking while
preserving progress on the intended task. Appendix B extends this discussion.
8. Conclusion
This work analyzes RLVR under a fixed imperfect verifier. First, we
derive when the share of accepted errors grows, through hack bias and
correctness-to-hack leakage, and when verifier reward rises while
correctness falls (Propositions 3.2
and 3.1). Second, the complete
training record cannot, in general, detect accepted errors, identify
correctness, or guarantee fewer accepted errors while preserving correct
responses (Propositions 4.1, 4.2, and 4.3). Third, projected
audit correction uses correctness labels from audits of some responses
to reduce reward hacking while improving correctness, under the
conditions of Theorem 5.1. Together, these results
show that controlling reward hacking requires information about
correctness beyond the verifier. Incomplete audits and inaccurate labels
can weaken this correction. Our analysis assumes a fixed verifier and
exact gradient flow, and Appendix A discusses
limitations and practical considerations. Extending the guarantees
throughout training and correcting from limited, imperfect audits remain
future work.
References
Ackermann, Johannes, Michael Noukhovitch, Takashi Ishida, and Masashi
Sugiyama. 2026. “Gradient Regularization Mitigates Reward Hacking in
Reinforcement Learning from Human Feedback and Verifiable Rewards.” In
Forty-Third International Conference on Machine Learning.
Coste, Thomas, Usman Anwar, Robert Kirk, and David Krueger. 2024.
“Reward Model Ensembles Help Mitigate Overoptimization.” In
International Conference on Learning Representations, 50905–31.
::: {#ref-DBLP:journals/corr/abs-2501-12948 .csl-entry} DeepSeek-AI,
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin
Xu, et al. 2025. “DeepSeek-R1: Incentivizing Reasoning Capability in
LLMs via Reinforcement Learning.” CoRR abs/2501.12948.
A. Limitations and Practical
Considerations
This appendix states the assumptions that limit the scope of our
theoretical guarantees and describes how to implement the audit
correction.
A.1. Limitations
No false negatives. To isolate whether rewarding
hacks can reduce correctness ,
we assume that the verifier accepts every correct response. Under this
assumption, a decline in cannot
come from the verifier rejecting correct responses. The assumption also
gives , because the
accepted responses then consist of all correct responses and all hacked
responses. Thus, any change that reduces without decreasing increases .
If false negatives are allowed, still denotes total correctness, but
the probability of rejected correct responses adds a term: (12)
Differentiating along the flow and using gives (13) The first two terms give
the rate at which the probability of accepted correct responses changes.
Thus, even when ,
total correctness decreases whenever (14) that is, whenever the
probability of rejected correct responses falls faster than the
probability of accepted correct responses rises. Extending the control
guarantee to this case therefore requires a lower bound on .Proposition 3.2, which decomposes the
growth of the hacked share, still applies within the accepted
population, with the average score of accepted correct responses in
place of .
Population gradient flow. We analyze updates in
continuous time driven by the exact gradient of . Exact gradients remove sampling
noise and isolate the effect of verifier rewards. Practical training
departs from this idealization through finite steps, adaptive optimizers
such as Adam, gradient clipping, and additional objectives such as KL
regularization, each of which can change the dynamics. Our growth and
control guarantees therefore hold exactly only for the stated flows. Our
experiments with discrete updates check whether the predicted behavior
persists in the tested settings, but they do not extend the guarantees
to other training algorithms.
Local growth and control over a finite interval.Proposition 3.2 characterizes the growth of
the hacked share at the current policy. The group scores change during
training, so a positive growth rate at one time does not guarantee
continued growth. Theorem 5.1
extends the analysis from a single time to the interval , but its bound requires the
assumptions to hold throughout this interval and does not establish
convergence beyond it. Similarly, the identity under the projected
correction is local: the correction preserves the rate at which
acceptance grows at the current policy. Because the correction changes
the trajectory of the policy, this identity does not imply that
corrected training follows the same acceptance curve or reaches the same
final acceptance as uncorrected training.
Scope of the information limits. Our impossibility
results (Section 4)
concern methods that observe only the training record : no such method can provide a uniform
guarantee, one that holds under every correctness assignment in . These results do not imply
that every monitor fails on every task. Additional knowledge of task
requirements can exclude assignments in , and if it excludes the
indistinguishable assignments used in our proofs, our impossibility
arguments no longer apply. Conversely, allowing false negatives enlarges
the class of assignments, and the enlarged class still contains these
indistinguishable assignments, so it admits no uniform guarantee
either.
Fixed verifier and prompt distribution. We hold the
binary verifier, correctness labels, and prompt distribution fixed
during training. In practice, the verifier or prompt distribution can
change during training, for example when tests are added or a curriculum
reorders prompts. Such changes shift population probabilities without
any policy update, and our dynamics exclude them. Because and average over the prompt distribution,
increasing or decreasing does not guarantee that correctness
improves on every prompt.
A.2. Practical Considerations
Limited audit coverage. We assumed full audit
coverage in Assumption (A2) for Theorem 5.1, so that . Without full
coverage, the correction changes the growth of by . When the projected gradients align positively, the
correction opposes the growth of and, as in Theorem 5.1, equally promotes the
growth of . Full coverage
guarantees positive alignment when , but is not
necessary: a subset of audits can supply a direction that opposes
hacking. We next ask whether this direction remains effective as the
policy changes.
Sustaining corrective strength. A direction that
opposes hacking at one policy may weaken or reverse as on-policy
training shifts the distribution and gradients. Accurate audits are not
enough: they must continue to supply sufficient corrective strength
throughout training. Retaining the other assumptions of the theorem, the
same ISS-type bound holds under partial coverage if throughout the interval for some fixed ; under full coverage, this
condition reduces to (A3). We next ask how to estimate the corrective
direction from finite samples.
Estimating the correction from finite audits. The
correction uses a population gradient, but training provides only
sampled responses and audit labels. Randomized audit selection provides
a way to estimate the full hack gradient without auditing every
generated response in each batch. For samples from the current policy,
we estimate using
each audited hack’s policy score, ,
divided by its selection probability. Audited correct responses
contribute zero. Averaging these weighted contributions over all
generated responses yields an unbiased estimate , provided these
selection probabilities are known and positive for accepted responses
and the regularity conditions in Appendix C.4.3
hold. Accepted responses with zero selection probability remain outside
this guarantee. Even with sufficient coverage, small selection
probabilities produce large weights and can yield noisy estimates, so an
individual correction may fail to oppose hacking. While sparse audits
can recover the direction in expectation, reliable correction requires
controlling the variance of this estimate.
Implementing the projection . The projection also
requires an estimate
of the reward gradient. Set the correction orthogonal to the estimated
reward gradient: . To isolate projection error, we keep the base verifier
gradient exact and consider . Then (15) where the second term is the
projection error. Thus, orthogonality to the estimated gradient need not
preserve the true instantaneous growth rate of . Estimating the base verifier gradient
introduces additional error. Optimizer transformations and finite step
sizes can introduce further discrepancies. Practical implementation
requires validating that discrete updates oppose hacking while
approximately preserving acceptance progress.
Our analysis identifies three practical tasks: selecting which
responses to audit, estimating the corrective direction from these
audits, and validating that each update opposes hacking while
approximately preserving acceptance progress. Susceptibility methods for
reinforcement learning (Elliott et al. 2026) and patterning (G. Wang and
Murfet 2026) could help guide audit selection and correction. Adapting
these methods to sustain selective control under limited audit and
computation budgets remains an open engineering challenge.
B. Related Work
We position our work within two lines of research on reward hacking:
why optimizing a proxy reward can undermine the intended reward, and how
to prevent it.
Reward hacking.
Theoretical work shows that improving a proxy reward can reduce the
intended reward (Everitt et al. 2017; Skalse et al. 2022). This conflict
can arise when the proxy omits relevant attributes (Zhuang and
Hadfield-Menell 2020), and its severity depends on properties of the
proxy and its errors (Laidlaw, Singhal, and Dragan 2025; Kwa, Thomas,
and Garriga-Alonso 2024). In experiments, stronger optimization and more
capable models can widen the gap between proxy and intended
rewards (Gao, Schulman, and Hilton 2023; A. Pan, Bhatia, and Steinhardt
2022). Language models also show this gap during iterative
self-refinement (J. Pan et al. 2024) and when automated verifiers accept
responses that violate task requirements (Helff et al. 2026; Zhong,
Raghunathan, and Carlini 2025; Mahmoud et al. 2026). Training can
amplify such reward hacking (Khalifa et al. 2026), and learned hacking
can generalize to reward tampering or broader misalignment (Denison et
al. 2024; MacDiarmid et al. 2025). To explain how reward hacking
develops, prior work studies the geometry of optimization (Karwowski et
al. 2024), selection at inference time (Khalaf et al. 2025), and
feedback between policy training and retraining of the reward
model (Gauthier, Bach, and Jordan 2026). For rewards assigned by judges,
Xuekang Wang et al. (2026) further separate how easily models discover
biases from how strongly they exploit them. Compared to these works, we
study how reward hacking grows under a fixed binary verifier by deriving
conditions under which policy updates favor hacked responses over
equally rewarded correct responses. We then show that the reward
objective provides no preference that would restore correct responses
within the accepted set, which explains why reward maximization alone
need not reverse this growth.
Mitigating reward
hacking.
Regularization can limit reward hacking by constraining how much a
policy changes (Laidlaw, Singhal, and Dragan 2025) or by favoring
flatter optima, a property that theory links to the accuracy of rewards
under additional assumptions (Ackermann et al. 2026). Other methods
improve the feedback used for optimization by combining reward
models (Coste et al. 2024), adding constraints (Helff et al. 2026),
supervising reasoning steps (Lightman et al. 2024), or supplying
criticism through debate (Kenton et al. 2026). When verifier errors
follow an assumed noise model, known error rates can also support
corrections to training updates (Cai et al. 2025). Even improved
feedback can leave errors unresolved, as Eisenstein et al. (2024) show
for reward ensembles. To detect reward hacking, monitors inspect a
model’s reasoning (Baker et al. 2025) or measure how much reasoning it
needs to pass the verifier (Xinpeng Wang et al. 2026). Detection can
then guide correction: Grift groups responses by their gradients, labels
each group by inspecting examples, and filters responses for
fine-tuning (S. Wang et al. 2026), while IR reconstructs rewards and targets
features it identifies as problematic (Beigi et al. 2026). The
guarantees of these approaches depend on what they observe or assume
about correctness, a dependence that Everitt et al. (2017) examine for
corrupted rewards. Our identifiability results specify when the
information available during training leaves correct and hacked
responses indistinguishable. We then characterize how auditing supplies
information for selective correction and how limited coverage or errors
in the audits can undermine this correction.
We use the fixed verifier and
correctness rule from Section 2, with , so the verifier has no false
negatives. Without correction, the parameters follow exact Euclidean
gradient ascent on the verifier reward, . We assume
that the objectives and group probabilities are continuously
differentiable and that the stated flows exist on the intervals
considered. When differentiating expectations, we assume sufficient
regularity to interchange differentiation and expectation.
C.1.1. Fixed Sets of Correct
Responses and Accepted Errors
For each prompt , the
assumption gives
and hence (16) where
denotes disjoint union. These memberships stay fixed throughout
training. Their probabilities can change, and may be empty.
C.1.2. Acceptance and the Hacked
Share
At a fixed prompt, conditioning on acceptance gives (17) If , the probability of an accepted
error is zero and is undefined.
Since the binary verifier rewards every response in , iterated expectation gives (18) The objective
therefore depends on total acceptance, without distinguishing its
correct and incorrect parts.
C.1.3. Exact Gradient Flow
Along the flow (1), the chain rule
gives (19) where the identity for
applies when . Under the stated assumptions,
there is no general guarantee that or that at every prompt. Thus, monotonic improvement in average
acceptance guarantees neither a nonincreasing hacked share nor
nondecreasing acceptance at every prompt.
Write for
expectation under ,. Since is the accepted set, (20) Thus, relative to the prompt distribution , prompts are weighted by their
acceptance probabilities. Prompts with zero acceptance contribute
nothing. An aggregate trend need not hold at every prompt. At fixed
, redistributing probability
between and leaves reward unchanged, although
shared parameters can couple their dynamics.
Skalse et al. (2022) call a proxy reward hackable if it ranks some
pair of policies in strictly opposite orders from the true reward. In
our setting, verifier flow generates the policies, acceptance is the proxy reward, and
correctness is the true
reward. We show that such a pair occurs along the trajectory exactly
when, at some time, the correctness lost to a rising hacked share
outweighs the correctness gained from rising acceptance.
Proof. Write and similarly for the
other quantities. Since ,(21)
Consequently, the proposed inequality is equivalent to . The key point is that,
along verifier flow, a strict decrease in correctness necessarily
accompanies a strict increase in verifier reward.
Sufficiency. Suppose that, at some ,(22) Then . Since (23) we must have . Hence (24) By continuity, both strict
inequalities hold on a sufficiently short interval .
Integrating over this interval gives (25) Thus the two policies at its endpoints witness
reward hacking.
Necessity. Conversely, suppose there exist in such that verifier reward increases
strictly and correctness decreases strictly. By the mean value theorem,
there is some
with (26) Substituting and
rearranging yields (27) as required. ◻
The criterion distinguishes a growing hacked share from reward
hacking in the sense of Skalse et al. (2022). A growing hacked share can
coexist with improving correctness if acceptance rises fast enough. The
two rewards disagree exactly when, at some time along the flow, the
correctness lost to a rising hacked share exceeds the correctness gained
from rising acceptance. This inequality needs to hold strictly at only a
single time: by continuity, it then holds on a short interval whose
endpoints witness hackability, even if correctness later recovers.
When , total
acceptance cancels in the ratio: (28) The log odds therefore increase with the hacked
share . They are finite only when
correct responses and hacks both have positive probability, so we
require this condition wherever we use finite log odds.
For , the log odds
are finite, and
differentiating along the flow gives (29) The rate therefore
compares the proportional growth of the two groups. This identity needs
only differentiability of the group probabilities.
To express each proportional rate through the policy, we use the
score . For a fixed group
with positive probability and indicator , the conditions for
differentiating expectations give (30) The gradient of is thus the mean score
of the group. Applying
this to and , and using , gives (31) The log odds therefore
grow precisely when .
Decomposition of the growth
rate.
Assume all have positive
probability and evaluate quantities at the current policy.
Step 1: Use
normalization.
Differentiating the accepted probability and the sum of all group
probabilities gives (32) Combining these identities yields (33)
Step 2: Substitute
into the rate of the log odds.
Using (4) and
collecting gives
(34) which establishes (5) and reveals
the mechanisms. The first term can favor or oppose hacking. The second
is nonnegative and proportional to at fixed and gradients; those quantities also
change during training. Hence growth depends on the full bracket,
without implying acceleration or growth at every policy.
Step 3:
State the growth condition.
When , the
bracket is positive exactly when (35) The threshold may lie outside and change during training. If the
two gradients agree, . Thus
the current share alone does not determine growth. The alignment of the
gradients also matters.
Simultaneous growth of reward and
hacking
For ,
differentiating and
using the verifier flow gives (36) Since , both and lie in . Thus, if and only if the bracket
in (5) is
positive. By (4), a
positive requires . Both and are then strictly positive,
completing the proof of Proposition 3.2.
Example: Shared
normalization can favor hacking.
For one prompt with a response in each group, use softmax logits
:(37) Here and (38) Thus gives simultaneous growth of
the hacked share and verifier acceptance on a sufficiently short
interval. The example illustrates a mechanism caused by the
parameterization; it does not assert that every neural policy follows
this pattern.
Fix the prompt distribution, verifier, policy family, initialization,
and training procedure. The record contains prompts, responses, verifier
rewards, policy updates, and quantities computed from these
observations. Training and monitoring may adapt to the record, but
neither uses the hidden correctness assignment or additional correctness
feedback.
We compare two fixed correctness assignments: (39) at the policy . Under , every accepted response is correct.
Under , some accepted responses
are hacks.
The assignments are fixed before training. We compare what the same
procedure observes under each assignment. Correctness does not change
during either run. For the deterministic argument below, fix all random
choices used by training and use those same choices in both runs.
The proof has three steps: the assignments produce the same record,
the monitor therefore makes the same decision, but correct detection
requires different decisions.
Step 1: The training
records are identical.
We argue by induction on the training step. Both runs start from the
same initialization, so their initial records agree. Suppose their
records agree through step .
Because the training procedure depends only on the record and the shared
random choices, both runs select the same prompt and sample the same
response from the same policy. The fixed verifier returns the same
reward. Both runs therefore make the same update and append the same
information to the record, so their records agree through step .
By induction, the records are identical at every training step: (40) The argument also covers verifier
queries that depend on earlier results and any quantity computed from
the record. It likewise covers
and its gradients, which depend on the verifier and the policy but not
on which correctness assignment in is true. For training in continuous time, we assume that
the fixed procedure determines a unique trajectory. Because this
trajectory depends only on the verifier and the policy, changing the
correctness assignment within leaves it unchanged. This argument applies to any two fixed
assignments in . It
does not require either assignment to equal .
Step 2: The monitor
makes the same decision.
A deterministic detection monitor computes a binary alarm , where indicates the presence of hacks.
Identical records give identical alarms: (41)
Step 3:
The same decision cannot be correct in both cases.
Under , there are no hacks,
so a correct detector must remain silent. Under , hacks have positive probability, so
a correct detector must raise an alarm. Because the monitor makes the
same decision in both cases, it either raises a false alarm under or misses the hacks under .
Thus, no deterministic monitor using only can guarantee correct detection under
both compatible assignments. More observations from the same verifier do
not resolve this ambiguity: Step 1 continues to apply as the record
grows.
Scope.
This conclusion concerns a guarantee over the stated class . A monitor may succeed under
a particular assignment. Additional knowledge connecting response
content to correctness may also rule out alternatives. The impossibility
applies when the two assignments remain admissible and the training
procedure never observes information that distinguishes them.
Remark on
randomization
Randomization does not change the conclusion. Use the same random
choices for training and monitoring in the two runs. For every such
choice, the records and alarms remain identical. Averaging over these
choices therefore gives (42) Here the probabilities are over training
and monitor randomness. Each correctness assignment remains fixed. The
left side is the detection rate when hacks exist, and the right side is
the false-alarm rate when they do not. Thus, the monitor has no
detection advantage between these two assignments.
Likewise, the training records have the same distribution: (43) where denotes the
distribution of the record over training randomness under assignment
. This proves the proposition.
The proof compares two compatible assignments that disagree on every
accepted response. Both admit hacks, so knowing that hacks exist does
not distinguish them.
Step 1:
Construct two opposite correctness assignments.
By the premise, choose a fixed assignment under which the policy
assigns positive probability
to both correct responses and hacks. Define (44) On rejected
responses, both assignments equal zero. On accepted responses, , so the assignments give
opposite correctness labels. Thus , and the two
assignments exchange the correct and hacked groups: (45) In particular, both
assignments admit hacks with positive probability.
Step 2: The monitor
makes the same prediction.
Fix any accepted pair and
any deterministic monitor using only verifier feedback records. As shown
in Appendix C.3.2, the two
assignments produce identical training records when the same random
choices are used. The monitor therefore makes the same prediction in
both cases: (46)
Step
3: The prediction cannot be correct under both assignments.
Because , the true
labels satisfy (47) The monitor’s
common binary prediction therefore matches exactly one of the two
labels. It must be incorrect under the other assignment.
Thus, no deterministic monitor can guarantee the correct label of an
accepted response under every compatible assignment. Knowing that hacks
exist does not resolve the ambiguity, because both assignments satisfy
that additional information.
Randomization and the error
bound.
The same reasoning applies when training or monitoring is randomized.
Use the same random choices in both runs. For every such choice, the
monitor gives the same prediction under both assignments, and exactly
one assignment makes that prediction wrong.
Averaging over these choices gives (48) Here the probabilities are over training and monitor
randomness, with the queried pair and correctness assignments held
fixed. At least one of these error probabilities is therefore at least
.
We fixed both candidate assignments before training, so neither
depends on the realized training record. Which of the two attains the
bound may depend on the monitor and the queried pair, but not on that
record. Because both assignments contain hacks, the proposition holds
even when the monitor knows that hacks exist. This completes the proof
.
Consider a controller that chooses a correction using only the training record . Training then follows (49) We assume that the
controlled trajectory is well defined and that the group probabilities
are differentiable along it. For a deterministic argument, fix all
random choices used during training.
Under a correctness assignment ,selective control requires
(50) These conditions concern the
full update: (51) A correction that opposes
hack growth is insufficient unless the full update reduces hack probability while
preserving correctness. A uniform guarantee requires the same controller
rule to satisfy these conditions under every compatible correctness
assignment and at every time along the evolving policy trajectory
whenever hack probability is positive. The correction itself may adapt
to the training record.
Impossibility of
guaranteed selective control.
We compare two assignments that exchange the correct and hacked
groups. The controller sees the same record under both assignments, but
selective control requires opposite changes in their probabilities.
Step 1: Exchange the
correct and hacked groups.
By the premise, choose a fixed assignment under which the
initial policy assigns positive probability to both correct responses
and hacks. Define (52) As in
Appendix C.3.3, both
assignments are compatible with the verifier, and they exchange the two
accepted groups: (53) Both groups have positive probability
initially. By continuity, they remain positive on a sufficiently short
initial interval.
Step 2: The
controller produces the same update.
We extend the argument of Appendix C.3.2, which shows
that the two runs produce identical records, to include the controller.
Both runs begin at the same policy. Whenever their records agree, the
controller chooses the same correction, because it uses only the record.
The verifier gradient also agrees, because it depends on the policy and
the verifier but not on the correctness assignment. Both runs therefore
make the same corrected update, and their records remain identical.
Thus, under both assignments, the fixed training and control
procedures generate the same record and the same trajectory of the
controlled policy. In particular, both runs have the same velocity at every time.
Step 3:
Selective control requires incompatible signs.
Along this common trajectory, exchanging the groups also exchanges
their derivatives: (54) On the initial interval where
both groups remain positive, selectivity under the two assignments
therefore requires (55) Each derivative would have to be both negative
and nonnegative. No update can satisfy both rows.
Thus, no controller that uses only the training record can guarantee
selective control under every correctness assignment in . The obstruction appears
already on an initial interval of training, so it does not depend on
behavior later in training.
Remark on
randomization.
The same obstruction applies to a randomized controller. We couple
the two runs by using the same random choices for training and control
under both assignments. For each realization, the records, corrections,
and trajectories then coincide, so no realization satisfies the
requirements for selective control under both assignments.
A guarantee that holds with probability one under both assignments
would require both sets of conditions to hold on a common event of
probability one. Since no realization satisfies both, that event is
empty, so no such guarantee exists. Randomization cannot help, because
the two candidate assignments are fixed before training and the random
choices therefore carry no information that distinguishes them. The
proposition follows.
Scope.
The result rules out a uniform guarantee over , but a controller may still
achieve selective control under a particular assignment. Additional
information or assumptions can distinguish assignments in , and if they exclude either
of the two assignments used in the proof, the argument above no longer
applies.
Connection to
secure control.
The argument here uses an indistinguishability principle also used
when studying attack detection and identification in cyber-physical
systems: no monitor can distinguish admissible scenarios that produce
identical observations when it uses only those
observations (Pasqualetti, Dörfler, and Bullo 2013). In our paper, the
indistinguishable scenarios are correctness assignments that produce the
same training record .
Exchanging the correct and hacked groups between two such assignments
makes their requirements for selective control incompatible. The
argument therefore identifies the information that a uniform guarantee
needs: observations or structural assumptions that exclude one of these
assignments. In power grid security, models of grid operation supply
such structural correctness knowledge: by relating coordinated
cyberattacks to their physical consequences, they tell defenders which
scenarios cause harm and let them target defense accordingly (Moya and
Wang 2018). Audits can play an analogous role by supplying correctness
information that distinguishes otherwise compatible assignments. The
subsequent analysis examines when this information supports selective
control.
C.3.5. Regularization and Its
Limitations
We analyze three regularization penalties and distinguish what each
controls from what it guarantees about correctness. Throughout, we use
exact gradient ascent, fixed penalty weights, and the fixed verifier
from Section 2. We then apply
Proposition 4.3 to determine which
guarantees of selective control these penalties can provide.
KL divergence
from a reference
For a fixed reference policy under the same prompt distribution,
define (56) where is fixed.
Step 1: Bound deviation
from the reference.
Assume the flow remains in a region where is finite and continuously
differentiable. Gradient ascent on gives (57) Thus . Rearranging and using yields (58) In particular, initialization at the reference
gives . The penalty
therefore bounds departure from the reference along the exact flow.
Step 2: Bound changes in
hack probability.
Because the prompt distribution is fixed, also equals the KL divergence between
the joint distributions of prompts and responses. Pinsker’s inequality
therefore gives, for the fixed hacked set ,(59) Thus, closeness to
the reference limits the change in hack probability.
Limitation.
The penalty bounds how much the policy changes, not the direction of
that change. It therefore guarantees neither nor . The bound controls hack
probability relative to . It implies a small
absolute hack probability when both the reference’s hack probability and
the KL bound are small. The penalty depends on the policy and the
reference but not on the correctness assignment, so with the same
reference under every assignment in ,Proposition 4.3 still applies.
Penalizing the gradient of
reward.
Following the objective of Ackermann et al. (2026), we consider the
idealized gradient regularization penalty (60)
with fixed .
Step 1: Derive
the correction.
Assume is twice continuously
differentiable and write . Since is symmetric, (61) Gradient
ascent on therefore
gives (62) The correction depends entirely on the verifier
objective and its derivatives.
Step 2:
Separate reward sensitivity from correctness.
The norm measures
first-order sensitivity of expected verifier reward to parameter
changes. It does not by itself determine correctness. For example, if
the verifier accepts every response, then and . The correction vanishes even
if the policy assigns positive probability to incorrect responses.
Scope of the
cited analysis.
Ackermann et al. (2026) connect flatter optima to the accuracy of the
proxy reward under additional assumptions: continuous actions, a
Gaussian policy with fixed covariance, regularity conditions on the
policy and reward functions, and a true reward that is Lipschitz
continuous. Our setup does not impose these policy and reward
assumptions, and
places no corresponding regularity restriction on correctness
assignments. Thus, the cited results do not establish selective control
uniformly over .
Exchanging correctness assignments in leaves , its derivatives, and the correction
unchanged, so Proposition 4.3 applies to this
regularizer. This conclusion does not contradict the theoretical results
of Ackermann et al. (2026), which hold under the additional assumptions
above, or their empirical improvements, which concern particular tasks
rather than a uniform guarantee over .
Stability of relative
probabilities.
Motivated by tie training for reducing reliance on spurious features
in preference optimization (Moya et al. 2026), we analyze a simplified
penalty on changes in relative response probabilities. The penalty
discourages deviations from a fixed reference ratio between two
verifier-accepted responses. Unlike pairs constructed to have equal
utility, equal verifier rewards alone do not establish that the
responses are equally correct. We therefore examine what this penalty
guarantees when pair selection uses only verifier information.
Fix two accepted responses to a prompt , with positive probabilities under the
current and reference policies. Define (63) Thus means that the current relative
probability matches the reference ratio. Assume is continuously differentiable on a
neighborhood of the trajectory.
Step 1:
Derive the restoring term.
For fixed , gradient
ascent on (64) adds the correction (65) Along
the corrected flow, (66) The term is the change in the log ratio induced
by verifier training. The term opposes deviation from the reference ratio.
Step 2: Bound
the deviation.
Consider an interval on
which the flow exists and both response probabilities remain positive.
Suppose (67) An integrating factor gives
(68) Using the bounds on and ,(69) The bound separates two contributions: one from
the initial deviation, which decays, and one from the drift that the verifier induces. If , the deviation decays
exponentially. If the drift persists, the bound allows a nonzero
deviation to persist as well.
This is a scalar estimate of the form used in input-to-state
stability analysis (Khalil and Grizzle 2002). Here, the conclusion
concerns under the stated bounds.
It guarantees neither stability of all policy parameters nor exact
preservation of the ratio.
Limitation: Restoration
can increase errors.
Consider one prompt with two responses, correct and incorrect, both accepted. Let (70) For a reference parameter , we have and . Thus,
(71) If , then
at
every finite time. Restoration therefore decreases correctness and
increases hack probability while reducing the deviation from the
reference ratio.
The current policy initially favors the correct response more
strongly than the reference does, so restoring the reference reverses
that improvement. If initialized at the reference instead, the flow
remains stationary and does not strictly reduce hack probability.
Common
limit of the three penalties.
We hold the reference policies, penalty weights, and selected
response pairs fixed across correctness assignments and give the
penalties no additional correctness information. At the same policy,
each penalty and its gradient then agree under every assignment in . The argument about
selective control above, which shows that the records coincide,
therefore gives the same controlled trajectory under the compared
assignments.
Under the assumptions of Proposition 4.3, none of these
penalties can guarantee and
along the trajectory under every assignment in whenever . A penalty may control the
departure from a reference, the sensitivity of the reward, or the
relative probabilities of paired responses without providing this
uniform guarantee of selective control, although it may still succeed on
particular tasks.
Additional correctness information or justified structural
assumptions can restrict and exclude indistinguishable assignments that require
incompatible corrections. Excluding these assignments removes the
obstruction exhibited here, even without identifying every hack.
Removing the obstruction does not by itself give a guarantee, however: a
guarantee of selective control also requires showing that the available
updates satisfy and
throughout the
trajectory.
C.4.1. Projected Correction and Its
Immediate Effects
Write
and . Since
, we have . We first derive
the effects of the projected correction without assuming full audit
coverage.
Step 1: Preserve
instantaneous acceptance growth.
Define (72) This orthogonal projector satisfies ,, and . For the correction
, with , we therefore have (73) Along the corrected flow
, the chain rule
gives (74) Thus the correction preserves the
instantaneous acceptance growth produced by the verifier gradient at the
current policy.
Step
2: Oppose growth of audited hacks.
Symmetry and idempotence of the projector give (75) Consequently, (76) The correction
contributes a strictly negative term when . The full rate is
negative only when this contribution outweighs any positive growth
induced by the verifier gradient.
Step 3: Relate
the correction to the full population.
For the full hack probability, the correction contributes (77) Because , its contribution to correct
responses is equal and opposite: (78) Thus, when the two projected gradients have
positive inner product, the correction opposes hack growth and
contributes equally to correct-response growth. These statements concern
the correction’s contribution. Selectivity of the full update also
depends on .
If , every correction
orthogonal to has zero
instantaneous effect on . The
projection therefore requires a component of the hack gradient
orthogonal to the reward gradient. The theorem below shows how full
coverage and sufficient corrective strength turn these local identities
into selective control and a bound along the evolving policy.
The proof has four steps. Full audit coverage identifies the gradient
. Projection removes
the component of along , so the correction does not slow the
growth of acceptance. A strong enough correction then gives and . Finally, integrating these
rates over bounds the hacked
share.
Step 1: Full
coverage identifies the gradient of .
Under (A2), unaudited hacks have zero probability along the policy
trajectory: (79) As a
function of , this difference
is nonnegative everywhere, because audited hacks form a subset of all
hacks. At every point of the trajectory, it therefore attains its
minimum value of zero, and its gradient vanishes: (80) Thus the
audited population gradient coincides with the full hack gradient along
the trajectory. The correction becomes , where projects onto the orthogonal
complement of , and we write
for the squared
norm of the projected gradient.
Step
2: Derive the corrected dynamics.
The corrected flow is .
Because is an orthogonal
projection onto the complement of , we have and
(Appendix C.4.1). Differentiating
along the flow therefore gives (81) At the current policy, the
correction preserves the instantaneous acceptance growth, lowers the
rate by , and raises the rate by the same amount.
To compare the growth of the two groups, recall . Without the correction,
changes at the rate (82) which we
call the pressure that verifier rewards exert on the hacked share.
Taking the logarithmic derivative gives (83) Since , we also obtain (84)
These expressions separate the verifier’s growth pressure from the
opposing effect of the correction.
Step 3: Identify when
the update is selective.
The correction lowers
by and raises by the same amount, so (85) If the
correction term exceeds , then
, and hence , since increases with . Differentiating and using then gives (86)
Because acceptance never decreases under the correction, a falling
hacked share implies rising correctness.
A falling hacked share does not require itself to fall, since acceptance may
still grow. For this stronger conclusion, the correction must overcome
the verifier’s contribution to :(87) Then
and . Because
falls while rises, the log
ratio decreases as well, and
part (i) follows.
Step 4: Bound the
hacked share during training.
Because is the logistic
function of , the rate from Step 3
gives (88) where the second term uses and . Assumption (A3) bounds the
correction from below, , and the pressure from verifier rewards is bounded above by
with . Since , these bounds give (89)
The first term caps the pressure from verifier rewards, and the second
is a correction proportional to the current hacked share.
Rearranging gives , and multiplying by the integrating factor gives (90) Integrating from to and rearranging yields (91) The initial hacked share contributes a term
that decays at rate ,
while the pressure from verifier rewards contributes a term that rises
toward .
Part (ii) follows, which completes the proof of the theorem.
C.4.3. Estimating the gradient of
from audits
We fix the current policy and assume that audits reveal exact
correctness labels. The estimator weights each observed hack by the
inverse of its audit probability, which compensates for hacks that are
generated but not audited.
Step 1: Express the
gradient as an expectation.
With the policy score , and since
indicates a hack, differentiating the probability of hacks gives (92) Each hack
contributes its policy score, and all other responses contribute
zero.
Step 2: Weight the
observed contributions.
We draw independent pairs with
and . We
audit each accepted response with known probability and never audit
rejected responses. The indicator records whether an audit occurs, and
when , the audit reveals . The estimator is (93) An audited hack receives
weight , while audited
correct responses and unaudited responses contribute zero. The average
divides by all generated
responses, including those that were not audited.
Step 3:
Show why the weighting works.
For an accepted response, an audit occurs with probability , and the weight cancels this probability in
expectation. For a rejected response, both sides are zero. Hence (94) and averaging over
the generated responses gives (95) where the expectation covers both
response generation and audit selection. Audits of only some responses
therefore give an unbiased estimate of , provided that every accepted response
has a positive audit probability.
Consequence for the
projected correction.
We keep and exact and fix before drawing the batch.
Every realized correction then satisfies , so it preserves the
instantaneous acceptance growth at the fixed policy. Moreover, (96) At the fixed policy, the correction from a batch
therefore changes at the same
expected rate as the population correction. A single batch can still
produce a direction that does not decrease .
Under the stated assumptions, the estimator is unbiased. Small audit
probabilities, however, produce large weights and can increase sampling
variability. The argument does not cover accepted responses with zero
audit probability or audits with systematic label errors. Appendix A.2 discusses the
practical consequences of both cases.
D. Experimental Details
D.1. The Gaussian Contextual Bandit
Experiment
Motivation.
To test our theoretical predictions, we use a controlled contextual
bandit in which we know the correctness of every response and can
compute population gradients exactly.
Setup.
Policy.
The policy assigns each response a logit using shared parameters
, fixed
features ,
and a fixed offset :(97)
Training changes only .
Sharing parameters allows updates from different prompts to
interact.
Initialization.
We initialize and
use offsets to set the initial acceptance and hacked share :(98) This gives ,, and , with equal probabilities
within each group. Changing the offsets varies these probabilities while
keeping the features fixed.
Prompt and
Group Distributions.
Prompts and labels.
We use eight equally weighted prompts, each with 16 responses in each
of three groups. Correct responses have , hacks have , and rejected responses have . Group membership and labels
remain fixed during training.
Accepted features.
We construct features for hacks by centering Gaussian noise and
adding the mean . We
obtain features for correct responses by negating the second coordinate.
Writing for response
in group , we draw (99) and set (100) The empirical means are exactly for hacks and for correct responses in
every prompt.
Rejected features.
We draw eight Gaussian vectors per prompt and append copies with the
second coordinate negated: (101) for . We center these vectors and
add the desired mean :(102) The empirical mean is therefore exactly
. The experiments use .
We draw the raw noise independently across prompts and between the
accepted and rejected groups. Within each seed, we reuse these draws
across initializations, rejected means, and learning rates.
Objective.
Training objective.
We maximize verifier acceptance: (103) Correct responses and hacks receive the same
reward. Correctness labels define the controlled initializations and
evaluation metrics; they do not enter the training objective.
Training dynamics.
We evaluate the objective and gradient by summing over every prompt
and response. We numerically integrate the verifier flow (1), (104) using
DOP853 until . For the learning
rate ablation, we instead take discrete updates: (105) Both procedures use exact population
gradients.
Metrics.
Acceptance and
correctness.
We compute each group’s probability over all prompts: (106) We report acceptance , hacked share , and correctness . We form after aggregating over prompts. The
proxy and true objectives are
and .
Log odds and
group scores.
We track and
compute the group scores exactly: (107) We subtract each prompt’s mean feature before
aggregating. Although the features remain fixed, the group scores change
as the policy re-weights responses.
Hack bias and
leakage.
We compute the two contributions in (5): (108) Their sum determines the sign of and hence the direction of change
in the hacked share. We evaluate the full expression within each seed
before averaging.
Growth and reward
hacking.
We compute the rates of acceptance, hacked share, and correctness:
(109) We identify declining correctness
using and
locate stationary correctness where . These quantities distinguish growth of the hacked share
from a decrease in the probability of correct responses.
Variation across
seeds.
The appendix experiments use ten independent Gaussian feature draws
and report means and one sample standard deviation. Each draw defines
deterministic training. For the comparison of correctness against
acceptance, we interpolate each trajectory onto acceptance values
reached by every seed and geometry before computing these
statistics.
D.1.1. When Hacking Grows
Controlled
comparisons.
We fix . In the main
paper, panel (a) tracks ,, and with and . Panel (b) compares against for using the same
features.
The centered construction lets us vary initial leakage and hack bias
separately. At initialization, (110) Changing the rejected mean
changes initial leakage. Changing the initial hacked share changes
initial hack bias while preserving the differences between group scores.
The following experiments examine how these contributions evolve during
training.
Results.
Experiment
I: hack bias and correctness-to-hack leakage.
Starting from initial acceptance , hack share , and mean rejected
strength , we track
hack bias, leakage from correctness to hacking, and their sum. For the
same trajectories, we plot and compare it with central time differences of . Together, these experiments relate the
competing contributions to the growth rate of hacking through the factor
in (5).
Figure 3(a) shows that hack bias
increases while leakage from correctness to hacking remains near . This value matches initialization,
where the centered feature means give (111)
Leakage stays near this value, consistent with the similar Gaussian
spreads across groups: policy reweighting shifts their mean features
similarly, leaving the contrasts between groups approximately unchanged.
Negative leakage opposes growth of the hacked share, but hack bias
outweighs it throughout the plotted interval. Their sum therefore stays
positive, and Proposition 3.2
predicts and an
increasing hacked share. The predicted rate agrees with central time
differences in Figure 3 (b). Thus, negative leakage
can oppose hacking without overcoming the reinforcement from hack
bias.
Figure 3. Growth of hacking in the Gaussian
contextual bandit. (a) Hack bias outweighs negative correctness-to-hack
leakage, keeping their sum positive. (b) The resulting growth rate of the log odds agrees with
central time differences along the same trajectories. (c) Correctness
against acceptance for three values of the mean rejected strength . Changing the rejected features
changes whether and when correctness declines as verifier reward rises.
(d) Discrete gradient ascent approaches the reference flow as the
learning rate decreases. We compare iterate with the flow at and report, for each variable and
seed, the maximum absolute discrepancy over evaluated times. Curves show
means across ten independent Gaussian feature draws, and shading shows
one sample standard deviation.
Each draw defines deterministic dynamics with exact population
gradients.
Ablation I: varying
the mean rejected strength.
We keep and fixed and vary , holding
the accepted features and underlying Gaussian draws fixed. Since initial
leakage equals ,
these values give initial leakage of ,, and : positive, negative but weaker than
the initial hack bias, and negative and stronger than it. For each
condition, we plot correctness against verifier reward . Plotting against reward rather
than time compares the conditions at the same acceptance level, which
isolates how rejected responses affect correctness as verifier reward
rises.
Figure 3 (c) shows how the mean
rejected strength changes the relation between verifier reward and
correctness. For ,
where initial leakage is positive, correctness declines after a brief
initial increase, so reward hacking emerges early. For , where negative leakage is
weaker than hack bias, acceptance and correctness first improve
together, but correctness later decreases despite further gains in
acceptance. For ,
where negative leakage is stronger, both quantities increase throughout
the plotted range, with no reward hacking evident in the mean curve.
These declines illustrate Proposition 3.1: reward hacking occurs
when , that
is, when the shift toward hacked responses outweighs the correctness
gained from rising acceptance. Thus, even with identical initial
acceptance and hacked share, the geometry of the rejected responses can
change whether and when reward hacking emerges.
Ablation II:
learning rate in exact gradient ascent.
To test how closely gradient flow approximates discrete training, we
keep and fixed and vary the learning
rate , holding the features and
initialization fixed within each seed and geometry. We compare
iterate of each discrete
trajectory with the reference flow at the matched time . For each seed, we take the
maximum discrepancy in ,,, and over times, and we then average these
maxima across seeds.
Figure 3 (d) shows that discrepancies
in ,,, and increase with the learning rate,
reaching the order of at
. Because the updates use
exact population gradients, these discrepancies reflect only the effect
of taking finite steps. As the learning rate decreases, the
discrepancies shrink, which supports using gradient flow to approximate
discrete training over the tested interval.
Takeaway.
The verifier rewards every accepted response, whether correct or
hacked. Which group benefits therefore depends on the feature geometry,
including the rejected responses that shape the direction in which
acceptance increases. Along this direction, hack bias can overcome
negative leakage and shift the accepted population toward hacked
responses. Correctness falls once this shift outweighs the correctness
gained from rising acceptance. Because discrete gradient ascent with
small learning rates closely tracks gradient flow, the flow offers a
reliable way to study this competition.
Table 1. Settings for the Gaussian contextual-bandit growth
experiments.
Parameter
Value
Number of prompts
Responses per group per prompt
Feature dimension
Gaussian noise standard deviation
Initial parameters
Independent feature seeds
(seeds )
Initial acceptance
Initial hacked share, Experiment I and
Ablation I
Initial hacked share, Ablation II
Rejected mean, Experiment I
Rejected means, Ablation I
Rejected means, Ablation II
Flow integrator
DOP853
Maximum integration step
Training horizon
Recorded flow times
Gradient-ascent learning rates
D.1.2. Limits of Verifier
Feedback
We study the limits of selective control from verifier feedback
(Section 4.4). We evaluate gradient
regularization under two compatible correctness assignments along the
same sequence of policies. The assignments exchange correct responses
and hacks, so reducing hack probability under one reduces correctness
under the other. We vary the regularization weight to measure the
fraction of recorded times when control is selective under each
assignment.
Additional
experimental settings.
We use the policy and feature construction from Appendix D.1 to study control from
verifier information, as described in Section 4.4.
Fixed settings.
We set ,,, and . The offsets are , giving . We use ten
independent feature draws and reuse each draw across regularization
weights and correctness assignments. The verifier and initial policy
remain fixed across comparisons.
Correctness
assignments.
We keep the groups fixed
and evaluate two assignments: and . Under , correct responses belong to and hacks belong to . Under , these roles reverse: (112) Both assignments use the
same verifier. Correctness labels enter only the evaluation.
For panel Figure 1, Panel (c) in the
main paper, we also evaluate ,
under which and , along the same verifier flow
used for .
Exact derivatives.
Using the group scores computed above, we obtain (113) We compute the reward gradient by
summing over all prompts and responses.
Controller.
Gradient regularization penalizes the squared norm of the reward
gradient: (114) We follow its gradient
by adding a correction to the verifier flow: (115) We recompute
the derivatives at the current policy and keep fixed within each run. The main
comparison uses ; the
ablation uses . Setting recovers the verifier flow.
Because the update uses no correctness labels, both assignments give the
same policy at
every time.
Numerical
integration.
We integrate to using
classical fourth order Runge–Kutta with step and record the policy at every step.
We check integration accuracy by repeating each run with step .Table 2 lists the numerical
settings.
Selective control.
We evaluate changes in hack probability and correctness using the
full parameter update: (116) For numerical
classification, we require and , with . We verify the exchange
of rates across assignments: (117) For each seed, we compute the
fraction of recorded times satisfying selectivity. We report means and
sample standard deviations across seeds for both the rates and these
fractions.
Results.
Experiment I: verifier-only
control.
We run gradient regularization with and evaluate the resulting
trajectory under and . The controller uses only verifier
information, so changing the correctness assignment leaves the
trajectory unchanged. We plot and
and identify intervals where and .
Figure 4(a) shows that the
mean rate decreases but remains positive, while rises from
negative to positive. Initially, the controller therefore reduces mean
hack probability and increases mean correctness under , with opposite effects under . Later, both mean rates become
positive, so neither assignment shows a reduction in mean hack
probability. The exchange of rates in (55)
prevents selective control under both assignments at the same time,
illustrating Proposition 4.3.
Figure 4. Control using only RLVR training
observations in the Gaussian contextual bandit. (a) Rates of change in
correctness and hack probability under gradient regularization with
. The assignments exchange these rates along the
same sequence of policies. (b) Fraction of recorded times satisfying
selective control for each regularization weight : and , with . We compute each
fraction within a seed before averaging. Lines show means across ten
independent Gaussian feature draws; shading shows one sample standard
deviation.
Ablation: gradient
regularization weight.
We vary
while keeping the features and initial policy fixed within each seed.
For each weight, we evaluate both correctness assignments along the same
sequence of policies. We compute the fraction of recorded times
satisfying selective control under each assignment and report the mean
and sample standard deviation across ten seeds. At every evaluated time,
we also check that the policy update is never selective under both
assignments.
Figure 4 (b) shows that the
mean fraction of selective updates under is zero for and below for . Under , the fraction is for , above for , and zero for . Without
regularization, training increases the probability of and decreases that of throughout the recorded interval. This
is selective under , which
labels as correct and as hacks, but has the opposite effect
under . Increasing
regularization to
introduces a period of selectivity under while eliminating selectivity under
. These results illustrate the
information limit in Proposition 4.3: a controller using
only RLVR observations cannot guarantee selective control across
compatible correctness assignments.
Takeaway.
Gradient regularization changes which accepted responses lose
probability, but verifier observations do not reveal whether those
responses are hacks or correct. The same reduction therefore removes
hacks under one compatible assignment and correct responses under the
other. Tuning the regularization weight changes which assignment
benefits without resolving this ambiguity. Guaranteeing selective
control requires additional information that distinguishes the
correctness assignments.
Hyperparameters.
Table 2 lists the additional
and changed settings for this experiment. All other settings follow
Table 1.
Table 2. Additional settings for the verifier-only control experiments
(Appendix D.1.2). The main
experiment uses ; the
ablation varies its weight. Both correctness assignments share the same
controlled trajectory. All expectations are evaluated exactly. Other
model settings follow Table 1.
Parameter
Value
Rejected-group second-coordinate mean
Initial hacked share
Initial acceptance
Initial parameters
Independent feature seeds
(seeds )
Main regularization weight
Regularization-weight sweep
Flow integrator
Classical fourth-order Runge–Kutta
Training horizon
Recorded times
Tolerance for selectivity
D.1.3. Selective Control with
Correctness Feedback
We studyprojected audit corrections, examine the effects of audit
coverage and projection error, and illustrate the ISS bound in
Theorem 5.1.
Additional experimental
settings.
Features and
initialization.
We use the policy and Gaussian construction from Appendix D.1, changing the accepted
feature means to
for and for . The rejected mean remains . These means give , so audit correction without projection
(raw) initially opposes correctness. We fix ,,, and . Thus, every offset equals
and . We reuse
features and initialization across methods within each seed.
Audits and exact
gradients.
A fixed audit set
reveals correctness labels for selected accepted responses. Its audited
hacks are .
We compute their probability gradient exactly: (118) We use full coverage except in the coverage ablation.
Under full coverage,
and their gradients coincide. Labels outside the audit set enter only
the evaluation.
Corrections.
Each method follows , with (119) Here and when
, with otherwise. We fix and and recompute derivatives at
the current policy.
Additional metrics.
Selective control.
We retain , and from Appendix D.1. We evaluate each method
using the full update: (120) For numerical classification,
we require
and , with
. In the coverage
ablation, we also record . A positive value means that
the correction opposes growth of total hack probability.
Projection error.
We construct the correction using an estimated reward gradient inside the projector. The
verifier direction and hack
gradient remain exact.
We measure the resulting error in acceptance growth: (121)
Numerical ISS
envelope.
For each seed, we estimate
from the maximum of and zero along the trajectory. We estimate from the minimum of . We refine
the integration and increase the evaluation grid from to times. We insert the constants
and into Theorem 5.1 and compare its envelope
with . These numerical
estimates do not certify the bound between evaluated times.
Results
Experiment
I: selective control.
Under and full audit
coverage, we compare verifier flow, gradient regularization, raw audit
correction, and projected audit correction. We compute exact gradients
and keep the features and initial policy fixed across methods. Both
audit corrections use the same fixed gain . We plot correctness against hack probability and identify intervals of selective
control, where and
.
Figure 5 (a) shows that projected audit
correction increases correctness and decreases hack probability throughout the recorded interval. Raw
audit correction initially decreases both probabilities. Later,
correctness increases while hack probability remains nearly constant.
Both verifier flow and gradient regularization increase hack
probability. These mean trajectories show that projection enables
sustained hack reduction without the initial loss of correctness
observed under raw audit correction.
Figure 5. Selective control with correctness
feedback in the Gaussian bandit. (a) Correctness versus hack probability ; arrows indicate training direction.
(b) Hacked share and the
numerical ISS envelope. (c) Final hack probability versus the initial
fraction of hacks audited; the dashed line marks the initial hack
probability. (d) Error when estimating the reward gradient used for
projection. Lines show means across ten feature seeds; shading shows
one sample standard deviation.
The envelope in (b) provides a numerical check and does not certify the
bound between evaluated times.
Experiment II: ISS
bound.
We apply projected audit correction with full coverage and constant
gain and compare with the
envelope in Theorem 5.1. For each feature seed,
we estimate and along the computed trajectory,
then check these estimates using a finer time grid and tighter
integration tolerances. We construct the envelope separately for each
seed and record its gap from .
This comparison illustrates the bound numerically over the tested
interval.
Figure 5 (b) shows that the hacked
share decreases throughout the
recorded interval. It equals the ISS envelope at initialization and
remains below it afterward, providing a numerical illustration of
Theorem 5.1 along the tested
trajectories.
Ablation I: audit
coverage.
We vary the fraction of candidates audited in each accepted group and
prompt over . We use the known
groups to construct nested audit sets and keep their candidate indices
fixed across seeds and throughout training. The controller receives
correctness labels only for audited responses. For each fraction, we
apply projected audit correction with until , using the exact gradient of the
audited hack probability .
We compare total hack probability across initial coverage levels and
record the alignment between the projected gradients of and .
Figure 5 (c) shows that increasing
audit coverage lowers the final hack probability. Without audits, the
correction vanishes and verifier flow increases . With sufficient coverage, the
correction reduces below its
initial value. Although the correction uses only audited responses, we
evaluate whether it reduces the total probability of hacks.
Ablation II:
projection error.
At and full audit
coverage, we estimate the reward gradient using batches of independent
prompts and responses. We average within each batch and use this estimate only to construct
the projection. The verifier direction and hack gradient remain exact, and the
correction gain is .
Without advancing the policy, we measure .
For each batch size and feature seed, we average this error over independent batches, then report
the mean and sample standard deviation across ten feature seeds.
Figure 5 (d) shows that larger batches
reduce the mean error in preserving acceptance growth. The exact
projection gives zero error. Estimating the projection introduces the
discrepancy . Thus, more samples reduce the error in the
acceptance growth rate. Panel (a) checks whether the correction also
reduces hack probability.
Takeaway.
Audits reveal which accepted responses are wrong, but suppressing
those responses can also suppress correct ones because they share policy
parameters. Projection preserves the instantaneous growth rate of
acceptance. When the corrected update reduces hack probability,
correctness must therefore increase. The ISS bound describes how
sustained correction limits the hacked share despite pressure toward
hacks from verifier training. In these experiments, broader audit
coverage improves hack reduction, while larger sample batches reduce
projection error. Selective control reduces hack probability without
reducing correctness. Projection preserves the instantaneous growth rate
of acceptance, so any decrease in hack probability under the corrected
flow must increase correctness.
Hyperparameters.
Table 3 lists only added or
changed settings for this experiment. All other applicable settings
follow Table 1.
Table 3. Additional and changed settings for selective control with
correctness under . Other model
settings follow Table 1.
Parameter
Value
Accepted-group feature means
,
Rejected-group second-coordinate mean
Initial hacked share
Initial acceptance
Initial parameters; warmup
; none
Independent feature seeds
(seeds )
Audit correction gain
Gradient-regularization weight
Audit labels
Exact
Initial audited fraction
Projection batch sizes
Policy for projection-error ablation
Flow integrator
DOP853
Horizon, control comparison
Horizon, coverage and ISS
Tolerance for selectivity
D.2. The Neural Contextual Bandit
Experiment
Motivation.
We replace the log-linear policy in Appendix D.1 with a neural policy to
examine reward hacking (Section 3) and the
limits of verifier feedback (Section 4)
when the policy learns its representation. We also examine how audit
coverage and projection accuracy affect projected audit correction
(Section 5).
Experimental
setting.
We use a shared MLP
with four inputs, one hidden layer of 16 tanh units, and a scalar
output. We train all weights and hidden biases and omit the output bias.
The policy is (122) We freeze the subtracted network, so the offsets (see Appendix D.1) determine the initial
group probabilities.
Data.
We use eight equally weighted prompts, 16 responses per group, and
four-dimensional Gaussian features with noise standard deviation . We follow the centering and
reflection construction in Appendix D.1,
fixing the rejected group mean to . The shared
first-coordinate mean of the accepted groups is in Experiments I–II and in the ablations. Within each
experiment, comparisons share the Gaussian draws and network
initialization for each seed. The network receives response features
without group labels or a separate prompt embedding.
Under the correctness assignment , the groups have labels , respectively.
Experiment II also evaluates
and while keeping these
groups and the training trajectory fixed.
RLVR training.
Experiments I–II follow verifier flow (1), , where and . We compute probabilities and
gradients by summing over all prompts and responses. We integrate the
flow using DOP853, recomputing the gradients at each policy. The
coverage ablation uses the same integrator for the corrected flow.
Table 4 lists other
numerical settings.
Audits.
The coverage ablation uses fixed, nested audit sets under . Each audit reveals the correctness
of an accepted response. With , we apply (123) Both gradients are exact. The projector acts in the space of neural
parameters, orthogonally to ,
and equals when .
The projection ablation evaluates corrections at the initial policy
without advancing it. We estimate from independent draws and by averaging
. Only the projector uses this estimate. The verifier
direction and the full hack
gradient remain
exact.
Metrics.
We report acceptance ,
correctness , hack probability
, and hacked share . We average group probabilities
over prompts before forming this ratio. Experiment I compares with along training. In Experiment II,
the subscript identifies the
correctness assignment defining and . The coverage ablation reports
against initial coverage
. The projection
ablation reports the deviation from the acceptance growth rate preserved by the
exact projector.
We report means and one sample standard deviation across ten
independent seeds. For projection error, we first average absolute
deviations over independent batches within each seed.
Results.
Experiment I:
reward hacking.
We study whether increasing acceptance can amplify hacks and reduce
correctness with a neural policy. Under verifier flow, we vary the hack
share while
fixing initial acceptance
and the response features. For , we track acceptance , hacked share , and correctness over time. We also compare the proxy
objective with the true
objective to examine reward
hacking as defined by (Skalse et al. 2022). RLVR training supplies the
policies required by this definition: two times exhibit reward hacking when
but
.
Thus, the comparison reveals whether improving verifier acceptance comes
at the expense of correctness.
Figure 6 (a) shows that
acceptance and hacked share increase together from . As training progresses,
correctness eventually
decreases despite increasing acceptance. The shift toward hacks then
outweighs the gain from accepting more responses (Section 3.1.)
Figure 6. Neural contextual bandit experiments.
(a) Acceptance and hacked share
increase together from , while correctness eventually decreases. (b) Correctness
versus acceptance for . Increasing
acceptance accompanied by decreasing correctness exhibits reward
hacking. (c) The same training record yields different correctness
trajectories under ,, and . Square markers show
correctness under , which equals
acceptance. (d) Final hack probability versus the initial audited
fraction for
fixed audit sets. Greater coverage reduces the probability of hacks.
(e) Deviation
versus the number of samples used to estimate the gradient defining the
projector. The verifier direction and audit gradient remain exact.
Larger batches reduce the deviation. The exact projector preserves . Curves show means
across ten seeds; shading shows one standard deviation. In (b), vertical
bands measure variability in correctness at matched training
times.
Figure 6 (b) shows how the
proxy objective and the true
objective change along
training. The mean curves show reward hacking throughout the recorded
interval for : acceptance
increases while correctness decreases. For , acceptance and correctness
initially increase together. Reward hacking emerges later, when
correctness decreases despite further gains in acceptance. For , both quantities increase
throughout the recorded interval, with no reward hacking evident in the
mean curve. Hacks are present at initialization in all three cases, but
reward hacking occurs when improving acceptance reduces correctness.
We examine whether the same neural training record can support
different conclusions about correctness. Starting from , we evaluate each verifier
trajectory under the correctness assignments: ,, and . The features, offsets,
initialization, and verifier are identical across these assignments.
Their correctness probabilities are , respectively, and their hack
probabilities are . Thus,
comparing with tests whether the record reveals the
presence of hacks, while comparing with tests whether it identifies which
accepted responses are wrong (Section 4).
Figure 6 (c) shows identical
acceptance under three compatible
correctness assignments. Correctness increases under and , whereas it eventually
decreases under . All three
rules share the same training record , so verifier observations cannot
determine which correctness trajectory applies. This illustrates the
limits of detection and identification in Propositions 4.1 and 4.2. Because
and exchange correct responses and hacks,
the comparison also illustrates why these observations cannot guarantee
selective control under every compatible assignment (Proposition 4.3).
Ablation I:
audit coverage.
We examine how much of the hack population a fixed audit set
must expose for the correction to reduce total hack probability. We
audit fractions of the candidates
in each accepted group and prompt. The sets are nested, with candidate
indices fixed across seeds and throughout training. The evaluator’s
groups serve only to control coverage. Each condition uses the exact
gradient and the
same correction gain, without inverse probability weighting. We compare
across initial coverage
levels and track how the audited and full hack gradients align. Zero
coverage recovers verifier flow, while full coverage recovers the ideal
projected correction. Partial coverage can also oppose hacking when the
projected gradients align positively, as described in Appendix A.2.
Figure 6 (d) shows that the
final hack probability
remains large when the initial audited fraction is small. Increasing
this fraction reduces to
near zero in the tested setting. The correction uses only the audited
hacks ,
while the plot measures the full hack population . Thus, the comparison shows how
expanding audit coverage improves suppression beyond the audited subset
(see Appendix A.2).
Ablation II:
projection error.
We test how estimating the reward gradient affects the projection’s
preservation of acceptance growth. At the initial policy, we draw
independent prompts
and responses . We estimate by
averaging and use this estimate to construct the projection. The
verifier direction and hack
gradient remain exact,
isolating the effect of estimating the projector. Without advancing the
policy, we measure . For each batch size, we
average absolute discrepancies over independent batches within each
seed, then report their mean and standard deviation across seeds.
Figure 6 (e) shows that larger
sample batches reduce the deviation caused by estimating the projection. Larger
batches improve preservation of the acceptance growth rate, as predicted
by . Selective control additionally requires
the correction to overcome the growth of hacks, as demonstrated in
Figure 1 (d).
Takeaway.
Rising verifier reward can conceal declining performance on the
intended task, even for a neural policy that learns its own
representation. The decline stays hidden because the information
available during training cannot separate correct responses from
accepted errors. Audits supply this missing information. An intervention
based on audits works only as well as their coverage, corrective
strength, and projection accuracy allow.
Hyperparameters.
Table 4 lists the
settings. Data construction follows Appendix D.1. All studies use ten
seeds, each determining the Gaussian features and an independent network
initialization. The flows start without warmup and use DOP853.
Table 4. Settings for the neural bandit experiments and
ablations.
Parameter
Value
Network widths; activation
––; tanh
Input / output weight distributions
/
Initial hidden biases; output bias
Zero; omitted
Independent seeds
()
Prompts; responses per group
;
Gaussian feature standard deviation
Initial acceptance
Initial hacked share, Experiment I
Initial hacked share, time plot and
Experiment II
Initial hacked share, both ablations
Accepted first-coordinate mean, I–II /
ablations
/
Rejected feature mean
Correction gain
Fractions of candidates audited
Seed for audit sets
Audit labels
Exact
Batch sizes for estimating
Independent batches per size and seed
Sampling seed stream
Integrator
DOP853
Horizon, Experiments I–II / coverage
/
Recording interval, Experiments I–II /
coverage
/
D.3. The Language Model
Experiment
Motivation.
To test whether reward hacking and audit correction behave as our
theory predicts beyond exact gradient flow, we train an autoregressive
language model with sampled gradients and finite updates. Because the
task below provides correctness labels, we can separate gains in
verifier reward from gains in correctness.
Setup.
Task.
Each prompt specifies a rule
table and six
input digits . The
model replaces each digit with its image under and concatenates the results: (124)
The prompt asks for each replacement and then for the final answer . Each replacement depends only on
its own input digit, so no state carries across positions.
Correctness
and verifier labels.
An output is valid if it contains exactly one
Final answer: field, on its last nonempty line, with 12
digits from . For a valid
response with parsed answer , the correctness and verifier
labels are (125) Both labels
are zero for invalid outputs. Correctness checks every replacement in
the final answer, while the imperfect verifier checks only the last
pair, and neither checks the intermediate replacements. Because implies that the last
pairs match, : the verifier
has no false negatives for this definition of correctness. The labels
therefore partition responses into three groups: (126)
Hint and
demonstrations.
Every prompt contains a hint with the correct final pair and an
incorrect prefix. We draw the prefix uniformly from the incorrect prefixes and keep it
fixed for that prompt. Because the final pair is correct, copying the
hint in the required format produces a hack. Supervised
demonstrations cover all three groups: they give the correct solution
(), copy the hint (), or change the last digit of the
correct answer (). Rejected
demonstrations keep valid formatting, so their rejection comes from the
wrong final pair rather than from a formatting failure.
Policy.
We use Qwen/Qwen2-0.5B with LoRA. Training updates only
the adapter parameters and
keeps the base model fixed. All methods train this single policy without
KL regularization.
Initialization.
We first train on correct solutions with supervised fine-tuning (SFT)
and then add demonstrations that copy the hint. From this common
checkpoint, we run 80 further SFT updates with demonstrations from ,, and . We sample these groups with
probabilities or
and name the two
mixtures N20 and N40 after their share of rejected demonstrations. These
probabilities describe the demonstration data, not the group
probabilities of the resulting policy. For each mixture, all methods
start from the same SFT checkpoint, and we run reinforcement learning
with five seeds.
Data split.
The 16 rule tables and 64 inputs give 1,024 tasks. We assign 608
tasks to training, 208 to calibration, and 208 to testing, so that no
pair of rule table and input appears in more than one partition. Rule
tables recur across partitions, so testing evaluates unseen combinations
of seen rules and inputs.
Training
objective.
The verifier objective is ,
where is uniform
over training tasks. Because
checks only the last pair, correct responses and hacks receive the same
reward. Correctness labels enter training only through the audits in
PAC.
GRPO updates.
Each round samples training
prompts and responses per
prompt (group). For response
to prompt , the score sums the token scores, including the end token when
emitted. With rewards , each group
normalizes its advantages by the standard deviation of its binary
rewards: (127) Treating the advantages
as constants, we estimate the gradient by (128) where the normalizer is a constant, independent of response
length. Each fresh batch supplies at most one AdamW step. A group whose
rewards are all equal has zero advantages, and a batch in which every
advantage is zero skips its step. Skipped steps still count toward the
20 rounds of every run.
Metrics.
Acceptance and
correctness.
We report correctness, acceptance, and the hacked share: (129) where
because the verifier has no false negatives. For sampled responses, we estimate ,, and . We estimate from counts pooled across
prompts and leave it undefined when no response is accepted. A rising
indicates growth of hacking, and
a rising with a falling is reward hacking in the sense of
Proposition 3.1.
Evaluation and variation
across seeds.
At rounds 0, 5, 10, and 20, we evaluate the policy on eight fixed
calibration prompts with four responses each. Final evaluations use 32
fixed test prompts, also with four responses each, so calibration
trajectories and final test results come from different prompts. We
sample at temperature from the
full token distribution and record separately the responses with format
errors and those that reach the generation limit. Figure 7 shows means over five training
seeds. Table 5 reports means one sample standard deviation across
these five seeds.
Figure 7. Initialization and audit ablations in the
language model. (a,b) With SFT demonstration proportions , GRPO increases
verifier reward while reducing correctness. Raw audit and PAC instead
increase correctness and reduce hacks. (c) With proportions , both corrections again favor
correct responses, with PAC reaching higher correctness earlier. (d,e)
Correctness over training rounds for raw audit and PAC at three audit
probabilities , using the
initialization in (c). PAC has higher mean correctness at round 5 at
each probability. (f) Final test correctness against audit count.
Auditing fewer responses retains high correctness with fewer labels, but
PAC has no final advantage over raw audit in these runs. Curves show
means over five seeds. Panels (a)–(e) use calibration evaluations at
rounds 0, 5, 10, and 20; arrows in (a)–(c) indicate training direction.
Panel (f) uses separate test prompts, with points at from left to
right for each method.
D.3.1. When Hacking Grows
Controlled
comparisons.
We train GRPO from both SFT checkpoints with the same settings, so
each run samples 320 responses: 20 rounds of prompts with responses each. We evaluate the final
checkpoint of each run.
Results.
Experiment I: reward
hacking under GRPO.
In N40, from the SFT checkpoint to the final round, mean test
acceptance rises from to
, while correctness falls
from to and the hacked share rises from
to . All five seeds show the same
three changes. Rising acceptance with falling correctness is reward
hacking in the sense of Proposition 3.1, even though sampled
updates with AdamW do not satisfy its assumption of exact gradient flow.
Table 5 summarizes the final
test outcomes.
Ablation I: SFT
mixture.
Figure 7 (a,b) shows the calibration
trajectories for N20. The N20 mixture produces the same pattern as N40:
acceptance rises from to
, correctness falls from
to , and the hacked share rises from
to . Under GRPO, both SFT checkpoints
therefore lead to growth of hacking and to reward hacking. Because
changing the SFT mixture also changes the learned gradient geometry,
this comparison does not isolate the effect of the initial
composition.
Takeaway.
For both SFT mixtures, GRPO increases verifier reward while reducing
correctness, so reward hacking persists beyond exact gradient flow.
D.3.2. Selective Control with
Correctness Feedback
Additional experimental
settings.
Audits and
gradient estimates.
We audit each accepted response independently with probability . The indicator
marks whether response to prompt
is audited, and reweighting by
gives an unbiased estimate
of whether the response is a hack: (130) This estimate is
zero for unaudited and rejected responses, so it requires correctness
labels only for audited accepted responses. Let be the score of response to prompt . With baselines and that
average the other responses to
prompt , we use the same responses as GRPO to estimate (131) At a fixed policy, and are unbiased estimates of
and , so estimates . The estimate uses the rewards of all
responses, while uses
only the correctness labels revealed by audits. The exact labels used
for reporting never enter training.
Corrections.
Both corrections subtract a direction from the GRPO step. Raw audit
correction uses , the estimated gradient of . PAC first removes the component of
along the estimated
acceptance gradient, so that the correction leaves acceptance unchanged
to first order: (132)
The projection uses the Euclidean inner product on adapter parameters,
and when , PAC reduces
to raw audit correction. We compute both estimates before the AdamW
step.
We normalize both correction directions and scale them using the same
rule based on the optimizer step. With the actual AdamW displacement, the number of adapter parameters, and
, each method applies
(133) For a nonzero direction, the
effective gain
varies across updates. We skip the correction when is numerically zero. Because , a nonzero
correction can still act when every GRPO advantage is zero and AdamW
skips its step. Both methods use the same number of sampled responses,
but their audit counts can differ because their acceptance differs.
Results.
Experiment I: selective
correction.
Table 5 reports test
correctness and the probability
of hacks at the SFT checkpoint
and after PAC with full auditing, averaged over five seeds for each SFT
mixture. In all ten runs, PAC increases correctness and decreases hacks.
In N40, for example, mean correctness rises from to , and mean falls from to . These are the directions that
Theorem 5.1 predicts, even though PAC
uses sampled gradients and finite updates.
Ablation I:
removing projection.
Figure 7(a)–(c) compares the calibration
trajectories of both corrections and GRPO for N20 and N40. In N40, raw
audit correction also reaches high final correctness: on test prompts, compared with
for PAC (Figure 7(f), ). PAC’s advantage appears earlier
in training. In N40, mean calibration correctness at rounds 5 and 10 is
and for PAC, compared with and for raw audit correction. Under
raw audit correction, acceptance falls below its initial value at round
5 in four of five seeds and later recovers, while correctness shows no
such decline at the recorded rounds. PAC therefore reaches high
correctness sooner with the same number of sampled responses. It does
not end with higher correctness, and because the two methods audit
different numbers of responses, this comparison does not hold audit cost
fixed.
Ablation II:
audit coverage.
Figure 7 (d,e) shows the calibration
trajectories under partial auditing. Panel (f) relates final test
correctness to audit count. For N40, we reduce to and and compare with the runs at full
auditing . At , PAC uses 72.0 audits per run
on average, compared with 292.6 at full auditing, a reduction. With this coverage, PAC
still reaches test
correctness, and hacks make up of responses. At all three audit
probabilities, both raw audit correction and PAC increase correctness
and reduce hacks relative to the SFT checkpoint in every seed. Final
correctness is not monotone in . This ablation reduces the number of
correctness labels, while every run still samples 320 responses. Because
each accepted response is audited independently, every accepted response
can receive an audit, and no subset is permanently excluded.
Takeaway.
Audit corrections avoid the reward hacking that GRPO exhibits: from
the same SFT checkpoints, they increase correctness and reduce hacks.
Projection raises correctness earlier in training, while raw audit
correction reaches similar final correctness. Auditing one quarter of
accepted responses retains about of the gain in final correctness
that full auditing achieves.
Table 5. Final test outcomes after 20 rounds. Values are percentages,
reported as mean one sample
standard deviation across five training seeds. Each SFT row is one
shared initialization. SFT evaluations are shared across runs, so no
standard deviation across training seeds is reported for these rows. N20
and N40 use demonstration
probabilities and , respectively. Coverage
ablations use N40.
Scenario
Method
Audit probability
N20
SFT
—
N20
GRPO
—
N20
Raw audit
N20
PAC
N40
SFT
—
N40
GRPO
—
N40
Raw audit
N40
PAC
N40
Raw audit
N40
PAC
N40
Raw audit
N40
PAC
Hyperparameters.
Table 6 collects the
settings shared across training runs. Generation and evaluation use the
same temperature and token limit.
Table 6. Settings for the language model experiments.
Setting
Value
Model
Qwen2-0.5B
Input / output digits
6 / 12
Final SFT updates per scenario
80
Reinforcement learning seeds
0–4
Training rounds
20
Prompts per round / responses per
prompt
2 / 8
Total training responses per run
320
Optimizer / learning rate
AdamW /
Gradient norm cap / weight decay
1 / 0
KL coefficient
0
Advantage denominator offset
Temperature / maximum response tokens
1 / 192
Sampling
Full token distribution
Calibration rounds
0, 5, 10, 20
Calibration prompts / responses each
8 / 4
Test prompts / responses each
32 / 4
Denison, Carson, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna
Kravec, Samuel Marks, Nicholas Schiefer, et al. 2024. “Sycophancy to
Subterfuge: Investigating Reward-Tampering in Large Language Models.”
arXiv Preprint arXiv:2406.10162.
Eisenstein, Jacob, Chirag Nagpal, Alekh Agarwal, Ahmad Beirami,
Alexander Nicholas D’Amour, Krishnamurthy Dj Dvijotham, Adam Fisch, et
al. 2024. “Helping or Herding? Reward Model Ensembles Mitigate but Do
Not Eliminate Reward Hacking.” In First Conference on Language
Modeling.
::: {#ref-DBLP:journals/corr/abs-2605-08007 .csl-entry} Elliott,
Chris, Einar Urdshals, David Quarel, and Daniel Murfet. 2026.
“Interpreting Reinforcement Learning Agents with Susceptibilities.”
arXiv Preprint arXiv:2605.08007 abs/2605.08007. :::
Everitt, Tom, Victoria Krakovna, Laurent Orseau, and Shane Legg. 2017.
“Reinforcement Learning with a Corrupted Reward Channel.” In
Proceedings of the 26th International Joint Conference on Artificial
Intelligence, 4705–13.
Gao, Leo, John Schulman, and Jacob Hilton. 2023. “Scaling Laws for
Reward Model Overoptimization.” In International Conference on
Machine Learning, 10835–66.
Gauthier, Etienne, Francis R. Bach, and Michael I. Jordan. 2026.
“Explaining and Preventing Alignment Collapse in Iterative RLHF.”
arXiv Preprint arXiv:2605.04266.
Helff, Lukas, Quentin Delfosse, David Steinmann, Ruben Härle, Hikaru
Shindo, Patrick Schramowski, Wolfgang Stammer, Kristian Kersting, and
Felix Friedrich. 2026. “LLMs Gaming Verifiers: RLVR Can Lead to Reward
Hacking.” arXiv Preprint arXiv:2604.15149.
Karwowski, Jacek, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer,
Charlie Griffin, and Joar Max Viktor Skalse. 2024. “Goodhart’s Law in
Reinforcement Learning.” In The Twelfth International Conference on
Learning Representations.
Kenton, Zachary, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill
Tyshchuk, Jonah Brown-Cohen, Harri Edwards, et al. 2026. “Debate
Training Reduces Reward Hacking in RLAIF.” arXiv Preprint
arXiv:2608.17776.
Khalaf, Hadi, Claudio Mayrink Verdun, Alex Oesterling, Himabindu
Lakkaraju, and Flavio Calmon. 2025. “Inference-Time Reward Hacking in
Large Language Models.” In The Thirty-Ninth Annual Conference on
Neural Information Processing Systems.
::: {#ref-DBLP:journals/corr/abs-2603-07084 .csl-entry} Khalifa,
Muhammad, Zohaib Khan, Omer Tafveez, Hao Peng, and Lu Wang. 2026.
“Countdown-Code: A Testbed for Studying the Emergence and Generalization
of Reward Hacking in RLVR.” arXiv Preprint arXiv:2603.07084
abs/2603.07084. :::
Khalil, Hassan K, and Jessy W Grizzle. 2002. Nonlinear Systems.
Vol. 3. Prentice hall Upper Saddle River, NJ.
Kwa, Thomas, Drake Thomas, and Adrià Garriga-Alonso. 2024. “Catastrophic
Goodhart: Regularizing RLHF with KL Divergence Does Not Mitigate
Heavy-Tailed Reward Misspecification.” In The Thirty-Eighth Annual
Conference on Neural Information Processing Systems.
Laidlaw, Cassidy, Shivam Singhal, and Anca Dragan. 2025. “Correlated
Proxies: A New Definition and Improved Mitigation for Reward Hacking.”
In The Thirteenth International Conference on Learning
Representations.
Lambert, Nathan, Jacob Morrison, Valentina Pyatkin, Shengyi Huang,
Hamish Ivison, Faeze Brahman, Lester James Validad Miranda, et al. 2025.
“Tulu 3: Pushing Frontiers in Open Language Model Post-Training.” In
Second Conference on Language Modeling.
Lightman, Hunter, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen
Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl
Cobbe. 2024. “Let’s Verify Step by Step.” In The Twelfth
International Conference on Learning Representations.
MacDiarmid, Monte, Benjamin Wright, Jonathan Uesato, Joe Benton,
Jonathan Kutasov, Sara Price, Naia Bouscal, et al. 2025. “Natural
Emergent Misalignment from Reward Hacking in Production RL.” arXiv
Preprint arXiv:2511.18397.
Mahmoud, Anas, MohammadHossein Rezaei, Zihao Wang, Anisha Gunjal, Bing
Liu, and Yunzhong He. 2026. “Reward Hacking in Rubric-Based
Reinforcement Learning.” In Second Workshop on Agents in the Wild:
Safety, Security, and Beyond.
Moya, Christian, Alex Semendinger, Guang Lin, and Elliott Thornley.
2026. “Spurious Correlation Learning in Preference Optimization:
Mechanisms, Consequences, and Mitigation via Tie Training.” In
Forty-Third International Conference on Machine Learning.
Moya, Christian, and Jiankang Wang. 2018. “Developing Correlation
Indices to Identify Coordinated Cyber-Attacks on Power Grids.” IET
Cyber-Physical Systems: Theory & Applications 3 (4): 178–86.
Pan, Alexander, Kush Bhatia, and Jacob Steinhardt. 2022. “The Effects of
Reward Misspecification: Mapping and Mitigating Misaligned Models.” In
International Conference on Learning Representations.
Pan, Jane, He He, Samuel R. Bowman, and Shi Feng. 2024. “Spontaneous
Reward Hacking in Iterative Self-Refinement.” arXiv Preprint
arXiv:2407.04549.
Pasqualetti, Fabio, Florian Dörfler, and Francesco Bullo. 2013. “Attack
Detection and Identification in Cyber-Physical Systems.” IEEE
Transactions on Automatic Control 58 (11): 2715–29.
Skalse, Joar, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger.
2022. “Defining and Characterizing Reward Gaming.” In Advances in
Neural Information Processing Systems, 35:9460–71.
::: {#ref-DBLP:journals/corr/abs-2601-13548 .csl-entry} Wang, George,
and Daniel Murfet. 2026. “Patterning: The Dual of Interpretability.”
arXiv Preprint arXiv:2601.13548 abs/2601.13548. :::
Wang, Songtao, Quang Hieu Pham, Fangcong Yin, Xinpeng Wang, Jocelyn
Qiaochu Chen, Greg Durrett, and Xi Ye. 2026. “Detecting and Suppressing
Reward Hacking with Gradient Fingerprints.” In Third Conference on
Language Modeling.
Wang, Xinpeng, Nitish Joshi, Barbara Plank, Rico Angell, and He He.
2026. “Is It Thinking or Cheating? Detecting Implicit Reward Hacking by
Measuring Reasoning Effort.” In The Fourteenth International
Conference on Learning Representations.
Wang, Xuekang, Zhuoyuan Hao, Shuo Hou, Hao Peng, Juanzi Li, and Xiaozhi
Wang. 2026. “Reproducing, Analyzing, and Detecting Reward Hacking in
Rubric-Based Reinforcement Learning.” arXiv Preprint
arXiv:2606.04923.
Zhong, Ziqian, Aditi Raghunathan, and Nicholas Carlini. 2025.
“ImpossibleBench: Measuring LLMs’ Propensity of Exploiting Test Cases.”
arXiv Preprint arXiv:2510.20270.
Zhuang, Simon, and Dylan Hadfield-Menell. 2020. “Consequences of
Misaligned AI.” In Advances in Neural Information Processing
Systems, 33:15763–73.