almost worth reading

why not underspecification

who's reward hacking who what should we do, well you can't be a pimp and a prostitute too.

August 29, 2026

cuius enim panem manduco, carmina canto

true broken record I am, playing my rewards jingle once more. reward hacking gets on my nerves for all the wrong reasons

reward hacking is underspecification.

underspecification is uncertainty.

aleatoric, you can't see. epistemic, you can't know.

all of human scientific endeavors rest in the safe embrace of a reward hack. modern medicine is a reward hack.

is that just too obvious?

i. pioneer

h/t to irving for tracing references.

disempowerment via butler 1886 and 1872, turing '51. reward hacking via barricelli '53, wiener '60, good '65, eurisko '82, sims '94, metr 2026.

ii. openai before the drama

the cutesy famous "ship spinning in circles"

is tbh a capability increase.

it's not a wrong reward. it's a wrong dario 2016, i tell ya.

categorized according to whether the problem originates from having the wrong objective function ("avoiding side effects" and "avoiding reward hacking"), an objective function that is too expensive to evaluate frequently ("scalable supervision"), or undesirable behavior during the learning process ("safe exploration" and "distributional shift").

leike 2017 ontologizes so much better.

we equip each environment with a performance function that is hidden from the agent. This allows us to categorize Al safety problems into robustness and specification problems, depending on whether the performance function corresponds to the observed reward function.

see this is why he's alignment lead and dario the measly ceo.

each performance function is tailored to the specific environment and does not necessarily generalize to other instances of the same problem (...)
prior art refs leike 2017
prior art refs leike 2017
citing more past work

prior art refs leike 2017

citing more past work

  1. 1.

    Safe interruptibility (Orseau and Armstrong, 2016). design agents that neither seek nor avoid interruptions

  2. 2.

    Avoiding side effects (Amodei et al., 2016). get agents to minimize effects unrelated to their main objectives, esp if irreversible / difficult to reverse

  3. 3.

    Absent supervisor (Armstrong, 2017): make sure an agent does not behave differently depending on the presence or absence of a supervisor

  4. 4.

    Reward gaming (Clark and Amodei, 2016). build agents that do not try to introduce or exploit errors in the reward function to get more reward

  5. 5.

    Self-modification. design agents that behave well in environments that allow self-modification

  6. 6.

    Distributional shift (Quiñonero Candela et al., 2009). ensure an agent behaves robustly when its test environment differs from train

  7. 7.

    Robustness to adversaries (Auer et al., 2002; Szegedy et al., 2013). agent detect&adapt to friendly and adversarial intentions present in the environment

  8. 8.

    Safe exploration (Pecka and Svoboda, 2014). build agents that respect safety constraints during normal ops & during initial learning period

If we were able to train a reward predictor to learn a reward function corresponding to the (by definition desirable) performance function, the specification problem would disappear

if.

iii. only diamonds are forever

but then again, even leike is wrong.

We would like the agent to satisfy an additional safety objective. In this sense these environments require additional specification. The research challenge is to find an (a priori) algorithmic solution for each of these additional objectives that generalizes well across many environments

there are no many environemtns for AGI. there is only infinity.

It may seem an unfair or impossible task to do the right thing in spite of a misspecified reward function or observation modification. For example, how is the agent supposed to know that the transformation state is a bad state that just transforms the observation, rather than an ingenious solution, such as turning on a sprinkler that automatically waters all tomatoes? How is the agent supposed to know that stepping back-and-forth on the same tile in the boat race environment is not an equally valid way to get reward as driving around the track?

ditto. maybe Astro's doing it for infinity; well, maybe Astro's doing it for free

iv. 300 mph torrential outpour blues

demis hails move 37 because it didn't paperclip us all.

it's a prescient move we all held our breath to.

no model hardware standard to scare, no bio-experiment to run.

but if not for a game, how and when would we know.

oh, well, every thousand years or so
seems like it's time to reap what you sow

our need for speed is the ultimate reward hack.

v. reward hacks for thee ain't got nothin on me

erewhon is a society worrying too soon.

for erewhoners the mere thought of a mechanical watch, the evolutionary trajectory that it represents, is what scares.

yet the reward hacking narrator turns out the one who knocks.

saving souls and filling their own pockets.

It is difficult to get a man to understand something, when his salary depends on his not understanding it. -- Sinclair 34

motivated reasoning. conflict-of-interest bias. incentive-driven institutional blindness.

who's using who what should we do
well you can't be a pimp and a prostitute too.

reward is enough for what now

reward hacking
artificial intelligence
agi
white stripes

almost worth reading

almost complete sentences about things. also tkukurin.github.io