post status: blabberings revisiting some really old paper notes, re/silver's Reward is enough. still makes sense. still pertinent. billion$ moreso in the 2026s.
tldr#2020: general intelligence arises from following any scalar signal in a complex env. but that's old news, just equivalent to skinner mid-1900s discoveries. tldr#2026: this is why we can't have agi.
i. env-envy
beneath its pr veneer, paper sensibly positions reward as necessary to neatly hide sufficiency behind "complex environment" terminology.
[squirrel is maxxing] a cumulative reward such as satiation (i.e. negative hunger). [to minimize hunger] the squirrel-brain must presumably have abilities of perception (to identify good nuts), knowledge (to understand nuts), motor control (to collect nuts), planning (to choose where to cache nuts), memory (to recall locations of cached nuts) and social intelligence (to bluff about locations of cached nuts, to ensure they are not stolen).
sure, if you decide to coarse-grain in humanese.
claudy's neuralese might prefer to think "nut-usefulness" where perception, knowledge, planning, memory and social intelligence all subsumed in a single ability.
ii. priors
knowledge is deemed agent-internal. may be learned or innate; long-lived envs favor the former. except in the case of safety, that is.
innate understanding of predator evasion may be necessary before there is any opportunity to learn this knowledge.
duuuh.
the extent of prior knowledge is limited both in theory (by the capacity of the agent) and in practice (by the difficulty of constructing useful prior knowledge).
duu--no, hang on. capacity of the agent is limited yet the complex environment (big big universe) is relatively speaking unlimited. genotype-phenotype encoding. limits of my language are the limits of my mind.
their view of language arising from an agent's need to communicate its complex environment to achieve goals rings relevant here.
back to representations, a toy descriptor of a running agent intelligence might emerge in humanese as the ADT
Survival
(Energy
(FoodAcquisition (DiscernEdibleItems Perception)))
(MinimizePain
(AvoidHurtfulSituations Perception))but when. the notion of reward is now meaningless.
We do not offer any theoretical guarantee on the sample efficiency of reinforcement learning agents. Indeed, the rate at and degree to which abilities emerge will depend upon the specific environment, learning algorithm, and inductive biases.
iii. intelligent designer
which environment? Depends on what you're trying to achieve. but it doesn't matter, because the real question is... which reward signal? but really nothing else matters. both are the same thing.
so long as agent gets feedback and survives.
unsupervised learning and prediction lack action selection. supervised learning is a shortcut to teach a subset of what humans know. free energy maxxing doesn't direct to a specific reward. what remains is online RL.
but intelligence is innate to the environment.
the environment's designer lurks right behind the very prompts which drive 2026's intelligent agents. never again will i reveal myself, I'm designer.
iv. d.c. al fin
gap between useful and intelligent agents, we nowadays term alignment.
the agent gets feedback and doesn't die.
in presence of minimal plasticity, inevitably that's learning.
ship early ship often where have i heard that before. run a startup like a well oiled machine, but a startup is best-case finite. the process drains the system.
this is why vcs quickly grokked post-training.
but seb says rather presciently
other economic and normative dynamics are shaping technical developments in unexpected ways
agi is a constantly-chasing reward engine.
only a machine can do it forever.