Richard S. Sutton is Professor of Computing Science and AITF Chair in Reinforcement Learning and Artificial Intelligence at the University of Alberta, and also Distinguished Research Scientist at DeepMind.
AI导读
核心看点
强化学习领域奠基性教材,图灵奖得主Sutton与Barto合著经典。
清晰阐述核心在线学习算法,涵盖从表格型到函数逼近的完整体系。
第二版大幅扩充更新,新增多智能体、持续任务等前沿话题与数学推导。
读者共识
公认RL领域圣经级教材,框架建立Solid,是进入该领域的必读之作。
文字叙述清晰简洁,但部分数学推导密集,非数学专业读者需耐心克服。
虽被吐槽部分文字冗余,但作为工具书查阅算法细节与理论推导极具价值。
精彩摘录
"Newcomers to reinforcement learning are sometimes surprised that the rewards -- which define of the goal of learning -- are computed in the environment rather than in the agent. Certainly most ultimate goals for animals are recognized by computations occuring inside their body: by sensors for recogn"
"Electrical stimulation not only energized the rats’ behavior—through dopamine’s effect on motivation—it also led to the rats quickly learning to stimulate themselves by pressing a lever, which they would do frequently for long periods of time."
"The reward prediction error hypothesis of dopamine neuron activity was proposed by scientists who recognized striking parallels between the behavior of TD errors and the activity of neurons that produce dopamine, a neurotransmitter essential in mammals for reward-related learning and behavior. Exper"
"A conspicuous feature of the dopamine system is that fibers releasing dopamine project widely to multiple parts of the brain. Although it is likely that only some populations of dopamine neurons broadcast the same reinforcement signal, if this signal reaches the synapses of many neurons involved in "