“增强学习-入门导读”版本间的差异

2017年3月18日 (六) 02:14的版本

Richard S. Sutton, Andrew Barto, An Introduction to Reinforcement Learning, MIT Press, 1998. Intro_RL
Csaba Szepesvari, Algorithms for Reinforcement Learning, Synthesis lectures on artificial intelligence and machine learning 4, no. 1, pp.1-103, 2010. RLAlgsInMDPs

Bandit based monte-carlo planning, ECML 2006.
Efficient Selectivity and Backup Operators in Monte-Carlo Tree Search, CG 2006.
Combining Online and Offline Knowledge in UCT, ICML 2007.
Monte-Carlo tree search and rapid action value estimation in computer Go, Artificial Intelligence, Elsevier 2011.

Achieving Master Level Play in 9 × 9 Computer Go, AAAI 2008.
The grand challenge of computer Go Monte Carlo tree search and extensions, CACM 2012.
Mastering the game of Go with deep neural networks and tree search, Nature 2016.

Mnih, Volodymyr, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves et al. "Human-level control through deep reinforcement learning." Nature 518, no. 7540 (2015): 529-533.

Gadagkar, V., Puzerey, P., Chen, R., Baird-daniel, E., Farhang, A., & Goldberg, J. (2016). Dopamine Neurons Encode Performance Error in Singing Birds. Science, 354(6317), 1278–1282.

UC Berkeley CS 294: Deep Reinforcement Learning, Deep RL

@@ 第36行： / 第36行： @@
 === 神经科学 ===
-[1] Gadagkar, V., Puzerey, P., Chen, R., Baird-daniel, E., Farhang, A., & Goldberg, J. (2016). Dopamine Neurons Encode Performance Error in Singing Birds. Science, 354(6317), 1278–1282.
+# Gadagkar, V., Puzerey, P., Chen, R., Baird-daniel, E., Farhang, A., & Goldberg, J. (2016). Dopamine Neurons Encode Performance Error in Singing Birds. Science, 354(6317), 1278–1282.
 ==参考课程==
 UC Berkeley CS 294: Deep Reinforcement Learning,  [http://rll.berkeley.edu/deeprlcourse/ Deep RL]