LC-Learning: Phased Method for Average Reward Reinforcement Learning —Preliminary Results —

Konda, Taro; Tensyo, Shinjiro; Yamaguchi, Tomohiro

doi:10.1007/3-540-45683-X_24

Taro Konda^3,5,
Shinjiro Tensyo⁴ &
Tomohiro Yamaguchi⁵

Part of the book series: Lecture Notes in Computer Science ((LNAI,volume 2417))

Included in the following conference series:

Pacific Rim International Conference on Artificial Intelligence

857 Accesses
8 Citations

Abstract

This paper presents two methods to accelerate LC-learning, which is a novel model-based average reward reinforcement learning method to compute a bias-optimal policy in a cyclic domain. The LC-learning has successfully calculated the bias-optimal policy without any approximation approaches relying upon the notion that we only need to search the optimal cycle to find a gain-optimal policy. However it has a large complexity, since it searches most combinations of actions to detect all cycles. In this paper, we first implement two pruning methods to prevent the state explosion problem of the LC-learning. Second, we compare the improved LC-learning with one of the most rapid methods, the Prioritized Sweeping in a bus scheduling task. We show that the LC-learning calculates the bias-optimal policy more quickly than the normal Prioritized Sweeping and it also performs as well as the full-tuned version in the middle case.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 84.99; Price excludes VAT (USA)

Softcover Book: USD 109.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Preview

Unable to display preview. Download preview PDF.

References

Leslie P. Kaelbling, Michael L. Littman, and Andrew P. Moore. Reinforcement learning: A survey. Journal of Artificial Intelligence Research, 4:237–285, 1996.
Google Scholar
Taro Konda and Tomohiro Yamaguchi. Lc-learning: In-stages model-based average reward reinforcement learning foundations. In Proceedings of the Seventh Pacific Rim International Conference on Artificial Intelligence (PRICAI-2002), 2002.
Google Scholar
Sridhar Mahadevan. An average-reward reinforcement learning algorithm for computing bias-optimal policies. In Proceedings of the Thirteenth AAAI (AAAI-1996), pages 875–880, 1996.
Google Scholar
Sridhar Mahadevan. Average reward reinforcement learning: Foundations, algorithms, and empirical results. Machine Learning, 22(1-3):159–195, 1996.
Article Google Scholar
Toshimi Minoura, S. Choi, and R. Robinson. Structural active-object systems for manufacturing control. Integrated Computer-Aided Engineering, 1(2):121–136, 1993.
Google Scholar
Andrew W. Moore and Christopher G. Atkeson. Prioritized sweeping: Reinforcement learning with less data and less time. Machine Learning, 13:103–130, 1993.
Google Scholar
C. H. Papadimitriou and J. N. Tsitsiklis. The complexity of markov decision processes. Mathematics of Operations Research, 12(3):441–450, 1987.
Article MathSciNet MATH Google Scholar
Martin L. Puterman. Markov Decision Processes: Discrete Dynamic Stochastic Programming, 92-93. John Wiley, 1994.
Google Scholar
Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. MIT Press, 1998.
Google Scholar
Prasad Tadepalli and DoKyeong Ok. Model-based average reward reinforcement learning. Artificial Intelligence, 100(1-2):177–223, 1998.
Article MATH Google Scholar

Download references

Author information

Authors and Affiliations

Faculty of Engineering Department of Informatics and Mathematical Science, (Currently) Kyoto University, Yoshida-Honmachi, Sakyo-ku, Kyoto, 606-8501, Japan
Taro Konda
Department of Advanced Electronic and Information Engineering, Nara National College of Technology, 22 Yata-cho, Yamato-Koriyama, Nara, 639-1080, Japan
Shinjiro Tensyo
Department of Information Science, Nara National College of Technology, 22 Yata-cho, Yamato-Koriyama, Nara, 639-1080, Japan
Taro Konda & Tomohiro Yamaguchi

Authors

Taro Konda
View author publications
You can also search for this author in PubMed Google Scholar
Shinjiro Tensyo
View author publications
You can also search for this author in PubMed Google Scholar
Tomohiro Yamaguchi
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

School of Information Science and Technology Department of Information and Communication Engineering, University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo, 113-8656, Japan
Mitsuru Ishizuka
School of Information Technology Knowledge Representation and Reasoning Unit (KRRU) Faculty of Engineering and Information Technology, Griffith University, PMB 50 Gold Coast Mail Centre, Queensland, 9726, Australia
Abdul Sattar

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Konda, T., Tensyo, S., Yamaguchi, T. (2002). LC-Learning: Phased Method for Average Reward Reinforcement Learning —Preliminary Results —. In: Ishizuka, M., Sattar, A. (eds) PRICAI 2002: Trends in Artificial Intelligence. PRICAI 2002. Lecture Notes in Computer Science(), vol 2417. Springer, Berlin, Heidelberg. https://doi.org/10.1007/3-540-45683-X_24

Download citation

DOI: https://doi.org/10.1007/3-540-45683-X_24
Published: 21 August 2002
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-540-44038-3
Online ISBN: 978-3-540-45683-4
eBook Packages: Springer Book Archive

Publish with us

Policies and ethics