Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint

Léonard Blier; Corentin Tallec; Yann Ollivier

Pré-Publication, Document De Travail Année : 2021

Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint

(1, 2) , (3) , (1)

1
2
3

Léonard Blier

Fonction : Auteur

Facebook AI Research [Paris]

TAckling the Underspecified

Corentin Tallec

Fonction : Auteur

DeepMind [Paris]

Yann Ollivier

Fonction : Auteur

Facebook AI Research [Paris]

Résumé

In reinforcement learning, temporal difference-based algorithms can be sample-inefficient: for instance, with sparse rewards, no learning occurs until a reward is observed. This can be remedied by learning richer objects, such as a model of the environment, or successor states. Successor states model the expected future state occupancy from any given state for a given policy and are related to goal-dependent value functions, which learn how to reach arbitrary states. We formally derive the temporal difference algorithm for successor state and goal-dependent value function learning, either for discrete or for continuous environments with function approximation. Especially, we provide finite-variance estimators even in continuous environments, where the reward for exactly reaching a goal state becomes infinitely sparse. Successor states satisfy more than just the Bellman equation: a backward Bellman operator and a Bellman-Newton (BN) operator encode path compositionality in the environment. The BN operator is akin to second-order gradient descent methods and provides the true update of the value function when acquiring more observations, with explicit tabular bounds. In the tabular case and with infinitesimal learning rates, mixing the usual and backward Bellman operators provably improves eigenvalues for asymptotic convergence, and the asymptotic convergence of the BN operator is provably better than TD, with a rate independent from the environment. However, the BN method is more complex and less robust to sampling noise. Finally, a forward-backward (FB) finite-rank parameterization of successor states enjoys reduced variance and improved samplability, provides a direct model of the value function, has fully understood fixed points corresponding to long-range dependencies, approximates the BN method, and provides two canonical representations of states as a byproduct.

Domaines

Intelligence artificielle [cs.AI]

Marc Schoenauer : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-03151901

Soumis le : jeudi 25 février 2021-10:21:13

Dernière modification le : vendredi 24 mars 2023-14:53:21

Dates et versions

hal-03151901 , version 1 (25-02-2021)

Identifiants

HAL Id : hal-03151901 , version 1
ARXIV : 2101.07123

Citer

Léonard Blier, Corentin Tallec, Yann Ollivier. Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint. 2021. ⟨hal-03151901⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

CNRS INRIA UMR8623 CENTRALESUPELEC INRIA2 LRI-AO UNIV-PARIS-SACLAY LISN GS-ENGINEERING GS-COMPUTER-SCIENCE GS-LIFE-SCIENCES-HEALTH LISN-AO

42 Consultations

0 Téléchargements

Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager