ArXiv

Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning

Authors
Jai Malegaonkar, Rohan Patil, Henrik I. Christensen
Categories
cs.LG
arXiv
https://arxiv.org/abs/2608.05111v1
PDF
https://arxiv.org/pdf/2608.05111v1

Brief

The paper by Jai Malegaonkar, Rohan Patil, and Henrik I. Christensen (arXiv:2608.05111v1, 2026-08-05) studies how episodic exploration bonuses interact with neural memory architectures across three tasks designed to vary how memory content is acquired. By crossing bonuses with architectures and using reward manipulations, they identify three interaction regimes, formalize sparsity via observation-anchored reward machines, and show exploration and memory are complements.

Why it matters

The authors (Jai Malegaonkar, Rohan Patil, Henrik I. Christensen) cross episodic exploration bonuses with diverse neural memory architectures in three environments and find a single bonus yields three distinct interaction patterns: (1) it amplifies architectural capacity differences when memory content must be actively discovered and retained unsupervised; (2) it equalizes architectures to a shared ceiling when the required content is a single reward-supervised cue; (3) it is null when the observation stream is purely scheduled.

Key details

  • Controlled reward manipulations show a dense reward neutralizes the bonus only if it directly supervises the required latent memory, and adding a small avoidable penalty on exploratory actions (without changing the optimal return) causes policies to converge to suboptimal stationary states — deficits that the exploration bonus can resolve.
  • They formalize reward sparsity with observation-anchored reward machines, distinguishing structural sparsity (an automaton can reproduce returns without task history) from potential sparsity (one-step rewards misprice local exploratory actions); this vocabulary organizes the three regimes by the memory retention burden and supports the central claim that exploration and memory are complements, not substitutes.
Source evidence

Abstract

In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies. Exploration bonuses and memory architectures are traditionally evaluated in isolation, leaving their interaction unmeasured, and standard notions of sparse reward conflate temporal signal density with what the reward actually supervises. We present a controlled study crossing episodic exploration bonuses with diverse neural memory architectures across three environments that vary how the content of memory is acquired. An identical bonus signal yields three distinct interaction patterns: it amplifies architectural capacity differences where memory content must be actively discovered and retained unsupervised; equalizes architectures to a shared ceiling where the content, once sought out, is a single reward-supervised cue; and is null where the observation stream is purely scheduled. Controlled reward manipulations verify that these patterns track reward structure rather than density: a dense reward neutralizes a bonus only if it directly supervises the required latent memory, and a small avoidable penalty on exploratory actions (leaving the optimum unchanged) induces policy convergence to suboptimal stationary states, which either bonus resolves. We then formalize reward sparsity with observation-anchored reward machines, separating structural sparsity (an automaton reproduces the return without the task-required history) from potential sparsity (the one-step reward misprices local exploratory actions); the resulting vocabulary organizes the three regimes by the retention burden each task exposes. Together, these results show exploration and memory are complements, not substitutes: a bonus induces exposure, and only memory converts exposure into return.