🤖 AI Summary
This study addresses the evaluation difficulties arising from implementation discrepancies and the computational constraints limiting experimental scale in multi-task reinforcement learning (RL) with linear temporal logic (LTL). To this end, it proposes a unified, high-performance benchmark suite. Methodologically, by leveraging the JAX framework to precompile symbolic task representations into static arrays, the approach enables an end-to-end just-in-time (JIT) compiled pipeline for training and evaluation, accompanied by standardized protocols and novel task sets. This framework achieves up to a 220-fold speedup, substantially enhancing statistical reliability. Furthermore, large-scale systematic evaluations reveal complementary strengths and weaknesses of existing methods regarding reasoning depth and scalability. Ultimately, this work establishes an efficient and reliable experimental foundation for advancing research in LTL-based multi-task RL.
📝 Abstract
Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to agents, and has been successfully adopted for training generalist multi-task policies. However, differences in implementations, task distributions, and evaluation protocols make existing methods difficult to compare, while high computational costs limit the scale and statistical reliability of experiments. We introduce Jaxolotl, a unified high-performance benchmark suite for multi-task LTL-RL to address these concerns. Jaxolotl provides a modular, end-to-end JAX implementation of six representative algorithms and four environments, together with newly curated task suites and a standardised, statistically robust evaluation protocol. By precompiling symbolic task representations into static arrays, Jaxolotl enables fully JIT-compiled training and evaluation, achieving end-to-end speedups of up to $220\times$ and supporting controlled comparisons at substantially greater experimental scale. We use this framework to systematically evaluate existing approaches, revealing complementary strengths and limitations: general methods capable of non-myopic reasoning struggle as the number of propositions grows, while methods with stronger scaling rely on environment-specific assumptions and suffer from myopia.