Can LLMs Translate Human Instructions into a Reinforcement Learning Agent's Internal Emergent Symbolic Representation?

📅 2025-10-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates whether large language models (LLMs) can accurately map natural language instructions to the emergent internal symbolic representations of hierarchical reinforcement learning (HRL) agents—thereby assessing the alignment between linguistic input and agent-internal structured semantics. Methodologically, it leverages HRL to generate interpretable, compositional internal symbols and establishes a standardized translation-and-verification evaluation framework across multiple granularity levels of symbol segmentation and varying task complexities. We systematically evaluate leading LLMs—including GPT, Claude, DeepSeek, and Grok—on cross-modal instruction-to-symbol translation. Results reveal that while LLMs exhibit baseline mapping capability, performance degrades markedly under fine-grained symbol segmentation and high-complexity tasks, exposing a fundamental limitation in precise semantic-to-structural alignment. This study provides the first empirical benchmark and diagnostic insights for language–agent representation alignment, highlighting critical gaps in grounding natural language in hierarchical, agent-internal symbolic structures.

Technology Category

Natural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: Symbolic RepresentationsMachine Learning: Large Multimodal Models (LMMs)

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Emergent symbolic representations are critical for enabling developmental learning agents to plan and generalize across tasks. In this work, we investigate whether large language models (LLMs) can translate human natural language instructions into the internal symbolic representations that emerge during hierarchical reinforcement learning. We apply a structured evaluation framework to measure the translation performance of commonly seen LLMs -- GPT, Claude, Deepseek and Grok -- across different internal symbolic partitions generated by a hierarchical reinforcement learning algorithm in the Ant Maze and Ant Fall environments. Our findings reveal that although LLMs demonstrate some ability to translate natural language into a symbolic representation of the environment dynamics, their performance is highly sensitive to partition granularity and task complexity. The results expose limitations in current LLMs capacity for representation alignment, highlighting the need for further research on robust alignment between language and internal agent representations.
Problem

Research questions and friction points this paper is trying to address.

LLMs translate human instructions to RL agent representations
Evaluate translation performance across symbolic partition granularities
Assess LLM sensitivity to task complexity and representation alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLMs translate human instructions into symbolic representations
Evaluation framework measures translation across symbolic partitions
Performance sensitive to partition granularity and task complexity
🔎 Similar Papers
No similar papers found.