🤖 AI Summary
This study addresses the performance instability of language models during structured domain adaptation, which often stems from insufficient representational compatibility. To this end, we propose an auxiliary supervision mechanism for representation alignment based on environment-derived tasks, revealing that semantically equivalent yet formally distinct inputs significantly influence model processing. Specifically, our method leverages multimodal chess representations (FEN and ASCII) alongside environment dynamics tasks to construct auxiliary training signals, achieving deep compatibility with pretrained models. Experimental results demonstrate that this mechanism substantially improves optimal move prediction accuracy, enhances cross-representational transferability, and effectively elevates the quality of open-ended commentary generation. Overall, this work establishes a novel paradigm for structured domain adaptation in language models.
📝 Abstract
Language models exhibit strong reasoning capabilities, yet adapting them to structured domains remains challenging and can yield inconsistent outcomes. We identify representation compatibility, the extent to which a model effectively processes a representation for a structured task, as a key factor in adaptation. We study this in chess, which provides a controlled testbed with precise semantics, computable optimal actions, and multiple state representations, including a symbolic encoding (FEN) and a spatial format (ASCII). We find that models often process semantically equivalent inputs substantially differently, affecting both learning and generalization. Building on this observation, we propose representation-aligned auxiliary supervision, which uses environment-derived tasks expressed in compatible representations to improve adaptation to structured domains. Across models and representations, auxiliary supervision consistently improves optimal-move prediction relative to target-only training under identical target data. Tasks that expose environment dynamics provide larger and most consistent gains than surface-level or static supervision, while remaining competitive with substantially increasing the amount of target-task data. Moreover, ASCII-trained models transfer more effectively to FEN than FEN-trained models do to ASCII, even surpassing the FEN target-only baseline on FEN evaluation. The gains also extend beyond optimal-move prediction to open-ended, factually grounded commentary generation. Overall, our results show that auxiliary supervision in model-compatible representations can enable effective adaptation in structured domains.