π€ AI Summary
This study investigates whether large language models can transcend surface-level frequencies to identify latent sequential structures, thereby addressing conditional dependency failures in behavioral simulation. To this end, it proposes an evaluation framework that distinguishes inference from generation, decoupling distribution matching from rule adherence through controlled games such as Rock-Paper-Scissors and N-gram continuation tasks. Combined with Markov chain analysis, this approach systematically assesses the modelsβ reasoning and generative capacities regarding higher-order dependencies. The findings reveal the mechanisms by which long-context ineffectiveness and higher-order dependencies cause significant degradation in rule recovery. Furthermore, this work demonstrates that correct identification does not guarantee faithful simulation, highlighting that superficial behavioral fidelity may obscure erroneous underlying generative mechanisms.
π Abstract
Large language models (LLMs) are increasingly used as interactive agents and simulators, yet it remains unclear whether they can recover latent sequential structure beyond surface action frequencies. This distinction is critical for behavioral simulation, where actions are often shaped by prior context rather than marginal frequencies alone. We study this question using controlled two-player Rock--Paper--Scissors interactions and a one-player stochastic n-gram continuation task. Across these experiments, we test whether LLMs can identify latent strategies, follow simple Markov rules, and sustain higher-order conditional dependencies. Our framework separates distribution matching from conditional rule following. Results show that longer context does not improve identification, correct recognition does not ensure faithful simulation, and higher-order dependencies substantially degrade rule recovery. Apparent behavioral fidelity can therefore mask incorrect generative mechanisms.