Do LLMs Understand Sequential Structure? A Controlled Study of Inference and Generation
This study investigates whether large language models can transcend surface-level frequencies to identify latent sequential structures, thereby addressing conditional dependency failures in behavioral simulation. To this end, it proposes an evaluation framework that distinguishes inference from generation, decoupling distribution matching from rule adherence through controlled games such as Rock-Paper-Scissors and N-gram continuation tasks. Combined with Markov chain analysis, this approach systematically assesses the models’ reasoning and generative capacities regarding higher-order dependencies. The findings reveal the mechanisms by which long-context ineffectiveness and higher-order dependencies cause significant degradation in rule recovery. Furthermore, this work demonstrates that correct identification does not guarantee faithful simulation, highlighting that superficial behavioral fidelity may obscure erroneous underlying generative mechanisms.