π€ AI Summary
This work addresses the lack of rigorous mathematical formalization in decomposition-based reasoning for multi-hop question answering. It introduces operad theory into the analysis of large language model reasoning by constructing a question operad $Q$, where question templates are treated as operations and sub-answers as inputs to be composed. The question-answering model is thereby interpreted as an algebra over this operad, providing a formal framework for problem decomposition and composition. Building on this foundation, the paper proposes operadic consistencyβa novel metric for evaluating the coherence of multi-step reasoning. Experiments across 12 large language models and 4 multi-hop question answering benchmarks demonstrate that this metric exhibits strong correlation with answer accuracy and significantly outperforms temperature-based self-consistency baselines.
π Abstract
Question decomposition, i.e. breaking a complex query into simpler sub-queries whose answers are composed to produce a final answer, is a widely used strategy for improving LLM reasoning, yet it currently lacks a rigorous mathematical foundation. In this paper, we propose operads, mathematical structures that model many-in, one-out operations and compositions thereof, as a natural framework for describing question decomposition. We define the questions operad $Q$, in which operations correspond to question templates and composition corresponds to substitution of sub-answers, and show how QA models can be interpreted as algebras over $Q$. Beyond reframing existing practice, this operadic perspective points toward new methods, in particular a notion of operadic consistency, which measures whether a QA model's answers agree across the partial collapses of a question decomposition tree. Empirical evaluation of operadic consistency is reported in our companion paper (Bottman, Liu, and Richardson, 2026), which finds it strongly correlated with accuracy across twelve LLMs and four multi-hop QA datasets and outperforming standard temperature-based self-consistency baselines. We argue that operads are the natural mathematical home for question decomposition, and that invariants such as operadic consistency open new directions for analyzing and improving the reliability of multi-step reasoning.