Model Capability Dominates: Inference-Time Optimization Lessons from AIMO 3

📅 2026-03-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the effectiveness of inference-time optimization strategies for large language models on mathematical reasoning tasks, with a focus on International Mathematical Olympiad (IMO)-level problems. The central hypothesis is that reducing error correlation across multiple solution attempts can improve overall accuracy. We present the first large-scale empirical assessment of techniques such as diversified prompting, high-temperature sampling, and multi-model ensembling during the AIMO 3 competition. Our findings indicate that none of these interventions yield significant performance gains: high-temperature sampling alone suffices to decorrelate errors, while weakened prompts actually degrade single-attempt accuracy. Moreover, differences in base model capabilities exert an order-of-magnitude greater influence on performance than any inference-time optimization strategy examined.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Search and Optimization: Sampling/Simulation-based SearchConstraint Satisfaction and Optimization: Satisfiability Modulo Theories

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Majority voting over multiple LLM attempts improves mathematical reasoning, but correlated errors limit the effective sample size. A natural fix: assign structurally different reasoning strategies to different voters to decorrelate errors. We test this Diverse Prompt Mixer in the AIMO~3 competition: 3 models, 23+ experiments, and 50 IMO-level problems on a single H100 80 GB with a 5-hour limit. Every intervention fails. High-temperature sampling already decorrelates errors sufficiently; weaker prompt strategies reduce per-attempt accuracy more than they reduce correlation. Across a 17-point model capability gap and every inference-time optimization we tried, model capability dominates by an order of magnitude.
Problem

Research questions and friction points this paper is trying to address.

mathematical reasoning
error correlation
inference-time optimization
diverse prompting
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

inference-time optimization
model capability
diverse prompting
error decorrelation
mathematical reasoning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Natapong Nitarach