🤖 AI Summary
This study addresses the challenges of performance degradation, iterative instability, and unreliable candidate selection in multimodal autonomous research by proposing the MMResearch framework. This framework establishes a multimodal benchmark encompassing images and audio, employing a hierarchical memory mechanism to link media evidence with hypotheses. By integrating development set evaluation, it ensures cross-turn gain retention while suppressing non-target regressions. Furthermore, the framework incorporates LLM agents, code agent runtimes, and multimodal post-training techniques to systematically evaluate agents' capacity for continuous improvement under fixed budgets. Experimental results demonstrate that this approach yields accuracy improvements of 7.75 and 2.33 percentage points for Claude Opus 4.8 and GPT-5.6-sol, respectively.
📝 Abstract
Autonomous research seeks sustained model improvements through iterative experimentation and feedback. LLM agents show promise in automating machine learning and language-model post-training, but their ability to sustain multimodal improvement remains unclear. We introduce MMPostTrainBench, a benchmark spanning eight tasks in image, audio, video, and joint audio-video understanding and image-grounded software repair. Agents operate from a common base model within fixed budgets, using development feedback before independent evaluation of their submitted models. Evaluation covers target and non-target model outcomes, iterative model improvement and selection, and research integrity. Across all eight tasks, 52.1% of model--task means fall below the base, and evaluated submissions also exhibit non-target regressions. Model performance does not consistently improve across research iterations, and agents do not reliably select the best evaluated candidate for submission; final submissions trail that candidate by up to 5.38 percentage points. Extending autonomous research from text-only to multimodal tasks introduces additional sources of error in perception, cross-modal alignment, and temporal grounding. The observed regressions and selection gaps highlight the need to balance targeted improvements with non-target capability preservation and to retain gains across research iterations. These requirements motivate MMResearch, a multimodal research framework that connects media-grounded evidence to hypotheses and interventions, carries findings across rounds through hierarchical memory, and retains candidates using development evaluation. Added to existing code-agent runtimes, it improves submitted-model accuracy by up to 7.75 percentage points for Claude Opus 4.8 with Claude Code and 2.33 points for GPT-5.6-sol with Codex.