Score
Designs, builds, or analyzes systems that retrieve and select relevant context items (text, images, or other modalities) to condition a pretrained model at inference time, enabling retrieval-based, train-free in-context learning and visual or multimodal in-context learning. This includes LM-based and corpus‑conditioned retrieval components and context modeling/selection methods that operate at corpus or million‑token scale and are evaluated against dense retrievers and long‑context generalization.
Large language models (LLMs) excel at complex contextual understanding but exhibit pronounced capability asymmetry—struggling to stably generate long, equally sophisticated texts. Method: We systematically establish a unified “context engineering” framework, proposing a four-dimensional taxonomy encompassing retrieval, generation, processing, and management. Based on a systematic review and architectural analysis of 1,300+ papers, we construct the first comprehensive context engineering technology map; identify the intrinsic mechanisms underlying the understanding–generation capability mismatch; and delineate architectural integration pathways for four key application paradigms: retrieval-augmented generation, memory modeling, tool integration, and multi-agent coordination. Contribution/Results: The work delivers a standardized conceptual framework, a strategic technology roadmap, and identified critical breakthrough directions—providing both theoretical foundations and practical guidance for developing advanced context-aware AI systems.
This work investigates the behavioral limits of in-context learning (ICL) with thousands of demonstrations in ultra-long-context language models. Through systematic experiments across multiple models (e.g., Llama, Qwen) and datasets, and employing controlled analytical techniques—including random shuffling, label-based grouping, and demonstration subsampling—we find that: (1) ICL robustness to input ordering significantly increases with context length; (2) clustering examples by label degrades performance; and (3) gains do not arise from joint encoding of multiple demonstrations. Key contributions include: the first empirical demonstration that ICL performance scales continuously with demonstration count up to several thousand in large-label-space tasks; superior effectiveness over fine-tuning under low-to-moderate data regimes; and non-negligible gains achievable without fully utilizing available context capacity—challenging prevailing assumptions about ICL mechanisms.
Existing studies lack a systematic characterization of in-context learning (ICL) mechanisms for regression tasks in large language models (LLMs), particularly regarding the trade-off between internal knowledge retrieval and context-based example learning. Method: The authors introduce the first regression-oriented ICL mechanism evaluation framework, integrating controlled regression datasets, attribution-aware prompt ablation, and quantitative attribution analysis across three mainstream LLMs. Contribution/Results: They empirically demonstrate that LLM regression behavior lies on a “retrieval–learning” continuum, with the dominant mode modulated significantly by task priors, example type, and information richness. Based on these findings, they propose task-aware prompting principles for regression. Results exhibit strong cross-model and cross-dataset robustness, providing both theoretical grounding and practical guidance for efficient regression prompt engineering.
This work addresses the limitations of existing approaches in complex image retrieval, which suffer from insufficient fine-grained contextual modeling and entangled optimization objectives, thereby constraining the performance of multimodal large language models. To overcome these challenges, the authors propose an automatic pipeline for constructing a fine-grained multimodal quintuple dataset and introduce a two-stage decoupled fine-tuning strategy: first enhancing contextual reasoning capabilities and subsequently refining retrieval alignment. The proposed method achieves substantial improvements over current state-of-the-art approaches across five complex image retrieval benchmarks. Notably, even with a lightweight backbone model under zero-shot settings, it attains leading performance, demonstrating the efficacy of fine-grained representation learning and staged optimization in multimodal retrieval tasks.
This study investigates the effectiveness of multimodal in-context learning (ICL) in large multimodal models (LMMs) for image captioning, specifically addressing how to optimally configure image-caption exemplars (ICEs) to improve performance. Method: We propose a dual-perspective framework: an *external analysis* systematically evaluates ICE quantity, image retrieval strategies, and caption assignment schemes; an *internal attention mechanism* introduces a novel attention metric, complemented by visualization and ablation studies to characterize model reasoning behavior and assess representational compression feasibility. Contributions/Results: (1) First quantitative characterization of how distinct ICE configurations modulate attention responses; (2) Identification of attention-level mechanisms underlying performance disparities among LMMs sharing the same architecture; (3) Derivation of transferable, generalizable ICE configuration principles; (4) Empirical validation of attention-guided optimization for ICL. Collectively, these findings provide both theoretical foundations and practical guidelines for efficient, interpretable multimodal ICL.
To address key bottlenecks in multimodal RAG systems—including inaccurate user intent understanding, monolithic retrieval strategies, and weak inappropriate response filtering—this paper proposes an end-to-end, multi-stage framework. First, it introduces an image-context-enhanced intent refinement module to improve query semantic accuracy. Second, it designs an intent-driven, context-aware query generation mechanism coupled with heterogeneous API collaborative retrieval. Third, it pioneers a dynamic, organization-aware three-tier joint filtering mechanism—operating over images, text, and multimodal representations—to enable fine-grained relevance and safety control. The method integrates multimodal large language models (MLLMs), cross-modal classifiers, policy-aware dynamic filtering, and API-integrated architecture. Extensive evaluation on multiple public benchmarks—including knowledge-intensive visual question answering (VQA) and safety-oriented datasets—as well as real-world data demonstrates consistent superiority over state-of-the-art methods, setting new records across several key metrics.
This work addresses the challenge of efficiently retrieving local examples for in-context learning on edge devices, where constraints on computation, memory, and data privacy limit conventional approaches. The authors propose CoRA, a novel framework that enables task-conditioned retrieval without fine-tuning, backpropagation, or queries to the target model. CoRA constructs a task-conditional representation space using a frozen encoder, aligns representations via closed-form ridge regression, and incorporates optimal low-rank compression theory with a two-stage streaming indexing algorithm to support efficient multimodal example retrieval. Experiments demonstrate that CoRA significantly improves retrieval efficiency across ten text and four multimodal benchmarks and is successfully deployed on a Raspberry Pi 5, showing compatibility with models such as Llama-3.2-1B.
This work addresses the lack of systematic evaluation of multimodal in-context learning (ICL) capabilities, which hinders the identification of bottlenecks in vision-language models’ joint reasoning and knowledge acquisition. To bridge this gap, we propose CLBench-V, the first benchmark that formally defines multimodal ICL and introduces a three-dimensional evaluation framework encompassing context localization, application of new information, and acquisition of novel knowledge. We develop an automated pipeline to curate and generate high-quality data spanning diverse domains, including scientific reasoning, finance, long-document understanding, spatial reasoning, and web-based visual question answering. Evaluating six state-of-the-art multimodal models on 3,443 instances reveals that overall performance remains limited, with the best model achieving only a score of 0.2847. Among them, InternVL3.5-30B-A3B excels in context localization and knowledge learning, while Qwen3.5-Plus demonstrates superior performance in applying newly provided information.
This work investigates the feasibility of performing in-context retrieval directly within language models operating on million-token-long contexts, addressing the failure of conventional dense retrieval methods under length generalization scenarios. The authors identify attention dilution as the primary cause of performance degradation in long-context retrieval and propose BlockSearch, a novel approach incorporating length-aware softmax reweighting and document-level sparse attention mechanisms. Evaluated on a 0.6B-parameter model, BlockSearch matches or exceeds state-of-the-art dense retrieval systems on standard benchmarks such as MS MARCO and Natural Questions, while achieving over three times higher scores on the LIMIT benchmark. Notably, it outperforms contemporary multi-stage attention (MSA) models with only one-seventh of their parameters, demonstrating both the effectiveness and efficiency of in-context retrieval at scale.
Current multimodal models struggle to learn local rules, procedural steps, and empirical patterns from instructional contexts—such as images, videos, or manuals—and generalize them to novel visual instances. To address this gap, this work introduces MMCL-Bench, the first systematic benchmark for evaluating multimodal in-context learning (MMCL) capabilities, comprising 102 tasks across three categories: rule application, procedure execution, and empirical generalization. These tasks require models to locate relevant evidence within multimodal contexts and perform reasoning accordingly. Through rigorous scoring criteria and ablation analyses, we uncover critical bottlenecks in state-of-the-art models across the full pipeline—from contextual anchoring and visual evidence extraction to reasoning and response generation. Experimental results show that even the best-performing models solve fewer than one-third of the tasks on average, highlighting MMCL as a significant and underdeveloped capability in contemporary multimodal AI systems.
This study addresses the confounding effects of metadata, structured representations, and retrieval mechanisms in current RAG systems, which often combine multiple context-augmentation strategies, obscuring their individual contributions to answer quality. Through controlled experiments across six benchmarks, four models, and five augmentation levels—totaling over 24,000 evaluations—the work reveals that increased contextual richness does not necessarily improve accuracy. It introduces the “tractability hierarchy” theory, emphasizing that context must align with model capacity. The findings demonstrate that most augmentation strategies actually degrade performance; however, when metadata and retrieval strategies are carefully matched to a model’s capabilities, smaller models can outperform state-of-the-art large models by up to 19 F1 points on specific tasks, challenging the prevailing RAG design paradigm centered on stacking metadata.