Score
Design and build reusable semantic templates that structure and encode information across multiple data modalities (e.g., text, images, audio) to serve as prompts or inputs for multimodal models. This work defines slot schemas and phrasing, maps semantic descriptors or behavioral cues to signals across modalities, and produces template variants for controlled, personality- or trait-oriented prompting and evaluation.
This work addresses the challenge of inconsistent semantic representations across multimodal data—such as images, videos, and text—by proposing a language-centric atomic propositional representation framework. The approach transforms observations from any modality into sets of atomic propositions, which are then mapped via a global semantic codebook into a unified, interpretable shared semantic space. This enables compositional expression ranging from fine-grained facts to high-level concepts and facilitates cross-modal reasoning. Experimental results demonstrate that the framework substantially enhances complex multimodal understanding, structured retrieval, and high-quality data curation in autonomous driving and open-world scenarios, offering strong advantages in interpretability, compositionality, and cross-modal alignment.
This study addresses the challenge that multimodal models struggle to reliably bind textual prompts (e.g., “image”) to their corresponding input modalities, resulting in inadequate source modality tracking. The work formally defines and empirically investigates the “source modality monitoring” problem for the first time, introducing an evaluation paradigm grounded in target-modality information retrieval. By integrating syntactic manipulation with semantic perturbation, the authors systematically assess binding mechanisms across eleven prominent vision-language models. Their findings reveal that semantic cues dominate the binding process when modality distributions exhibit significant divergence, consistently outweighing syntactic signals. These insights offer critical evidence and a novel perspective for enhancing the reliability and robustness of multimodal agents.
Existing 3D scene understanding approaches rely on fixed multimodal combinations, which struggle to dynamically adapt modality usage according to the input query, often introducing semantic noise and limiting reasoning capabilities. To address this, this work proposes SmartMage, a unified multimodal large language model that introduces two key innovations: Semantic-guided Multimodal Adaptive Routing (SMART) and Modality-Aware Gating of Experts (MAGE). These mechanisms jointly enable adaptive activation of relevant modalities and expert modules based on query semantics, text-modality alignment, and modality quality. SmartMage achieves state-of-the-art performance across five 3D scene understanding benchmarks and demonstrates strong results on RGB video tasks. Ablation studies using ScanFacet confirm the effectiveness of its semantic-modality matching strategy.
Visually impaired users and those with limited domain expertise struggle to associate data groupings with real-world meaning and interpret field semantics, posing significant accessibility challenges. Method: This paper introduces “Semantic Scaffolding,” a dual-modality assistive framework leveraging large language models (LLMs) to automatically extract domain knowledge, generate semantic binning and context-aware data highlighting, and dynamically align data structures with real-world concepts. Contribution/Results: The approach enables the first co-design of semantic guidance and visualization accessibility while preserving user interpretive autonomy. In a qualitative study with 15 visually impaired participants, the method significantly improved data comprehension speed and prompted critical reflection on semantic guidance mechanisms—demonstrating the feasibility of jointly optimizing explainability and user agency in accessible data analysis.
Current text-to-image models often sacrifice diversity to strictly adhere to input prompts, yielding outputs confined to a single visual interpretation. This work proposes a novel paradigm termed “semantic browsing,” which decouples semantic decision-making from pixel generation to introduce controlled diversity while preserving prompt fidelity. By leveraging rich semantic text representations, vision-language models (VLMs), and a customized agent-based workflow, the method explicitly guides structured and interpretable semantic variations. The resulting navigable design space ensures that each generated image variant corresponds to a distinct, meaningful semantic choice, substantially enhancing both output diversity and user controllability.
This study addresses the challenge that existing data exploration tools struggle to accurately interpret users’ analytical intent when expressed in unstructured forms within spatiotemporal datasets. To bridge this gap, the authors propose a multimodal query system integrating freehand sketching, natural language, and visual annotations. Central to their approach is the concept of “proxemic semantics,” which captures how users disambiguate references through the relative spatial arrangement of multimodal elements within a unified interaction space. The system employs a hybrid architecture combining geometric sketch matching with vision-language models (VLMs), enabling joint pattern matching and semantic constraint-based query parsing. A user study with 20 participants empirically validates the stability of proxemic semantics, offering both empirical grounding and design implications for multimodal data exploration interfaces.
Traditional lexical semantic representations struggle to capture implicit dimensions such as the scenes, atmospheres, and emotions evoked by words in specific contexts. This work proposes the Scene Abstraction framework, which introduces a structured representation for context-sensitive word meanings by integrating contextual scenes—comprising events, entities, and environments—with expressive profiles that include associated events, generalized attributes, and elicited emotions. Leveraging few-shot prompting with large language models, the framework extracts embodied semantics directly from contextual usage. Based on this approach, we construct the COCA-Scenes dataset and conduct human evaluations demonstrating that annotators achieve 82.4% accuracy in identifying induced scenes—a relative improvement of 11.8 percentage points over standard text embeddings. Furthermore, across three semantic dimensions, 86.4% of participants significantly preferred our method’s generated scene profiles over those produced by the ATOMIC baseline.
This study investigates how the alignment between semantic grouping and spatial layout enhances visual search efficiency in user interfaces. Building upon a computational rationality framework, it introduces a hierarchical task representation into visual search modeling for the first time, developing a cognitive model that simulates how users integrate semantic structure and visual cues to adapt to task constraints. Through an integrated approach combining computational cognitive modeling, eye-tracking, and semantic categorization experiments, the model successfully replicates human search durations and eye movement patterns. Two user studies further demonstrate that search efficiency significantly improves when semantic groupings align with spatial arrangements. This work provides a computationally grounded evaluation tool for interface design and elucidates the cognitive mechanisms underlying the interplay between semantic and spatial consistency in visual search.
This study addresses whether the outputs of small language models (SLMs) in psychometric tasks stem from genuine semantic reasoning or are primarily driven by artifacts of prompt formulation. The authors propose the first diagnostic framework capable of disentangling the influence of such prompt artifacts, systematically manipulating role framing, instructions, item content, and option labels while employing controlled experiments and variance decomposition techniques to quantify the relative contributions of semantic signals versus prompt-induced artifacts. Findings reveal that prompt artifacts frequently dominate model responses, substantially undermining their psychometric validity. The proposed framework not only effectively identifies these confounding influences but also offers a novel pathway for evaluating and enhancing the semantic comprehension capabilities of large language models.