Score
Designs and implements models and pipelines that produce textual descriptions for images and image regions, from concise global captions to dense, attribute-rich region descriptions and structured image–text annotations. These systems integrate techniques such as LLM-guided or retrieval-augmented captioning, zero-shot and knowledge-grounded generation, reward-driven or reinforcement-learning optimization to balance coverage and factual accuracy, and modality- or task-specific adaptations (e.g., thermal/IR inputs or physics-aware constraints).
Existing image captioning datasets rely solely on unstructured text, failing to explicitly encode compositional structures and relational semantics among entities. To address this, we propose Graph-based Captioning (GBC), a novel paradigm where nodes represent entities, attributes, and relation phrases, while labeled edges explicitly model semantic connections—preserving linguistic flexibility while introducing hierarchical structure. We introduce the first graph-based annotation framework, enabling automatic construction of the large-scale GBC10M dataset (10 million samples). Moreover, we pioneer the use of graph structure as both a supervision signal in CLIP-style contrastive learning and an intermediate representation for text-to-image generation. Integrating multimodal large models, object detection, and graph modeling, our approach achieves significant improvements across VQA, referring expression comprehension (REC), and captioning benchmarks. Experiments demonstrate that graph-structured representations enhance both fidelity and fine-grained controllability in text-to-image synthesis. Code and the GBC10M dataset are publicly released.
Existing vision-language models lack granularity control, making it difficult to generate image descriptions at user-specified levels of detail. To address this, we propose FlexCap—the first vision-language model supporting length-controllable, multi-granularity region captioning. Our key contributions are: (1) a novel length-conditioned region captioning paradigm; (2) a large-scale, multi-length weakly supervised region caption dataset, coupled with a region-localization-guided knowledge distillation strategy for efficient training; and (3) joint modeling of visual features and target caption length. Experiments demonstrate that FlexCap achieves state-of-the-art (SOTA) performance on the Visual Genome dense captioning task and establishes new SOTA results on zero-shot VQA benchmarks—including GQA and VQAv2. Moreover, FlexCap seamlessly supports diverse downstream applications such as image annotation, fine-grained attribute recognition, and vision-language dialogue.
This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.
Large language models (LLMs) inherently lack native visual understanding, limiting their applicability in multimodal scenarios. This work presents a systematic survey of vision-language models (VLMs), introducing— for the first time—a three-tier taxonomy grounded in modal input/output capabilities: comprehension-only, generation-only, and full-modality VLMs. We unify analysis across architectural design, training data composition, robustness properties, and benchmark performance (e.g., VQAv2, COCO Caption). Through comprehensive literature review, architectural decomposition, and cross-benchmark evaluation, we analyze over 100 state-of-the-art works to construct a technology evolution map. Our key contributions are: (1) the first scalable, capability-aware classification and evaluation framework for VLMs; (2) a precise delineation of current performance boundaries; and (3) identification of three critical future directions—embodied intelligence, robust multimodal reasoning, and efficient scaling—establishing an authoritative reference for the VLM research community.
This work addresses the lack of a unified, quantitative standard for evaluating image caption quality by proposing the first information-theoretic, three-dimensional evaluation framework—comprising informational sufficiency, redundancy minimization, and human interpretability. To operationalize this framework, we introduce Pyramid-based Captioning (PoCa), a novel caption generation method that fuses multi-granularity visual features and enforces local–global alignment, thereby achieving theoretically provable improvements in information efficiency. PoCa incorporates a weighted optimization objective and a cross-model/cross-dataset consistency validation mechanism to ensure robustness and generalizability. Extensive experiments on multiple mainstream benchmarks demonstrate significant gains in BLEU-4, CIDEr, and human evaluation scores. The approach exhibits both task adaptability and theoretical rigor, establishing a new paradigm for interpretable assessment and controllable generation of image captions.
Existing image captioning methods struggle to generate context-rich descriptions that incorporate object attributes, event contexts, and deep semantic information. This work proposes a hierarchical multimodal article retrieval mechanism that leverages structure-aware text weighting and multidimensional similarity computation—encompassing content-visual, visual-visual, and discourse-position alignments—to accurately retrieve relevant articles from an external news knowledge base. By synergistically integrating a vision-language model (VLM) with a large language model (LLM), the framework generates news image captions enriched with contextual depth. The approach overcomes the limitations of conventional unimodal retrieval strategies, significantly enhancing both semantic richness and factual consistency. It achieved a competitive fifth place on the private test set of the ACM Multimedia EVENTA 2025 OpenEvent-V1 challenge, attaining a composite score of 0.2824.
This work addresses the inefficiency of traditional neural architecture search (NAS), which relies heavily on manual design or brute-force trial-and-error. We propose, for the first time, leveraging large language models (LLMs) as end-to-end neural architecture designers that generate executable image captioning models under strict Net API contractual constraints. Methodologically, we build a prompt-driven NAS pipeline based on Qwen3-8B to jointly synthesize CNN-based encoders and LSTM/GRU/Transformer decoders—including their hyperparameters and training strategies—while integrating automated BLEU-4 evaluation and code correction. Our contributions include: (1) establishing the first LLM-driven, API-compliant NAS paradigm; (2) open-sourcing the extended LEMUR dataset; and (3) experimentally generating dozens of models, over 50% of which train successfully, achieving a peak BLEU-4 score of 32.7—demonstrating the critical role of API constraints in ensuring generation quality.
This work addresses the limited semantic understanding and environmental interaction capabilities of agents in vision-language tasks by proposing a unified multimodal intelligence framework. The framework introduces three key innovations: a DETR-based mechanism for fusing grid and region visual features, a lightweight multi-input Transformer attention module (LTMI) enabling efficient visual dialogue with less than one-tenth the parameters of comparable models while maintaining performance, and a two-stage language-vision fusion decoding strategy to support embodied instruction execution. The resulting GRIT model achieves a favorable balance between accuracy and speed on image captioning and attains a state-of-the-art 8.37% success rate on unseen scenarios in the ALFRED dataset.
To address the challenges of high-resolution image synthesis and multimodal semantic understanding, this paper introduces VLM-RF, a vision-enhanced large language model. Methodologically, it pioneers a noise-aware learning algorithm and integrates a linear-path rectified flow (RF) mechanism with a cross-modal bidirectional tokenization strategy, enabling unified spatiotemporal feature embedding and hybrid sequence modeling across text, images, and video. The contributions are threefold: (1) substantial improvement in generation quality—image resolution and perceptual sharpness increase by 25%; (2) 20% reduction in computational overhead; and (3) consistent superiority over state-of-the-art diffusion models in both synthesis fidelity and cross-modal alignment. By unifying generative modeling and language understanding within a single scalable architecture, VLM-RF establishes a novel paradigm for efficient, high-fidelity multimodal generation.
本文提出一种数据驱动的选择方法,帮助选择适合文档语料库的RAG管道,通过评估文本和多模态管道来平衡检索性能与系统效率。