generate image captions

Designs and implements models and pipelines that produce textual descriptions for images and image regions, from concise global captions to dense, attribute-rich region descriptions and structured image–text annotations. These systems integrate techniques such as LLM-guided or retrieval-augmented captioning, zero-shot and knowledge-grounded generation, reward-driven or reinforcement-learning optimization to balance coverage and factual accuracy, and modality- or task-specific adaptations (e.g., thermal/IR inputs or physics-aware constraints).

generateimagecaptions

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.62
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions

Jul 09, 2024
YH
Yu-Guan Hsieh
🏛️ Apple | University of Washington | National Tsing Hua University

Existing image captioning datasets rely solely on unstructured text, failing to explicitly encode compositional structures and relational semantics among entities. To address this, we propose Graph-based Captioning (GBC), a novel paradigm where nodes represent entities, attributes, and relation phrases, while labeled edges explicitly model semantic connections—preserving linguistic flexibility while introducing hierarchical structure. We introduce the first graph-based annotation framework, enabling automatic construction of the large-scale GBC10M dataset (10 million samples). Moreover, we pioneer the use of graph structure as both a supervision signal in CLIP-style contrastive learning and an intermediate representation for text-to-image generation. Integrating multimodal large models, object detection, and graph modeling, our approach achieves significant improvements across VQA, referring expression comprehension (REC), and captioning benchmarks. Experiments demonstrate that graph-structured representations enhance both fidelity and fine-grained controllability in text-to-image synthesis. Code and the GBC10M dataset are publicly released.

Automating graph-based caption generationEnhancing image descriptions with graph structuresImproving multimodal model performance via GBC

FlexCap: Describe Anything in Images in Controllable Detail

Mar 18, 2024
DD
Debidatta Dwibedi
🏛️ Google Deepmind | Carnegie Mellon University

Existing vision-language models lack granularity control, making it difficult to generate image descriptions at user-specified levels of detail. To address this, we propose FlexCap—the first vision-language model supporting length-controllable, multi-granularity region captioning. Our key contributions are: (1) a novel length-conditioned region captioning paradigm; (2) a large-scale, multi-length weakly supervised region caption dataset, coupled with a region-localization-guided knowledge distillation strategy for efficient training; and (3) joint modeling of visual features and target caption length. Experiments demonstrate that FlexCap achieves state-of-the-art (SOTA) performance on the Visual Genome dense captioning task and establishes new SOTA results on zero-shot VQA benchmarks—including GQA and VQAv2. Moreover, FlexCap seamlessly supports diverse downstream applications such as image annotation, fine-grained attribute recognition, and vision-language dialogue.

Flexibility in DescriptionImage CaptioningVariable Detail Level

This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.

Addressing challenges in scalability, robustness, and cross-modal learningExploring integration of text, images, video, and audio for cross-modal understandingSurveying multimodal large language models' architectures and applications

Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions

Feb 20, 2024
AG
Akash Ghosh
🏛️ IIT Patna | Stanford University | Amazon AI

Large language models (LLMs) inherently lack native visual understanding, limiting their applicability in multimodal scenarios. This work presents a systematic survey of vision-language models (VLMs), introducing— for the first time—a three-tier taxonomy grounded in modal input/output capabilities: comprehension-only, generation-only, and full-modality VLMs. We unify analysis across architectural design, training data composition, robustness properties, and benchmark performance (e.g., VQAv2, COCO Caption). Through comprehensive literature review, architectural decomposition, and cross-benchmark evaluation, we analyze over 100 state-of-the-art works to construct a technology evolution map. Our key contributions are: (1) the first scalable, capability-aware classification and evaluation framework for VLMs; (2) a precise delineation of current performance boundaries; and (3) identification of three critical future directions—embodied intelligence, robust multimodal reasoning, and efficient scaling—establishing an authoritative reference for the VLM research community.

Classifying VLMs by capabilities in multimodal data processingIdentifying future research directions for vision-language advancementsSurveying current vision-language models' methodologies and limitations

What Makes for Good Image Captions?

May 01, 2024
DC
Delong Chen
🏛️ HKUST

This work addresses the lack of a unified, quantitative standard for evaluating image caption quality by proposing the first information-theoretic, three-dimensional evaluation framework—comprising informational sufficiency, redundancy minimization, and human interpretability. To operationalize this framework, we introduce Pyramid-based Captioning (PoCa), a novel caption generation method that fuses multi-granularity visual features and enforces local–global alignment, thereby achieving theoretically provable improvements in information efficiency. PoCa incorporates a weighted optimization objective and a cross-model/cross-dataset consistency validation mechanism to ensure robustness and generalizability. Extensive experiments on multiple mainstream benchmarks demonstrate significant gains in BLEU-4, CIDEr, and human evaluation scores. The approach exhibits both task adaptability and theoretical rigor, establishing a new paradigm for interpretable assessment and controllable generation of image captions.

Balancing information sufficiency, minimal redundancy, and human comprehensibilityEstablishing an information-theoretic framework for image captioningQuantitatively measuring and optimizing caption quality across diverse requirements

Latest Papers

What's happening recently
View more

Existing image captioning methods struggle to generate context-rich descriptions that incorporate object attributes, event contexts, and deep semantic information. This work proposes a hierarchical multimodal article retrieval mechanism that leverages structure-aware text weighting and multidimensional similarity computation—encompassing content-visual, visual-visual, and discourse-position alignments—to accurately retrieve relevant articles from an external news knowledge base. By synergistically integrating a vision-language model (VLM) with a large language model (LLM), the framework generates news image captions enriched with contextual depth. The approach overcomes the limitations of conventional unimodal retrieval strategies, significantly enhancing both semantic richness and factual consistency. It achieved a competitive fifth place on the private test set of the ACM Multimedia EVENTA 2025 OpenEvent-V1 challenge, attaining a composite score of 0.2824.

context-rich descriptionexternal knowledgeimage captioning

This work addresses the inefficiency of traditional neural architecture search (NAS), which relies heavily on manual design or brute-force trial-and-error. We propose, for the first time, leveraging large language models (LLMs) as end-to-end neural architecture designers that generate executable image captioning models under strict Net API contractual constraints. Methodologically, we build a prompt-driven NAS pipeline based on Qwen3-8B to jointly synthesize CNN-based encoders and LSTM/GRU/Transformer decoders—including their hyperparameters and training strategies—while integrating automated BLEU-4 evaluation and code correction. Our contributions include: (1) establishing the first LLM-driven, API-compliant NAS paradigm; (2) open-sourcing the extended LEMUR dataset; and (3) experimentally generating dozens of models, over 50% of which train successfully, achieving a peak BLEU-4 score of 32.7—demonstrating the critical role of API constraints in ensuring generation quality.

Addresses challenges like code hallucinations and API compliance through prompt engineering.Automates neural architecture search for image captioning models using LLMs.Generates runnable models under strict API constraints with CNN encoders and decoders.

This work addresses the limited semantic understanding and environmental interaction capabilities of agents in vision-language tasks by proposing a unified multimodal intelligence framework. The framework introduces three key innovations: a DETR-based mechanism for fusing grid and region visual features, a lightweight multi-input Transformer attention module (LTMI) enabling efficient visual dialogue with less than one-tenth the parameters of comparable models while maintaining performance, and a two-stage language-vision fusion decoding strategy to support embodied instruction execution. The resulting GRIT model achieves a favorable balance between accuracy and speed on image captioning and attains a state-of-the-art 8.37% success rate on unseen scenarios in the ALFRED dataset.

embodied AIimage captioninginstruction following

To address the challenges of high-resolution image synthesis and multimodal semantic understanding, this paper introduces VLM-RF, a vision-enhanced large language model. Methodologically, it pioneers a noise-aware learning algorithm and integrates a linear-path rectified flow (RF) mechanism with a cross-modal bidirectional tokenization strategy, enabling unified spatiotemporal feature embedding and hybrid sequence modeling across text, images, and video. The contributions are threefold: (1) substantial improvement in generation quality—image resolution and perceptual sharpness increase by 25%; (2) 20% reduction in computational overhead; and (3) consistent superiority over state-of-the-art diffusion models in both synthesis fidelity and cross-modal alignment. By unifying generative modeling and language understanding within a single scalable architecture, VLM-RF establishes a novel paradigm for efficient, high-fidelity multimodal generation.

Enhance high-resolution image synthesis using vision-augmented LLMsImprove generative performance under noisy input conditionsInterpret multimodal data via unified text-image-video tokenization

Hot Scholars

WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
CS

Cordelia Schmid

Research director INRIA
Computer visionobject recognitionvideo recognitionlearning
CH

Conghui He

Shanghai AI Laboratory
Data-centric AILLMDocument Intelligence
SJ

Sajid Javed

Assistant Professor, Khalifa University of Science and Technology, UAE
Computer VisionComputational Pathology