What, When, and How: Audio Description as Constrained Global Optimization

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing automatic audio description methods, which fail to decouple content selection, temporal scheduling, and text generation, thereby yielding suboptimal decisions. We formulate this task as a constrained global optimization problem and propose a hybrid framework integrating large language models (LLMs) with mixed-integer linear programming (MILP). Specifically, the LLM performs visual element grounding and saliency estimation, while the MILP jointly solves for the globally optimal strategy of what, when, and how to describe under spatiotemporal constraints. Experimental results demonstrate that our approach establishes new state-of-the-art performance on the REFRAMED benchmark in terms of narrative question answering and temporal metrics, significantly improving the appropriateness of description placement.
📝 Abstract
Audio Description (AD) makes movies accessible to blind and visually impaired audiences by narrating visual information in gaps between dialogue. Existing automatic AD systems largely treat generation as a local video-to-text problem, assuming that the content to describe and its temporal location are already provided. Realistic AD instead requires coupled decisions about what visual information is narratively important, when it can be spoken without interfering with dialogue, and how it should be formulated to fit within the available time. We formalize AD generation as a constrained optimization problem over these three decisions. Our hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations. A mixed-integer linear program then jointly selects and schedules descriptions across a scene subject to temporal constraints. When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics. Ablations show that explicit temporal constraints drive gains in placement, while salience estimation controls how much narratively useful content is retained. Improvements are concentrated on temporal and narrative measures rather than n-gram overlap, although a significant gap to professional describers remains.
Problem

Research questions and friction points this paper is trying to address.

Audio Description
Constrained Optimization
Video-to-Text
Temporal Constraints
Narrative Salience
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio Description
Constrained Optimization
Mixed-Integer Linear Programming
Large Language Models
Salience Estimation
🔎 Similar Papers
2024-09-15arXiv.orgCitations: 0