🤖 AI Summary
This study addresses the challenges of provenance detection and watermarking for visual content generated by large language models (LLMs) via two distinct pathways: direct generation and code-based rendering. It proposes a production-centric conceptual framework that organizes watermarking mechanisms according to production stages and introduces explicit verification specifications. Through interface documentation review, boundary case analysis, and multimodal theoretical synthesis, the work systematically compares detection discrepancies between these two pathways across images, videos, and code. The primary contribution lies in formulating a conceptual research agenda grounded in existing methodologies and interface analysis, thereby offering a systematic theoretical perspective and delineating future research directions for multimodal AI content provenance.
📝 Abstract
AI systems create images and videos with image/video generation models or by writing code and graphics descriptions that are then rendered. These routes can produce similar visible artifacts but expose different representations, intervention points, and provenance evidence. We develop a production-centered framework that compares detection and watermarking across both routes. An explicit verification specification distinguishes passive inference, message recovery, and authenticated provenance. We organize image, video, source-code, and rendering-aware watermarks by production stage. We examine the different requirements of generated images and video, plots and SVG, programmable video, and agent-composed workflows. Documented Claude, OpenAI, and rendering-tool interfaces connect the framework to concrete systems. We pose ten scoped research questions on identifiability, observability, fair comparison across stages, recoverable payload, reconstruction, synchronization, composition, hybrid local contribution, and private production-event authentication. The result is a conceptual research agenda grounded in published methods, inspected interfaces, and elementary boundary examples. It reports no experiments and claims no new theorems; its appendix results are elementary calculations, and documentation and source inspection establish interfaces, not empirical robustness.