An Evolutionary Agentic Approach for Open-ended Image Quality Perception
该研究针对现有图像质量评估模型难以扩展到开放感知维度的问题,提出了一种基于多代理协作进化的无需训练框架PACE,通过构建具体的视觉问答协议来改进图像质量评分。
该研究针对现有图像质量评估模型难以扩展到开放感知维度的问题,提出了一种基于多代理协作进化的无需训练框架PACE,通过构建具体的视觉问答协议来改进图像质量评分。
This work addresses the semantic channel gap in multimodal large language models, which struggle to effectively leverage visual text for reasoning about task instructions embedded in images—such as those found in screenshots. The study is the first to explicitly identify and quantify this gap, introducing an end-to-end prompt-region grounding method that operates without OCR or region-level metadata. By embedding task instructions directly into images, the authors construct a Visualized Task Semantics (VTS) benchmark and align question-relevant regions with their semantic meanings through masked image modeling and typed semantic representations, thereby recovering clean visual features from occluded views. Evaluated across four benchmarks, the proposed approach improves VTS accuracy from 58.0 to 66.3 (+8.3 percentage points) while preserving performance on original text-interface tasks.
This work addresses the computational and memory bottlenecks hindering efficient, cinematic camera-motion video generation (e.g., bullet time, dolly zoom) from images on mobile devices. The authors propose a three-stage optimization strategy: distillation-guided pruning to obtain a compact model, combined diffusion distillation and reinforcement learning to compress the generator into four denoising steps, and mixed post-training quantization to reduce the model size below 1 GB. Built upon the Diffusion Transformer architecture, this approach achieves the first on-device image-to-video diffusion model capable of generating cinematic camera motions. Compared to the Wan 2.1 teacher model, it offers a 40× speedup in inference, enabling generation of 49-frame 480p videos on a MediaTek Dimensity 8400 chipset with only 20 seconds per denoising step and a peak memory footprint of 1.8 GB, significantly advancing practical high-quality video synthesis on mobile platforms.
This work addresses the challenges of object removal in real-world scenarios, where modeling non-local effects and handling inaccurate user-provided masks remain difficult, while existing diffusion models incur prohibitive computational costs that hinder deployment on interactive or edge devices. The authors propose OSOR, a novel approach that achieves stable training within a single-step diffusion framework for the first time. OSOR integrates an occupancy-guided discriminator, a lightweight alpha head, and a Semantic Anchor Validation Pipeline (SAVP) to jointly optimize generation quality under imperfect masks. Evaluated on the large-scale CORNE dataset and the AnimeEraseBench/TextEraseBench benchmarks, OSOR surpasses multi-step diffusion baselines in perceptual quality while accelerating inference by 4–30×, substantially improving both efficiency and robustness for interactive and edge-based applications.
This work addresses the challenge in streaming zero-shot voice conversion of disentangling speaker identity from linguistic content, where existing approaches struggle to simultaneously suppress speaker leakage and preserve vocal expressiveness. The study introduces speaker anonymization into this task for the first time, proposing a novel perturbation mechanism that explicitly protects speaker identity while retaining prosodic information. Furthermore, it designs a strictly causal, non-lookahead generative network to enable truly zero-latency streaming conversion. Without relying on future-frame buffering, the method effectively balances speaker privacy and speech utility, significantly enhancing both the naturalness and real-time performance of converted speech.
该研究针对现有图像质量评估模型难以扩展到开放感知维度的问题,提出了一种基于多代理协作进化的无需训练框架PACE,通过构建具体的视觉问答协议来改进图像质量评分。
This work addresses the semantic channel gap in multimodal large language models, which struggle to effectively leverage visual text for reasoning about task instructions embedded in images—such as those found in screenshots. The study is the first to explicitly identify and quantify this gap, introducing an end-to-end prompt-region grounding method that operates without OCR or region-level metadata. By embedding task instructions directly into images, the authors construct a Visualized Task Semantics (VTS) benchmark and align question-relevant regions with their semantic meanings through masked image modeling and typed semantic representations, thereby recovering clean visual features from occluded views. Evaluated across four benchmarks, the proposed approach improves VTS accuracy from 58.0 to 66.3 (+8.3 percentage points) while preserving performance on original text-interface tasks.
This work addresses the computational and memory bottlenecks hindering efficient, cinematic camera-motion video generation (e.g., bullet time, dolly zoom) from images on mobile devices. The authors propose a three-stage optimization strategy: distillation-guided pruning to obtain a compact model, combined diffusion distillation and reinforcement learning to compress the generator into four denoising steps, and mixed post-training quantization to reduce the model size below 1 GB. Built upon the Diffusion Transformer architecture, this approach achieves the first on-device image-to-video diffusion model capable of generating cinematic camera motions. Compared to the Wan 2.1 teacher model, it offers a 40× speedup in inference, enabling generation of 49-frame 480p videos on a MediaTek Dimensity 8400 chipset with only 20 seconds per denoising step and a peak memory footprint of 1.8 GB, significantly advancing practical high-quality video synthesis on mobile platforms.
This work addresses the challenges of object removal in real-world scenarios, where modeling non-local effects and handling inaccurate user-provided masks remain difficult, while existing diffusion models incur prohibitive computational costs that hinder deployment on interactive or edge devices. The authors propose OSOR, a novel approach that achieves stable training within a single-step diffusion framework for the first time. OSOR integrates an occupancy-guided discriminator, a lightweight alpha head, and a Semantic Anchor Validation Pipeline (SAVP) to jointly optimize generation quality under imperfect masks. Evaluated on the large-scale CORNE dataset and the AnimeEraseBench/TextEraseBench benchmarks, OSOR surpasses multi-step diffusion baselines in perceptual quality while accelerating inference by 4–30×, substantially improving both efficiency and robustness for interactive and edge-based applications.
This work addresses the challenge in streaming zero-shot voice conversion of disentangling speaker identity from linguistic content, where existing approaches struggle to simultaneously suppress speaker leakage and preserve vocal expressiveness. The study introduces speaker anonymization into this task for the first time, proposing a novel perturbation mechanism that explicitly protects speaker identity while retaining prosodic information. Furthermore, it designs a strictly causal, non-lookahead generative network to enable truly zero-latency streaming conversion. Without relying on future-frame buffering, the method effectively balances speaker privacy and speech utility, significantly enhancing both the naturalness and real-time performance of converted speech.