Institution profile

Lovart AI

Industry researchnorthamerica · us
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

Oct 01, 2026

This study addresses the challenge of managing long-range spatial context caused by growing memory sequences in long-video world models. To this end, it proposes SMI, the first framework that leverages understanding models to systematically manage spatial memory. Specifically, SMI pioneers the integration of the spatial reasoning capabilities of multimodal large language models into world model memory management. It achieves efficient memory optimization through four synergistic atomic operations: spatial clustering, intra-cluster sparsification, action-aware retrieval, and reliability filtering. Experimental results demonstrate that SMI significantly improves memory sparsity, generation stability, and spatial consistency across multiple benchmarks and backbone architectures.

0 citationsRead paper

Post-Training Leaves Behavioral Shadows on Unrelated Decisions

Sep 24, 2026

This study addresses the limitation of existing knowledge transfer methods, which typically rely on target data or teacher parameters and thus struggle to achieve efficient cross-task capability transfer. To overcome this, we propose an active task-free distillation paradigm that leverages a common ancestor model to identify nearly unrelated word pairs for constructing single-word prompts. Through unsupervised knowledge distillation, this approach uncovers "behavioral shadows" within irrelevant texts, enabling the transmission of model capabilities using only isolated words under fully decoupled conditions. Crucially, the proposed method requires neither target data nor teacher parameters. Empirical results demonstrate significant performance improvements in tasks such as code generation and logical reasoning. Furthermore, the approach exhibits strong scalability across model generations and scales, offering a novel pathway for transferring model capabilities.

0 citationsRead paper

VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers

May 17, 2026

Existing video style transfer methods are often hindered by the scarcity of large-scale triplet data and effective modeling paradigms, leading to temporal inconsistency, fragile handling of occlusions, and flickering artifacts. To address these limitations, this work introduces VISTA-1000, the first large-scale synthetic dataset with aligned style, content, and motion, encompassing 1,000 distinct artistic styles. Building upon this dataset, we propose a context-aware transfer framework based on diffusion Transformers, augmented with a lightweight style adapter for robust style representation. By integrating joint modeling with a disentanglement strategy, our approach significantly outperforms existing methods in terms of style fidelity, temporal coherence, and content preservation, effectively suppressing flickering and drift artifacts.

0 citationsRead paper

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration

May 17, 2026

This work addresses the challenges of identity drift, background inconsistency, and semantic degradation in long-form stylized or actor-replaced cinematic content, which arise from frequent shot transitions and viewpoint changes. To tackle these issues, the authors propose a multi-agent collaborative framework that leverages a scene-level JSON script as a semantic backbone, integrating dynamic visual reference anchors with a grid-based batch keyframe generation mechanism. Building upon shared contextual modeling in latent space, the framework enables joint keyframe synthesis and incorporates closed-loop verification with selective regeneration for rigorous identity and alignment auditing. Its core innovation lies in the novel Dual-Bridge Consistency mechanism, which effectively enforces long-term language–vision coherence across hundreds of shots. Evaluated on the SoapBench benchmark, the method significantly outperforms existing commercial video generation APIs, demonstrating superior narrative fidelity and long-range consistency.

0 citationsRead paper

Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs

Mar 15, 2026

This work addresses the underexplored potential of large language models (LLMs) in symbolic visual representation by proposing SVE-ASCII, a framework that systematically investigates and enhances LLMs’ intrinsic ability to generate and comprehend visual content purely within textual space. Departing from conventional approaches that rely on external rendering or code execution, SVE-ASCII leverages a “Seed-and-Evolve” data synthesis strategy, context-aware style editing, and unified instruction tuning to jointly optimize text-to-ASCII art generation and ASCII-to-text understanding. The study introduces ASCIIArt-7K, a high-quality dataset, and ASCIIArt-Bench, a comprehensive benchmark. Experimental results demonstrate that training on generation substantially improves visual understanding performance, revealing a bidirectional enhancement between perception and generation in symbolic visual processing. All code, data, and models are publicly released.

0 citationsRead paper
Recent publications

Latest Papers

Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

Oct 01, 2026

This study addresses the challenge of managing long-range spatial context caused by growing memory sequences in long-video world models. To this end, it proposes SMI, the first framework that leverages understanding models to systematically manage spatial memory. Specifically, SMI pioneers the integration of the spatial reasoning capabilities of multimodal large language models into world model memory management. It achieves efficient memory optimization through four synergistic atomic operations: spatial clustering, intra-cluster sparsification, action-aware retrieval, and reliability filtering. Experimental results demonstrate that SMI significantly improves memory sparsity, generation stability, and spatial consistency across multiple benchmarks and backbone architectures.

0 citationsRead paper

Post-Training Leaves Behavioral Shadows on Unrelated Decisions

Sep 24, 2026

This study addresses the limitation of existing knowledge transfer methods, which typically rely on target data or teacher parameters and thus struggle to achieve efficient cross-task capability transfer. To overcome this, we propose an active task-free distillation paradigm that leverages a common ancestor model to identify nearly unrelated word pairs for constructing single-word prompts. Through unsupervised knowledge distillation, this approach uncovers "behavioral shadows" within irrelevant texts, enabling the transmission of model capabilities using only isolated words under fully decoupled conditions. Crucially, the proposed method requires neither target data nor teacher parameters. Empirical results demonstrate significant performance improvements in tasks such as code generation and logical reasoning. Furthermore, the approach exhibits strong scalability across model generations and scales, offering a novel pathway for transferring model capabilities.

0 citationsRead paper

VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers

May 17, 2026

Existing video style transfer methods are often hindered by the scarcity of large-scale triplet data and effective modeling paradigms, leading to temporal inconsistency, fragile handling of occlusions, and flickering artifacts. To address these limitations, this work introduces VISTA-1000, the first large-scale synthetic dataset with aligned style, content, and motion, encompassing 1,000 distinct artistic styles. Building upon this dataset, we propose a context-aware transfer framework based on diffusion Transformers, augmented with a lightweight style adapter for robust style representation. By integrating joint modeling with a disentanglement strategy, our approach significantly outperforms existing methods in terms of style fidelity, temporal coherence, and content preservation, effectively suppressing flickering and drift artifacts.

0 citationsRead paper

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration

May 17, 2026

This work addresses the challenges of identity drift, background inconsistency, and semantic degradation in long-form stylized or actor-replaced cinematic content, which arise from frequent shot transitions and viewpoint changes. To tackle these issues, the authors propose a multi-agent collaborative framework that leverages a scene-level JSON script as a semantic backbone, integrating dynamic visual reference anchors with a grid-based batch keyframe generation mechanism. Building upon shared contextual modeling in latent space, the framework enables joint keyframe synthesis and incorporates closed-loop verification with selective regeneration for rigorous identity and alignment auditing. Its core innovation lies in the novel Dual-Bridge Consistency mechanism, which effectively enforces long-term language–vision coherence across hundreds of shots. Evaluated on the SoapBench benchmark, the method significantly outperforms existing commercial video generation APIs, demonstrating superior narrative fidelity and long-range consistency.

0 citationsRead paper

Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs

Mar 15, 2026

This work addresses the underexplored potential of large language models (LLMs) in symbolic visual representation by proposing SVE-ASCII, a framework that systematically investigates and enhances LLMs’ intrinsic ability to generate and comprehend visual content purely within textual space. Departing from conventional approaches that rely on external rendering or code execution, SVE-ASCII leverages a “Seed-and-Evolve” data synthesis strategy, context-aware style editing, and unified instruction tuning to jointly optimize text-to-ASCII art generation and ASCII-to-text understanding. The study introduces ASCIIArt-7K, a high-quality dataset, and ASCIIArt-Bench, a comprehensive benchmark. Experimental results demonstrate that training on generation substantially improves visual understanding performance, revealing a bidirectional enhancement between perception and generation in symbolic visual processing. All code, data, and models are publicly released.

0 citationsRead paper